An intelligent propaganda and education system fusing multi-modal data analysis

By integrating multimodal data analysis into an intelligent education system, and combining image and speech recognition technologies, the system enables real-time assessment and dynamic education of patients' emotions. This solves the problems of high repetition and difficulty in evaluating the effectiveness of education in existing technologies, thereby improving the reliability and efficiency of education.

CN121146703BActive Publication Date: 2026-06-12ZHONGKE RUNHE (HANGZHOU) INFORMATION TECHNOLOGY CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHONGKE RUNHE (HANGZHOU) INFORMATION TECHNOLOGY CO LTD
Filing Date
2025-09-09
Publication Date
2026-06-12

AI Technical Summary

Technical Problem

Existing technologies are highly repetitive and time-consuming in the nursing process, and it is difficult to accurately assess the effectiveness of education by analyzing patients' micro-expressions, eye tracking, and body movements in real time.

Method used

An intelligent education system integrating multimodal data analysis is adopted, including modules for image recognition, speech recognition, emotion recognition, and recognition result output. It combines facial images, motion images, and speech data for emotion analysis, uses multimodal data for emotion assessment, and processes education through speech.

Benefits of technology

This improved the reliability and efficiency of patient education, enabled dynamic assessment of patients' understanding of the education content, and enhanced the accuracy and safety of the education effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121146703B_ABST
    Figure CN121146703B_ABST
Patent Text Reader

Abstract

The application provides an intelligent propaganda system and method fusing multi-modal data analysis, and belongs to the technical field of robots, and specifically comprises: an image recognition module responsible for acquiring and processing facial images and action images of propaganda objects.A speech recognition module is responsible for acquiring and processing speech data of the propaganda objects.An emotion recognition module determines a fusion processing object according to the deviation of emotion recognition results in each propaganda object and the propaganda duration, and performs emotion analysis processing on the fusion processing object by using multi-modal data.An emotion recognition result output module outputs the results of emotion analysis processing of each propaganda object to nursing staff, thereby improving the reliability of propaganda processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of robotics technology, and in particular relates to an intelligent education system that integrates multimodal data analysis. Background Technology

[0002] In the nursing process, it is often necessary to educate patients so that they can understand the precautions for their condition. The existing technical solutions often involve education by nursing staff, which has technical problems such as needing to educate each bed, high repetition, taking up a lot of time and energy, difficulty in retaining paper records, and easy omission of key information.

[0003] To address the aforementioned technical problems, invention patent application CN202510115130.8, "Robot-Assisted Diagnosis and Treatment System and Application Method for Nuclear Medicine Wards," utilizes AI robots for health education, significantly reducing the difficulty of the education process. However, the above technical solution has the following technical issues:

[0004] Existing technical solutions often neglect to analyze patients' micro-expressions, eye tracking, and body movements in real time when conducting health education, thus failing to evaluate the effectiveness of the education. This makes it difficult for medical staff to accurately understand the patient's understanding of the education content, resulting in education outcomes that often fall short of requirements.

[0005] To address the aforementioned technical issues, this application provides an intelligent education method and system that integrates multimodal data analysis. Summary of the Invention

[0006] To achieve the objectives of this invention, the following technical solution is adopted:

[0007] Specifically, this invention provides an intelligent education system that integrates multimodal data analysis, including:

[0008] Image recognition module, speech recognition module, emotion recognition module, recognition result output module;

[0009] The image recognition module is responsible for acquiring and processing the facial and motion images of the speaker.

[0010] The speech recognition module is responsible for acquiring and processing the speech data of the audience.

[0011] The emotion recognition module determines the fusion processing object based on the deviation of the emotion recognition results in each speech recipient and the speech duration, and performs emotion analysis processing using multimodal data in the fusion processing object;

[0012] The recognition result output module outputs the results of the emotion analysis of each speaker and uploads them to the nursing staff.

[0013] Furthermore, it also includes a missionary coefficient, which is responsible for delivering missionary messages to the target audience via voice.

[0014] Furthermore, sentiment analysis is performed using multimodal data, specifically including:

[0015] The original input speech signal (sampling rate 16kHz) is processed by Short Time Fourier Transform (STFT) to extract Mel-Spectrogram features, which are then converted into Mel-Frequency Cepstral Coefficients (MFCCs), i.e., Mel-Spectral Features.

[0016] Input facial video frames, detect facial coordinate points (68 coordinates) through MTCNN, and extract dynamic expression features;

[0017] By utilizing the dynamic facial expression features and Mel spectrum features, and employing a combination of early fusion and late fusion, the results of emotion analysis processing are obtained.

[0018] Furthermore, it also includes an identity recognition module, which uses a combination of infrared and visible light feature fusion for face verification and voiceprint-assisted verification for identity recognition.

[0019] Furthermore, it also includes a face-following module, where the mobile device's panel automatically follows the patient as they watch educational content, ensuring the patient can clearly see the broadcast information.

[0020] Secondly, this application provides an intelligent education method that integrates multimodal data analysis, applied to the aforementioned intelligent education system that integrates multimodal data analysis, specifically including:

[0021] S1 uses the facial expression and voice data recognition processing results as a basis to determine the deviation of the emotion recognition results in the target audience. Based on the deviation and the presentation duration of the target audience, it determines the multimodal data fusion recognition objects in the target audience. Based on the deviation of the emotion recognition results in each fusion recognition object, it determines the recognition deviation type of the emotion recognition results in each light intensity range.

[0022] Based on the recognition deviation type and recognition data in each light intensity range, S2 determines whether it is necessary to use multimodal data for emotion recognition processing for all the subjects to be addressed, based on the distribution interval data of the subjects to be addressed and each fused recognition object, whether the subjects to be addressed need to use multimodal data for emotion recognition processing under the given presentation duration.

[0023] Furthermore, the deviation of the emotion recognition results includes the deviation between the emotion recognition results of facial expressions and the emotion recognition results of voice data in different audiences.

[0024] Furthermore, the method for determining the multimodal data fusion and identification object in the target audience is as follows:

[0025] Based on the aforementioned deviation, the speakers whose emotion recognition results based on facial expressions are inconsistent with those based on voice data are identified, and these speakers are designated as emotion recognition deviation objects.

[0026] Based on the presentation duration of the target audience and the data on the object of emotion recognition deviation, the multimodal data of the target audience is used to determine the object of fusion recognition.

[0027] Further, determine whether multimodal data is needed for emotion recognition processing, specifically including:

[0028] Based on the data of the target audience, determine the number of integrated identification objects in the target audience. Based on the composition data of the integrated identification objects in the target audience, determine the proportion of the integrated identification objects in the target audience and use it as the proportion of integrated objects.

[0029] Based on the distribution interval data between the object to be presented and the fused identification object, the number of fused identification objects preceding the object to be presented is determined;

[0030] Based on the number of previously identified fusion objects, the proportion of fusion objects, and the presentation duration of the presenter, it is determined whether the presenter needs to use multimodal data for emotion recognition processing.

[0031] The beneficial effects of this invention are as follows:

[0032] When patients are watching the educational content, the mobile device's panel automatically follows them, ensuring that they can clearly see the broadcast information, thereby improving the reliability of the educational process.

[0033] Voiceprints are "non-contact and unique," and can be combined with facial recognition (results of infrared and visible light fusion) to achieve two-factor authentication, improving security and convenience.

[0034] By using audio, video, and images, patients can gain a more intuitive understanding of the educational content. During the educational process, the system analyzes patients' micro-expressions, eye movements, and body language in real time to generate dynamic assessment reports, which help medical staff understand patients' comprehension of the educational content.

[0035] Based on the distribution interval data of the target audience and each fusion recognition object, it is determined whether the target audience needs to use multimodal data for emotion recognition processing under the given presentation duration. This enables the use of multimodal data for emotion recognition processing in target audiences with longer presentation durations, even when the number of fusion recognition objects is small or their distribution is relatively scattered. This improves the reliability of emotion recognition processing and also allows for faster recognition processing of light intensity ranges where there is a risk of recognition failure.

[0036] Other features and advantages will be set forth in the following description, and the objects and other advantages of the invention are realized and obtained through the structures particularly pointed out in the description and the drawings.

[0037] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0038] The above and other features and advantages of the present invention will become more apparent from a detailed description of exemplary embodiments thereof with reference to the accompanying drawings.

[0039] Figure 1 This is a framework diagram of an intelligent education and outreach system that integrates multimodal data analysis;

[0040] Figure 2 This is a flowchart of an intelligent education method that integrates multimodal data analysis;

[0041] Figure 3 This is a flowchart illustrating the method for determining the target object through the fusion of multimodal data from the target audience to be presented.

[0042] Figure 4 This is a flowchart of a method for determining the type of recognition deviation in the light intensity range of emotion recognition results. Detailed Implementation

[0043] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.

[0044] Example 1

[0045] like Figure 1 As shown, this application provides an intelligent education system that integrates multimodal data analysis, specifically including:

[0046] Image recognition module, speech recognition module, emotion recognition module, recognition result output module;

[0047] The image recognition module is responsible for acquiring and processing the facial and motion images of the speaker.

[0048] The speech recognition module is responsible for acquiring and processing the speech data of the audience.

[0049] The emotion recognition module determines the fusion processing object based on the deviation of the emotion recognition results in each speech recipient and the speech duration, and performs emotion analysis processing using multimodal data in the fusion processing object;

[0050] The recognition result output module outputs the results of the emotion analysis of each speaker and uploads them to the nursing staff.

[0051] Furthermore, it also includes a missionary coefficient, which is responsible for delivering missionary messages to the target audience via voice.

[0052] Furthermore, sentiment analysis is performed using multimodal data, specifically including:

[0053] The original input speech signal (sampling rate 16kHz) is processed by Short Time Fourier Transform (STFT) to extract Mel-Spectrogram features, which are then converted into Mel-Frequency Cepstral Coefficients (MFCCs), i.e., Mel-Spectral Features.

[0054] Input facial video frames, detect facial coordinate points (68 coordinates) through MTCNN, and extract dynamic expression features;

[0055] By utilizing the dynamic facial expression features and Mel spectrum features, and employing a combination of early fusion and late fusion, the results of emotion analysis processing are obtained.

[0056] In one possible embodiment, after obtaining the Mel spectrum features and dynamic expression features, strategy 1: observe the input pattern of the fusion module.

[0057] Early fusion: The input to the fusion module is the raw / low-level features of the multimodal dataset, rather than the decision results.

[0058] For example:

[0059] The speech modality input is "Mel spectrum features (200-dimensional)" and the facial expression modality input is "facial landmark coordinates (66-dimensional)". The two are concatenated to form a 266-dimensional feature and then input into the Transformer. At this point, the input to the fusion module (concatenation operation) is a low-level feature, which belongs to early fusion.

[0060] Late-stage fusion: The input to the fusion module is the processing results of each modality (such as classification probability and regression value).

[0061] For example, the speech modality is processed by LSTM and outputs "emotion classification probability (e.g., happy: 0.8, sad: 0.2)", and the facial expression modality is processed by CNN and outputs "emotion classification probability (e.g., happy: 0.7, sad: 0.3)". The two are weighted (e.g., speech weight 0.6, facial expression weight 0.4) to obtain the final result. At this time, the input of the fusion module (weighted layer) is the decision result, which belongs to late fusion.

[0062] Strategy 2: Analyze the "branching" structure of the model architecture

[0063] Early fusion: The model architecture is usually "single-branch" - multimodal features are merged early and processed by only one main model (such as Transformer, CNN).

[0064] For example, the user mentioned "joint input of speech and facial expression features into Transformer", which means that speech and facial expression features are first concatenated and then input into the same Transformer. This is a typical early fusion (without independent speech / facial expression branch models). Late fusion: The model architecture is "multi-branch + final fusion layer" - each modality has an independent processing branch (such as LSTM for the speech branch and CNN for the facial expression branch), and finally the results are merged through a simple layer (such as a fully connected layer or a weighted layer).

[0065] For example, the user mentioned "LSTM processing temporal behavior data and then weighting the decision." If LSTM processes speech temporal data and facial expression temporal data separately, obtains their respective decision scores, and then weights them, then it belongs to late fusion (there are independent branches).

[0066] Strategy 3: Determine whether fusion depends on "early interactions" between modalities.

[0067] Early fusion: It is necessary to enable early interaction of multimodal features through operations (such as concatenation, element-wise multiplication, cross attention, etc.). For example, cross attention mechanism can be used to make speech features pay attention to key regions of facial expression features (such as binding "anger features of speech" with "frowning features of facial expression"). This interaction occurs at the feature layer and belongs to early fusion.

[0068] Late-stage fusion: There is no interaction during the processing of each modality. The fusion is only performed at the final result level (e.g., "If the voice says 'happy' and the facial expression shows 'smiling', then the overall judgment is 'happy'"), without the need for early feature interaction.

[0069] Furthermore, it also includes an identity recognition module, which uses a combination of infrared and visible light feature fusion for face verification and voiceprint-assisted verification for identity recognition.

[0070] Infrared images (thermal radiation imaging) are unaffected by visible light conditions (strong light, backlight, nighttime) and can stably capture target outlines; visible light images retain rich texture details (such as facial features and clothing textures), but are easily affected by lighting conditions. The fusion of these two technologies enables "all-weather robust recognition." The core is to achieve complementary advantages through feature layer or decision layer fusion. First, an infrared + visible light fusion model is used for facial verification, followed by secondary confirmation using voiceprint features. This rapid verification shortens voiceprint segments to 1-2 seconds (e.g., when a user says "login system") and enables real-time extraction using lightweight models (such as Tiny-CNN). It also exhibits dialect robustness by incorporating multi-dialect speech data during training, decoupling voiceprint features from language content (focusing on pronunciation habits rather than semantics). Furthermore, it features dynamic template updates: after each user login, the registration template is fine-tuned using newly collected facial / voiceprint features (e.g., sliding average updates) to adapt to long-term changes (such as facial aging and voiceprint changes).

[0071] Furthermore, it also includes a face-following module, where the mobile device's panel automatically follows the patient as they watch educational content, ensuring the patient can clearly see the broadcast information.

[0072] Core technologies:

[0073] ① Target tracking method:

[0074] def Tracking strategy selection (scenario complexity):

[0075] If the lighting is stable and there are no obstructions:

[0076] Return to KCF (High-Speed ​​Kernel Correlation Filter)

[0077] elif (fast motion):

[0078] return SiamRPN (Siam Network Region Proposal)

[0079] else:

[0080] return ByteTrack (Multi-target association)

[0081] ② Key point tracking

[0082] Adaptive model:

[0083] A simplified 68-point to 5-point model (eyebrow corner / nose tip / mouth corner) is used for real-time tracking;

[0084] 3D pose estimation to compensate for head rotation (based on PnP algorithm).

[0085] Example 2

[0086] Secondly, such as Figure 2 As shown, this application provides an intelligent education method that integrates multimodal data analysis, applied to the aforementioned intelligent education system that integrates multimodal data analysis, specifically including:

[0087] Based on the recognition and processing results of facial expression and voice data, the deviation of the emotion recognition results in the target audience is determined. Based on the deviation and the presentation duration of the target audience, the fusion recognition objects of the multimodal data in the target audience are determined. Based on the deviation of the emotion recognition results in each fusion recognition object, the recognition deviation type of the emotion recognition results in each light intensity range is determined.

[0088] Furthermore, the deviation of the emotion recognition results includes the deviation between the emotion recognition results of facial expressions and the emotion recognition results of voice data in different audiences.

[0089] Understandably, the emotion recognition results include happiness, confusion, surprise, and disgust.

[0090] It should be noted that the target audience for the presentation refers to those who have not yet received a presentation.

[0091] Specifically, such as Figure 3 As shown, the method for determining the multimodal data fusion and identification object in the subject to be presented is as follows:

[0092] Based on the aforementioned deviation, the speakers whose emotion recognition results based on facial expressions are inconsistent with those based on voice data are identified, and these speakers are designated as emotion recognition deviation objects.

[0093] Based on the presentation duration of the target audience and the data on the object of emotion recognition deviation, the multimodal data of the target audience is used to determine the object of fusion recognition.

[0094] Understandably, when there are no objects with emotion recognition bias, there is no need to determine the objects to be identified through the fusion of multimodal data.

[0095] It should be noted that when there are objects with emotion recognition bias, it is necessary to determine whether the number of objects with emotion recognition bias meets the requirements. If the number of objects with emotion recognition bias does not meet the requirements, it means that the reliability of the emotion recognition results is not good. Therefore, based on this, the fusion recognition objects of multimodal data are determined by using the presentation duration and preset ratio.

[0096] It should be noted that when there is no object to be identified through fusion, the emotion recognition result of the speaker is determined by using multimodal data according to a preset time period. In other cases, the emotion recognition result is determined by using facial expressions.

[0097] Additionally, it can be understood that when the number of objects with emotion recognition deviations meets the requirements, the fusion recognition objects of multimodal data are determined based on the lecture duration and the second preset ratio.

[0098] In one possible specific embodiment, when the number of emotion recognition deviation objects is not less than 4, it is determined that the number of emotion recognition deviation objects does not meet the requirements, and the preset ratio is one-fifth. If the presentation time of the object to be presented is in the top one-fifth among the objects to be presented, it is determined that the object to be presented is a multimodal data fusion recognition object, and the second preset ratio is less than the preset ratio. In one possible embodiment, the second preset ratio is one-seventh.

[0099] In one possible specific embodiment, the emotion recognition results of the audience are obtained by using multimodal data according to a preset time period. Specifically, this includes: extracting the facial expression data and voice data of the most recent 10 seconds every 30 seconds to determine the emotion recognition results, and using facial expression data to determine the emotion recognition results at other times.

[0100] Specifically, such as Figure 4 As shown, the method for determining the type of recognition deviation in the light intensity range for the emotion recognition result is as follows:

[0101] Based on the deviation of the emotion recognition results of each fusion recognition object, the recognition time when the emotion recognition results of facial expressions of each fusion recognition object are inconsistent with the emotion recognition results of speech data is determined and taken as the emotion recognition deviation time.

[0102] Using the aforementioned emotion recognition deviation time data, determine the emotion recognition deviation time of the fused recognition object in each light intensity range;

[0103] Based on the moment data of emotion recognition deviation within the light intensity range, the type of recognition deviation within the light intensity range is determined.

[0104] It is understandable that the type of recognition deviation in the light intensity range is determined based on the proportion of times when emotion recognition deviation occurs in each light intensity range.

[0105] It should be noted that the recognition deviation types include severe deviation intervals, moderate deviation intervals, other deviation intervals, and accurate recognition intervals. In one possible embodiment, when the proportion of emotion recognition deviation moments within the light intensity interval is greater than 0.1, the light intensity interval is determined to be a severe deviation interval; when the proportion of emotion recognition deviation moments within the light intensity interval is between 0.05 and 0.1, the light intensity interval is determined to be a moderate deviation interval; when the proportion of emotion recognition deviation moments within the light intensity interval is between 0 and 0.05, the light intensity interval is determined to be an other deviation interval; and when there are no emotion recognition deviation moments, the recognition deviation type within the light intensity interval is determined to be an accurate recognition interval.

[0106] For example, the light intensity range is divided into intervals of 50 Lux illuminance.

[0107] Based on the identification deviation type and identification data in each light intensity range, when it is determined that it is not necessary to use multimodal data for emotion recognition processing for all the subjects to be addressed, the distribution interval data between the subjects to be addressed and each fused identification object is used as a basis to determine whether the subjects to be addressed need to use multimodal data for emotion recognition processing under the given presentation duration.

[0108] Furthermore, it was determined that it was unnecessary to use multimodal data for emotion recognition processing for all the audience members to be addressed, specifically including:

[0109] Based on the recognition deviation type of each light intensity range, the data of the severe deviation range and the light intensity range at the moment when emotion recognition deviation exists are determined;

[0110] Based on the recognition data in each light intensity range, the recognition time in each light intensity range is determined;

[0111] Based on the recognition time data in various light intensity ranges, the data in the severe deviation range, and the light intensity range where there is a deviation in emotion recognition, it is determined whether it is necessary to use multimodal data to perform emotion recognition processing for all the audience to be addressed.

[0112] Understandably, when there is a severe deviation interval, in order to determine the true state of the reliability of the emotion recognition results within the severe deviation interval, and to enable timely and effective judgment of the consistency between the emotion recognition results of facial expressions and the emotion recognition results of voice data during the time period corresponding to the severe deviation interval, multimodal data is used for emotion recognition processing on all the audience to be addressed.

[0113] Furthermore, when there is no severe deviation interval, it is necessary to further determine the number of light intensity intervals at which emotion recognition deviation occurs. When the number of light intensity intervals at which emotion recognition deviation occurs does not meet the requirements, in order to achieve the true state of the recognition reliability of the emotion recognition results obtained by relying on facial expressions within the light intensity intervals at which emotion recognition deviation occurs, it is determined that multimodal data will be used for emotion recognition processing for all the audience to be addressed.

[0114] Additionally, it is understandable that when the number of light intensity intervals where there is emotion recognition deviation meets the requirements, it is still necessary to determine the recognition time data in each light intensity interval. Specifically, when there are multiple light intensity intervals where the number of recognition times does not meet the requirements and there is emotion recognition deviation, it is determined that multimodal data will be used to process the emotion recognition for all the audience to be addressed.

[0115] It should be noted that when the number of recognition moments in the light intensity range where there is an emotion recognition deviation is less than the preset threshold for the number of recognition moments, the reliability of the emotion recognition result obtained by relying on facial expressions in the light intensity range where there is an emotion recognition deviation is difficult to determine due to the small number of recognition moments. Therefore, the number of recognition moments in the light intensity range where there is an emotion recognition deviation does not meet the requirements.

[0116] Furthermore, if there are no light intensity ranges where the number of multiple recognition moments does not meet the requirements for emotion recognition deviation, then it is determined that it is not necessary to use multimodal data for emotion recognition processing on all the audiences to be addressed.

[0117] In one possible embodiment, if the number of light intensity intervals where there is an emotion recognition deviation is not less than 4, then it is determined that the number of light intensity intervals where there is an emotion recognition deviation does not meet the requirement; if the number of recognition moments within the light intensity intervals where there is an emotion recognition deviation is less than 50, then it is determined that the number of recognition moments within the light intensity intervals where there is an emotion recognition deviation does not meet the requirement.

[0118] Further, determine whether multimodal data is needed for emotion recognition processing, specifically including:

[0119] Based on the data of the target audience, determine the number of integrated identification objects in the target audience. Based on the composition data of the integrated identification objects in the target audience, determine the proportion of the integrated identification objects in the target audience and use it as the proportion of integrated objects.

[0120] Based on the distribution interval data between the object to be presented and the fused recognition object, the number of objects to be presented, which is the interval between the object to be presented and the adjacent fused recognition object, is determined.

[0121] Based on the number of subjects to be addressed at the specified interval, the proportion of subjects to be integrated, and the duration of the presentations by the subjects to be addressed, it is determined whether the subjects to be addressed need to undergo emotion recognition processing using multimodal data.

[0122] It should be noted that when the proportion of the fused objects meets the requirements, that is, when the proportion of the fused objects is greater than the preset proportion threshold, then regardless of the presentation duration of the presenter, it is not necessary to use multimodal data for emotion recognition processing.

[0123] Furthermore, when the ratio of the fused objects does not meet the requirements, if the number of objects to be presented at the interval is less than the preset interval number threshold, it is determined that regardless of the presentation duration of the objects to be presented, it is not necessary to use multimodal data for emotion recognition processing.

[0124] Additionally, it can be understood that if the number of people to be addressed at the specified interval is not less than a preset interval number threshold, and the speaking duration of the people to be addressed exceeds a preset speaking duration threshold, then it is determined that multimodal data will be used for emotion recognition processing in the people to be addressed. However, if the speaking duration of the people to be addressed is not greater than the preset speaking duration threshold, then it is determined that multimodal data is not needed for emotion recognition processing in the people to be addressed.

[0125] Example 3

[0126] Optionally, the method for determining the multimodal data fusion identification object in the target audience is as follows:

[0127] Based on the aforementioned deviation, the speakers whose emotion recognition results based on facial expressions are inconsistent with those based on voice data are identified, and these speakers are designated as emotion recognition deviation objects.

[0128] The proportion of the object with the emotion recognition deviation in the audience is taken as the emotion recognition deviation ratio.

[0129] Based on the presentation duration of the target audience and the proportion of emotion recognition deviation, the multimodal data fusion recognition objects among the target audience are determined.

[0130] It is understandable that if the speaking time of the person to be addressed is within the pre-emotion recognition deviation ratio among the people to be addressed, then the person to be addressed is determined to be a multimodal data fusion recognition object.

[0131] Example 4

[0132] Optionally, it can be determined that it is not necessary to perform emotion recognition processing using multimodal data on all the audience members to be addressed, specifically including:

[0133] Based on the recognition deviation type of each light intensity range, the data of the severe deviation range and the light intensity range at the moment when emotion recognition deviation exists are determined;

[0134] Based on the recognition data in each light intensity range, the recognition time in each light intensity range is determined. Based on the number of recognition times in the light intensity range where there is emotional bias, the average number of recognition times in the light intensity range where there is emotional bias is determined.

[0135] Based on the average number of recognition times in the light intensity range where emotional bias exists, the data of the severe bias range, and the light intensity range where emotional recognition bias exists, it is determined whether it is necessary to use multimodal data for emotion recognition processing for all the audience to be addressed.

[0136] Understandably, when there is a severe deviation interval, in order to determine the true state of the reliability of the emotion recognition results within the severe deviation interval, and to enable timely and effective judgment of the consistency between the emotion recognition results of facial expressions and the emotion recognition results of voice data during the time period corresponding to the severe deviation interval, multimodal data is used for emotion recognition processing on all the audience to be addressed.

[0137] Furthermore, when there is no severe deviation interval, if the average number of recognition times in the light intensity interval where there is emotional deviation is not less than the preset threshold for the number of recognition times, then it is determined that it is not necessary to use multimodal data for emotion recognition processing on all the audience to be addressed.

[0138] Additionally, it should be noted that if the average number of recognition times in the light intensity intervals where emotional deviation exists is less than a preset threshold for the number of recognition times, then if the number of light intensity intervals where emotional deviation exists is greater than a preset threshold for the number of intervals, then it is determined that multimodal data will be used for emotion recognition processing on all the audience members to be addressed. Conversely, if the number of light intensity intervals where emotional deviation exists is not greater than the preset threshold for the number of intervals, then it is determined that multimodal data will not be used for emotion recognition processing on all the audience members to be addressed.

[0139] Real-time Example 5

[0140] Further, determine whether multimodal data is needed for emotion recognition processing, specifically including:

[0141] Based on the data of the target audience, determine the number of integrated identification objects in the target audience. Based on the composition data of the integrated identification objects in the target audience, determine the proportion of the integrated identification objects in the target audience and use it as the proportion of integrated objects.

[0142] Based on the distribution interval data between the object to be presented and the fused identification object, the number of fused identification objects preceding the object to be presented is determined;

[0143] Based on the number of previously identified fusion objects, the proportion of fusion objects, and the presentation duration of the presenter, it is determined whether the presenter needs to use multimodal data for emotion recognition processing.

[0144] It should be noted that when the proportion of the fused objects meets the requirements, that is, when the proportion of the fused objects is greater than the preset proportion threshold, then regardless of the presentation duration of the presenter, it is not necessary to use multimodal data for emotion recognition processing.

[0145] In one possible embodiment, when the ratio of the fused objects is greater than 0.3, it is determined that the ratio of the fused objects meets the requirements.

[0146] Furthermore, when the proportion of the fused objects does not meet the requirements, it is necessary to further determine the distribution interval between the fused recognition objects. When the distribution interval meets the requirements, that is, the distribution interval between different adjacent fused recognition objects, where the number of objects to be presented in the distribution interval between different adjacent fused recognition objects is less than the preset threshold for the number of objects to be presented, then regardless of the presentation duration of the objects to be presented, it is not necessary to use multimodal data for emotion recognition processing.

[0147] In one possible embodiment, when the number of subjects to be addressed in the distribution interval between different adjacent fusion recognition subjects is less than 3, then regardless of the speaking time of the subjects to be addressed, it is not necessary to use multimodal data for emotion recognition processing.

[0148] Additionally, it can be understood that when the distribution interval does not meet the requirements, the number of fused recognition objects before the subject to be presented is determined. When the number of fused recognition objects before the subject to be presented meets the requirements, that is, when the number of fused recognition objects before the subject to be presented is large, the number of objects for recognition processing of recognition deviation types in different light intensity ranges is already large. On this basis, regardless of the presentation duration of the subject to be presented, it is not necessary to use multimodal data for emotion recognition processing.

[0149] In one possible embodiment, if the number of fused identification objects preceding the subject to be presented is not less than 5, then the number of fused identification objects preceding the subject to be presented is determined to meet the requirement.

[0150] Furthermore, when the number of previously identified fusion objects does not meet the requirements, it is also necessary to determine the presentation duration of the presenting object. If the presentation duration of the presenting object is greater than a preset presentation duration threshold, it is determined that multimodal data will be used for emotion recognition processing in the presenting object. If the presentation duration of the presenting object is not greater than the preset presentation duration threshold, it is determined that multimodal data is not needed for emotion recognition processing in the presenting object.

[0151] In one possible specific embodiment, when the speaking duration of the person to be addressed is greater than 10 minutes, it is determined that multimodal data is used to identify and process emotions in the person to be addressed.

[0152] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, devices, and non-volatile computer storage media are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0153] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0154] The above description is merely one or more embodiments of this specification and is not intended to limit this specification. Various modifications and variations can be made to the one or more embodiments of this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of one or more embodiments of this specification should be included within the scope of the claims of this specification.

Claims

1. An intelligent education and outreach system integrating multimodal data analysis, characterized in that, Specifically, it includes: Image recognition module, speech recognition module, emotion recognition module, recognition result output module; The image recognition module is responsible for acquiring and processing the facial and motion images of the speaker. The speech recognition module is responsible for acquiring and processing the speech data of the audience. The emotion recognition module determines the fusion processing object based on the deviation of the emotion recognition results in each speech recipient and the speech duration, and performs emotion analysis processing using multimodal data in the fusion processing object; Based on the recognition and processing results of facial expression and voice data, the deviation of the emotion recognition results in the target audience is determined. Based on the deviation and the presentation duration of the target audience, the fusion recognition objects of the multimodal data in the target audience are determined. Based on the deviation of the emotion recognition results in each fusion recognition object, the recognition deviation type of the emotion recognition results in each light intensity range is determined. Based on the identification deviation type and identification data in each light intensity range, when it is determined that it is not necessary to use multimodal data for emotion recognition processing for all the subjects to be addressed, the distribution interval data between the subjects to be addressed and each fused identification object is used as a basis to determine whether the subjects to be addressed need to use multimodal data for emotion recognition processing under the given presentation duration. The method for determining the multimodal data fusion and identification objects in the subject to be presented is as follows: Based on the aforementioned deviation, the speakers whose emotion recognition results based on facial expressions are inconsistent with those based on voice data are identified, and these speakers are designated as emotion recognition deviation objects. Based on the presentation duration of the target audience and the data on objects with emotion recognition bias, the fusion and recognition objects of multimodal data in the target audience are determined; The recognition result output module outputs the results of the emotion analysis of each audience member and uploads them to the nursing staff; Determining whether multimodal data is needed for emotion recognition processing includes: Based on the data of the target audience, determine the number of integrated identification objects in the target audience. Based on the composition data of the integrated identification objects in the target audience, determine the proportion of the integrated identification objects in the target audience and use it as the proportion of integrated objects. Based on the distribution interval data between the object to be presented and the fused identification object, the number of fused identification objects preceding the object to be presented is determined; Based on the number of fused recognition objects before the subject to be addressed, the proportion of fused objects, and the presentation duration of the subject to be addressed, it is determined whether the subject to be addressed needs to use multimodal data for emotion recognition processing. When the proportion of the fusion objects meets the requirements, then regardless of the presentation duration of the presenter, it is not necessary to use multimodal data for emotion recognition processing. When the ratio of the fusion objects does not meet the requirements, if the number of objects to be presented at the interval is less than the preset interval number threshold, it is determined that no matter how long the presentation of the objects to be presented is, it is not necessary to use multimodal data for emotion recognition processing. If the number of recipients to be addressed at the specified interval is not less than a preset interval number threshold, and the speaking duration of the recipients to be addressed exceeds a preset speaking duration threshold, then it is determined that multimodal data will be used for emotion recognition processing in the recipients to be addressed. If the speaking duration of the recipients to be addressed is not greater than the preset speaking duration threshold, then it is determined that multimodal data is not required for emotion recognition processing in the recipients to be addressed.

2. The intelligent education system integrating multimodal data analysis as described in claim 1, characterized in that, It also includes the missionary coefficient, which is responsible for delivering missionary messages to the target audience via voice.

3. The intelligent education system integrating multimodal data analysis as described in claim 1, characterized in that, Sentiment analysis using multimodal data specifically includes: The original speech signal is input, and the Mel spectral features are extracted through short-time Fourier transform, then converted into Mel cepstral coefficients, i.e., Mel spectral features: Input facial video frames, detect facial coordinate points through MTCNN, and extract dynamic expression features; By utilizing the dynamic facial expression features and Mel spectrum features, and employing a combination of early fusion and late fusion, the results of emotion analysis processing are obtained.

4. The intelligent education system integrating multimodal data analysis as described in claim 1, characterized in that, It also includes an identity recognition module, which uses a combination of infrared and visible light feature fusion for face verification and voiceprint-assisted verification for identity recognition.

5. The intelligent education system integrating multimodal data analysis as described in claim 1, characterized in that, It also includes a face-following module, where the mobile device's panel automatically follows the patient as they watch the educational content, ensuring the patient can clearly see the broadcast information.

6. The intelligent education system integrating multimodal data analysis as described in claim 1, characterized in that, The deviation of the emotion recognition results includes the deviation between the emotion recognition results of facial expressions and the emotion recognition results of voice data in different audiences.

7. The intelligent education system integrating multimodal data analysis as described in claim 1, characterized in that, The emotion recognition results include happiness, confusion, surprise, and disgust.

Citation Information

Patent Citations

  • Robot-assisted diagnosis and treatment system for nuclear medicine ward and application method

    CN119772894A

  • Psychological intervention method, system and equipment based on multi-modal emotion recognition and medium

    CN119153039A

  • A teaching optimization method based on big data informationization

    CN119741175A

Cited By

  • A monitoring system for an education system and a method of monitoring data analysis

    CN122201608A