Robustness enhancement method and device based on adaptive visual emotion recognition system
By generating adversarial samples to conduct adversarial training on the adaptive visual emotion recognition system and adjusting the parameters after adding noise to the data, the problem of the visual emotion recognition model being vulnerable to attacks is solved, the robustness and accuracy of the system are enhanced, and user privacy and security are protected.
Patent Information
- Application Number
- CN202510612478.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-10-14
AI Technical Summary
Existing artificial intelligence visual emotion recognition models that have passed the Turing test are easily exploited by attackers, resulting in serious deviations in emotion recognition results, affecting system decision-making and user privacy security.
By obtaining multimodal data of the surrounding environment to generate adversarial samples, the adaptive visual emotion recognition system is trained adversarially, noise is added to the data, and the system parameters are adjusted according to the recognition results to enhance the robustness of the system.
It improves the system's adaptability and robustness in complex and changing environments, effectively resists various attacks, avoids the leakage of user sensitive information, and improves the accuracy and reliability of emotion recognition.
Smart Images

Figure CN120783089A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of multi-modal emotion recognition, and in particular to a robustness enhancement method and device based on an adaptive visual emotion recognition system. BACKGROUND
[0002] In related technologies, visual emotion recognition models passing the Turing test not only can accurately capture multi-modal signals such as facial expressions, behaviors, and visualized texts, but also can understand and analyze human emotional states in real time. Compared with traditional emotion recognition methods and artificial intelligence models, these models have significantly improved in terms of accuracy, sensitivity, and adaptability to complex situations in emotion classification.
[0003] However, in related technologies, the powerful performance of artificial intelligence models passing the Turing test is often exploited by attackers to mislead emotion judgments by inputting specific adversarial samples. In addition, visual models also face threats from image attack models such as black-box attacks and white-box attacks. Due to the powerful learning ability of artificial intelligence models passing the Turing test, the perturbations input by attackers can cause serious deviations in emotion recognition results, thereby affecting the decision-making and behavior of the system and threatening the privacy and security of users, which needs to be improved. SUMMARY
[0004] The present application provides a robustness enhancement method and device based on an adaptive visual emotion recognition system to solve the problem in related technologies that due to the powerful learning ability of artificial intelligence models passing the Turing test, the perturbations input by attackers can cause serious deviations in emotion recognition results, thereby affecting the decision-making and behavior of the system and threatening the privacy and security of users.
[0005] The first aspect embodiment of the present application provides a robustness enhancement method based on an adaptive visual emotion recognition system, including the following steps: obtaining first multi-modal data based on environmental information of a surrounding environment; generating adversarial samples according to the first multi-modal data to perform adversarial training on an adaptive visual emotion recognition system using the adversarial samples, obtaining a trained adaptive visual emotion recognition system; adding generated noise to the first multi-modal data to obtain second multi-modal data; inputting the second multi-modal data into the trained adaptive visual emotion recognition system for emotion recognition, and adaptively adjusting at least one system parameter of the trained adaptive visual emotion recognition system according to the recognition result to perform robustness enhancement on the trained adaptive visual emotion recognition system, so that the robustness-enhanced adaptive visual emotion recognition system meets a preset adaptability condition.
[0006] By the technical solution, the embodiment of the present application can generate an adversarial sample by acquiring first multi-modal data of the surrounding environment, perform adversarial training on the adaptive visual emotion recognition system, and enhance the resistance of the system to the adversarial sample. After adding noise to the data, emotion recognition is performed, and the system parameters are adjusted according to the recognition result, so that the system optimizes itself according to the actual situation, thereby improving the adaptability and robustness of the system in a complex and variable environment, effectively resisting various attacks, avoiding leakage of user sensitive information, and improving the accuracy and reliability of emotion recognition.
[0007] Optionally, in an embodiment of the present application, the environment information based on the surrounding environment is used to acquire the first multi-modal data, including: acquiring visual data of at least one of facial expressions, behaviors and visualized texts from the surrounding environment, and determining the first multi-modal data.
[0008] By the technical solution, the embodiment of the present application can select at least one visual data such as facial expressions, behaviors and visualized texts from the surrounding environment to determine the first multi-modal data. These data can intuitively reflect the emotional state of human beings, are widely sourced and easy to acquire, and are helpful to improve the accuracy and comprehensiveness of subsequent emotion recognition.
[0009] Optionally, in an embodiment of the present application, the first multi-modal data is used to generate an adversarial sample, and the adversarial sample is used to perform adversarial training on the adaptive visual emotion recognition system: a multi-modal disturbance sample containing visual data of at least one of facial expressions, behaviors and visualized texts is created by using a preset adversarial network; and the multi-modal disturbance sample is added to a training data set to perform the adversarial training.
[0010] By the technical solution, the embodiment of the present application can create a multi-modal disturbance sample containing various visual data by using an adversarial network, simulate possible adversarial attack situations, and enrich the diversity of training data. Adding these disturbance samples to the training data set for adversarial training can enable the adaptive visual emotion recognition system to learn how to deal with various potential attacks during the training process, enhance the robustness of the system to adversarial samples, improve the anti-interference ability of the system in the face of complex and variable actual application scenarios, effectively reduce emotion recognition errors caused by adversarial attacks, and ensure stable and accurate operation of the system.
[0011] Optionally, in one embodiment of the present application, the adaptive adjustment of at least one system parameter of the trained adaptive visual emotion recognition system according to the recognition result to enhance the robustness of the trained adaptive visual emotion recognition system so that the adaptive visual emotion recognition system with enhanced robustness meets a preset adaptability condition, including: obtaining a deviation value according to the recognition result; adjusting the at least one system parameter according to the deviation value to enhance the robustness of the trained adaptive visual emotion recognition system.
[0012] Through the above technical solution, the embodiment of the present application can obtain a deviation value based on the recognition results, providing a quantitative basis for adjusting system parameters, making the adjustment more targeted and scientific. Adjusting system parameters based on the deviation value can accurately optimize the trained adaptive visual emotion recognition system, allowing the system to adaptively adapt to different environments and data changes.
[0013] The second aspect of the present application provides a robustness enhancement device based on an adaptive visual emotion recognition system, including: an acquisition module for acquiring first multimodal data based on environmental information of the surrounding environment; a training module for generating adversarial samples based on the first multimodal data, so as to use the adversarial samples to perform adversarial training on the adaptive visual emotion recognition system to obtain a trained adaptive visual emotion recognition system; a noise adding module for adding generated noise to the first multimodal data to obtain second multimodal data; an adaptive adjustment module for inputting the second multimodal data into the trained adaptive visual emotion recognition system for emotion recognition, and adaptively adjusting at least one system parameter of the trained adaptive visual emotion recognition system according to the recognition result to enhance the robustness of the trained adaptive visual emotion recognition system, so that the adaptive visual emotion recognition system after robustness enhancement meets the preset adaptability conditions.
[0014] Through the above technical solution, the embodiment of the present application can generate adversarial samples by acquiring first multimodal data of the surrounding environment, perform adversarial training on the adaptive visual emotion recognition system, and enhance the system's resistance to adversarial samples. Emotion recognition is performed after adding noise to the data, and system parameters are adjusted based on the recognition results, allowing the system to optimize itself according to actual conditions, thereby improving the system's adaptability and robustness in complex and changing environments, effectively resisting various attacks, preventing the leakage of user sensitive information, and improving the accuracy and reliability of emotion recognition.
[0015] Optionally, in one embodiment of the present application, the acquisition module includes: acquiring at least one visual data of facial expressions, behaviors and visual text from the surrounding environment to determine the first multimodal data.
[0016] By the technical solution, the first multi-modal data can be determined from at least one visual data such as facial expression, behavior and visualized text selected from the surrounding environment, the data can intuitively reflect human emotional state, is widely sourced and easy to obtain, and is helpful to improve the accuracy and comprehensiveness of subsequent emotion recognition.
[0017] Optionally, in an embodiment of the present application, the training module comprises: a creating unit configured to create multi-modal perturbation samples containing at least one visual data of facial expression, behavior and visualized text by using a preset adversarial network; and a training unit configured to add the multi-modal perturbation samples into a training data set for the adversarial training.
[0018] By the technical solution, the multi-modal perturbation samples containing various visual data can be created by the adversarial network, the adversarial attack situation that can occur is simulated, and the diversity of the training data is enriched. The adversarial training is performed on the training data set by adding the perturbation samples, the adaptive visual emotion recognition system can learn how to deal with various potential attacks in the training process, the robustness of the system to the adversarial samples is enhanced, the anti-interference ability of the system in the face of complex and changeable actual application scenarios is improved, the emotion recognition error caused by the adversarial attack is effectively reduced, and the stable and accurate operation of the system is ensured.
[0019] Optionally, in an embodiment of the present application, the adaptive adjustment module comprises: an obtaining unit configured to obtain a deviation value according to the recognition result; and an adjusting unit configured to adjust the at least one system parameter according to the deviation value to enhance the robustness of the trained adaptive visual emotion recognition system.
[0020] By the technical solution, the deviation value can be obtained according to the recognition result, a quantitative basis is provided for the adjustment of the system parameter, the adjustment is more targeted and scientific. The system parameter is adjusted based on the deviation value, the trained adaptive visual emotion recognition system can be accurately optimized, and the system can be adaptively adapted to different environments and data changes.
[0021] The third aspect embodiment of the present application provides an electronic device, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor executes the program to implement the robustness enhancement method of the adaptive visual emotion recognition system as described in the above embodiments.
[0022] The fourth aspect embodiment of the present application provides a computer readable storage medium, the computer readable storage medium stores a computer program, and the program is executed by a processor to implement the robustness enhancement method of the adaptive visual emotion recognition system as described above.
[0023] The fifth aspect of the present application provides a computer program product, which is executed to implement a robustness enhancement method based on an adaptive visual emotion recognition system.
[0024] The embodiments of the present application can determine first multimodal data by acquiring at least one visual data such as facial expressions, behaviors, and visual text in the surrounding environment, thereby generating multimodal perturbation samples containing multiple visual data, and conducting adversarial training on the adaptive visual emotion recognition system, enriching the diversity of training data, enhancing the system's resistance and robustness to adversarial samples, and reducing recognition errors. After adding noise to the data, emotion recognition is performed, and the system parameters are adjusted based on the deviation value obtained from the recognition results, so that the system can optimize itself according to the actual situation, thereby being more adaptable in complex and changing environments, effectively resisting various attacks, avoiding the leakage of user sensitive information, and improving the accuracy and reliability of emotion recognition.
[0025] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0027] Figure 1 A schematic diagram of the system architecture of a robustness enhancement method based on an adaptive visual emotion recognition system provided in accordance with an embodiment of the present application;
[0028] Figure 2 A flowchart of a robustness enhancement method based on an adaptive visual emotion recognition system provided according to an embodiment of the present application;
[0029] Figure 3 This is an example diagram of parameter adjustment of the adaptive adjustment recognition system according to a specific embodiment of the present application;
[0030] Figure 4 Schematic diagram of a robustness enhancement device based on an adaptive visual emotion recognition system provided according to an embodiment of the present application;
[0031] Figure 5 This is a diagram illustrating an example structure of an electronic device provided according to an embodiment of the present application. DETAILED DESCRIPTION
[0032] Embodiments of the present application are described below in detail, examples of which are shown in the drawings, wherein the same or similar notations represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by reference to the drawings are exemplary and are intended to explain the present application, and cannot be understood as a limitation of the present application.
[0033] The robustness enhancement method and device based on an adaptive visual emotion recognition system of embodiments of the present application are described below with reference to the accompanying drawings. In view of the problems in the related art mentioned above, due to the strong learning ability of the artificial intelligence model passing the Turing test, the perturbation input by the attacker can cause serious deviation of the emotion recognition result, thereby affecting the decision and behavior of the system and threatening the privacy security of the user. The present application provides a robustness enhancement method based on an adaptive visual emotion recognition system, in which the system can be trained against by generating adversarial samples from the first multi-modal data of the surrounding environment, enhancing the resistance of the system to adversarial samples. After adding noise to the data, emotion recognition is performed, and the system parameters are adjusted according to the recognition result, so that the system optimizes itself according to the actual situation, thereby improving the adaptability and robustness of the system in complex and variable environments, effectively resisting various attacks, avoiding the leakage of sensitive information of the user, and improving the accuracy and reliability of emotion recognition. Thus, the problem in the related art that due to the strong learning ability of the artificial intelligence model passing the Turing test, the perturbation input by the attacker can cause serious deviation of the emotion recognition result, thereby affecting the decision and behavior of the system and threatening the privacy security of the user is solved.
[0034] It should be noted that the adaptive visual emotion recognition system (system) used in the present application refers to a visual emotion recognition system that has passed the Turing test and has an adaptive mechanism. The system is actually an artificial intelligence model that has passed the Turing test and has mature visual multi-modal emotion recognition capability. Specifically, based on the adaptive visual emotion recognition system, the robustness enhancement method based on the adaptive visual emotion recognition system provided by the present application has a system architecture block diagram as shown in Figure 1 which includes data acquisition, adversarial sample generation, noise addition, emotion recognition, and system parameter adjustment. Figure 1 The overall process and the relationship between the various parts can be clearly understood.
[0035] Specifically, Figure 2 A flowchart of the robustness enhancement method based on the adaptive visual emotion recognition system provided by the embodiments of the present application is shown in
[0036] As shown in Figure 2 the robustness enhancement method based on the adaptive visual emotion recognition system includes the following steps:
[0037] In step S201, first multi-modal data is acquired based on environmental information of the surrounding environment.
[0038] It can be understood that the quality and richness of the data of the first multi-modal data acquired based on the environmental information of the surrounding environment directly affect the performance of the entire system.
[0039] Optionally, in an embodiment of the present application, acquiring the first multi-modal data based on the environmental information of the surrounding environment comprises: acquiring visual data of at least one of facial expressions, behaviors and visualized texts from the surrounding environment, and determining the first multi-modal data.
[0040] Visual data is not a general term for all visual information, but a specific term for visual data that can reflect emotional characteristics in a scene. In actual application scenarios, environmental information is diverse, but only data that can be associated with emotions has value for an emotion recognition system. Screening out these effective data can ensure that the system focuses on key information and avoids being disturbed by irrelevant information, thereby improving data processing efficiency and emotion recognition accuracy.
[0041] Specifically, facial expressions, as one of the most intuitive external manifestations of human emotions, contain rich emotional information. For example, a smile often represents joy and satisfaction, while a frown may indicate confusion or dissatisfaction. Different combinations of facial muscle movements can convey a variety of complex emotional states. By capturing and analyzing these facial expressions, the system can obtain important emotional clues.
[0042] Behavior data cannot be ignored either. People's behavior is often closely related to emotions, and body movements, posture changes, eye contact, etc. can provide key information for emotion recognition. For example, when excited, there may be large-scale body movements, while when depressed, there may be a dejected posture. These behavioral details play a unique role in the emotion recognition process and help the system better understand the user's emotions.
[0043] Visualized text is also an important source of data. In today's era of digital information explosion, the text content published by people on various platforms, such as social media updates and chat records in instant messaging software, contains rich emotional information. Positive or negative words and emotionally expressive sentences can all serve as strong evidence for emotion recognition.
[0044] Through the above technical solution, the embodiment of the present application can obtain the first multimodal data from the surrounding environment as the basis for system operation. Visual data such as facial expressions, behaviors and visual texts are selected to determine the multimodal data, which are widely available and easy to obtain. Facial expressions, behaviors, and visual texts reflect emotions from different dimensions. Facial expressions intuitively show complex emotions, behavioral details assist in fully understanding user emotions, and visual texts provide rich emotional information. Screening visual data that reflects emotional characteristics can avoid interference from irrelevant information, improve data processing efficiency and the accuracy of emotion recognition, lay a solid foundation for the subsequent system to accurately identify emotions, and enhance the practicality and reliability of the system.
[0045] In step S202, adversarial samples are generated according to the first multimodal data, so as to perform adversarial training on the adaptive visual emotion recognition system using the adversarial samples to obtain a trained adaptive visual emotion recognition system.
[0046] It is understandable that adversarial training of the system with the help of adversarial samples can effectively improve the system's ability to respond to potential attacks and enhance the system's stability and accuracy.
[0047] Optionally, in one embodiment of the present application, an adversarial sample is generated based on the first multimodal data to perform adversarial training on the adaptive visual emotion recognition system using the adversarial sample: a multimodal perturbation sample containing at least one visual data of facial expressions, behaviors, and visual text is created using a preset adversarial network; and the multimodal perturbation sample is added to the training data set for adversarial training.
[0048] Specifically, using GANs (Generative Adversarial Networks) to create perturbation samples can simulate attack methods and patterns that might occur in real-world scenarios. For example, using facial expression data, GANs might slightly alter the position of key facial muscles, making an otherwise neutral expression appear to have a hint of anger or sadness. For behavioral data, they might adjust the amplitude, speed, or sequence of movements to create perturbation samples that are slightly different from the original behavior but sufficient to confuse the system. For visual text, they might replace keywords and adjust sentence structure to blur or alter the emotional tendencies conveyed by the text.
[0049] After generating multimodal perturbation samples, they are added to the training dataset for adversarial training. This allows the system to be exposed to various possible interference situations during the training phase, so that it can learn how to identify and resist these interferences, thereby improving the robustness of the system.
[0050] The embodiment of the application can create multi-modal perturbation samples containing various visual data through GAN, accurately simulate attack means in real scenarios, and create interference from multiple dimensions such as facial expressions, behaviors, and visual texts. Adding these samples to the training set for adversarial training enables the system to cope with various interferences in the training stage, enhances the response capability to potential attacks, strengthens the system stability and accuracy, and improves the system robustness.
[0051] In step S203, the generated noise is added to the first multi-modal data to obtain second multi-modal data.
[0052] It can be understood that adding noise to the data through differential privacy technology means dynamically generating noise by using machine learning or deep learning methods and adding the noise to the data through differential privacy methods. Machine learning and deep learning methods have shown strong capabilities in today's technical field, and this is no exception in the aspect of generating noise. With the help of machine learning algorithms, the system can dynamically generate noise with specific statistical characteristics based on the learning and analysis of a large amount of data according to the characteristics and distribution of the data. Deep learning can generate noise that is more in line with actual needs by constructing complex neural network models to deeply mine and understand data.
[0053] After generating the noise, the noise is added to the data through the differential privacy method. Differential privacy is a strict privacy protection technology, and its core idea is to blur the details of the data by adding appropriate noise during data publishing or use, thereby protecting the privacy information of the data subject. In the present application, the differential privacy method ensures that the process of adding noise can effectively interfere with the analysis and use of the original data by potential attackers, while maximizing the usability of the data for the emotion recognition system. By carefully designing the way and parameters of adding noise, it is difficult for attackers to infer the true situation of the original data even if they obtain the data after adding noise, avoiding the leakage of sensitive information of users. At the same time, these data with added noise can still provide valuable information for the adaptive visual emotion recognition system, and the system can learn how to accurately recognize emotions in a noisy environment during processing of these data with noise, further improving the robustness and adaptability of the system.
[0054] The embodiment of the present application can add generated noise to the first multi-modal data to obtain second multi-modal data, dynamically generate noise by using a machine learning or deep learning method in combination with a differential privacy technology, simulate various interferences in an actual scene, adapt the adaptive visual emotion recognition system to complex situations in a training and recognition process, and improve the robustness and emotion recognition capability of the system in a noisy environment. Meanwhile, the use of the differential privacy method ensures the protection of data privacy when adding noise and avoids the leakage of sensitive information of users, enhances the performance of the system, guarantees the safety of data, and improves the reliability and practicality of the system.
[0055] In step S204, the second multi-modal data is input into the trained adaptive visual emotion recognition system for emotion recognition, at least one system parameter of the trained adaptive visual emotion recognition system is adaptively adjusted according to the recognition result, the robustness of the trained adaptive visual emotion recognition system is enhanced, and the adaptive visual emotion recognition system after the robustness is enhanced meets a preset adaptability condition.
[0056] It should be noted that the adaptive adjustment of the recognition system parameter includes adjustment of the adversarial training parameter, adjustment of the prediction model parameter, and adjustment of the noise addition strategy.
[0057] In actual execution, the second multi-modal data after adding noise is input into the system for emotion recognition, the system analyzes and processes these data, and outputs an emotion recognition result. Subsequently, the system intelligently detects the recognition result, and when it is detected that the emotion recognition result has obvious errors, the system corrects the recognition result in combination with the feedback of the user, such as but not limited to correction of the recognition result through an interface or a voice instruction. For example, in an actual application scenario, the system recognizes the facial expression of a person as “angry”, but the user feeds back through the interface that the expression actually expresses “surprise”, and the system corrects the recognition result according to the feedback, so as to ensure that the recognition result is more consistent with the actual situation.
[0058] In the above manner, the embodiment itself can improve and enhance the robustness of the trained adaptive visual emotion recognition system to adapt to different emotion recognition environments.
[0059] Optionally, in an embodiment of the present application, at least one system parameter of the trained adaptive visual emotion recognition system is adaptively adjusted according to the recognition result to enhance the robustness of the trained adaptive visual emotion recognition system, and the adaptive visual emotion recognition system after the robustness is enhanced meets a preset adaptability condition, including: obtaining a bias value according to the recognition result; and adjusting at least one system parameter according to the bias value to enhance the robustness of the trained adaptive visual emotion recognition system.
[0060] Specifically, there is a deviation between the recognition result and the actual situation, and the system can make adaptive adjustments based on the deviation, such as Figure 3 shown.
[0061] When adjusting adversarial training parameters, the weight of different modalities in adversarial training should be adjusted based on the weight of the different modalities in the current scenario. For example, in an emotion recognition scenario based on a video conference, if facial expression data dominates the multimodal data, then during adversarial training, the weight of adversarial samples related to facial expressions should be increased accordingly, so that the system can better defend against adversarial attacks based on facial expressions.
[0062] When adjusting prediction model parameters, adjust the weight of each modality in multimodal fusion based on the proportion of the modalities in the current scenario. The importance of multimodal data such as facial expressions, behaviors, and visual text for emotion recognition varies across different scenarios. For example, in a social media comment analysis scenario, visual text data may contribute more significantly to emotion recognition. In this case, it is necessary to increase the weight of visual text data in multimodal fusion so that the prediction model relies more heavily on the information provided by visual text when making emotion judgments, thereby improving the accuracy of emotion recognition.
[0063] When adjusting the noise addition strategy, the noise shape and placement are adjusted based on the attack behavior and its characteristics. If the system detects a specific attack targeting facial expression data, such as modifying certain pixels in a facial image to interfere with emotion recognition, the noise addition strategy can be adjusted accordingly, adding specific noise patterns to key areas of the facial expression data to enhance the system's resistance to this attack.
[0064] The embodiment of the present application can perform emotion recognition by inputting the second multimodal data with added noise into the initial system, and improve the recognition accuracy by combining the user feedback to correct the result. According to the recognition result, the deviation value is obtained to adjust the system parameters, and optimization is carried out from three aspects: adversarial training parameters, prediction model parameters and noise addition strategy. In the adjustment of adversarial training parameters, the proportion of adversarial samples is adjusted according to the proportion of scene modalities to enhance the system's ability to resist key modal adversarial attacks; the prediction model parameter adjustment highlights important modalities according to the scene to improve recognition accuracy; the noise addition strategy is adjusted according to the attack characteristics to enhance the system's anti-attack ability in a targeted manner. Ultimately, a system with enhanced robustness is obtained, which enables it to stably and accurately recognize emotions in complex and changing environments, ensure user experience, and improve the reliability and practicality of the system.
[0065] According to the robustness enhancement method of the adaptive visual emotion recognition system provided in the embodiments of the present application, the first multi-modal data of the surrounding environment can be acquired to generate an adversarial sample, the adaptive visual emotion recognition system is adversarially trained, and the resistance of the system to the adversarial sample is enhanced. After adding noise to the data, emotion recognition is performed, and the system parameters are adjusted according to the recognition result, so that the system optimizes itself according to the actual situation, thereby improving the adaptability and robustness of the system in a complex and variable environment, effectively resisting various attacks, avoiding leakage of user sensitive information, and improving the accuracy and reliability of emotion recognition.
[0066] Secondly, the robustness enhancement device of the adaptive visual emotion recognition system according to the embodiments of the present application is described with reference to the accompanying drawings.
[0067] Figure 4 is a block schematic diagram of the robustness enhancement device of the adaptive visual emotion recognition system according to the embodiments of the present application.
[0068] As shown in Figure 4 , the robustness enhancement device 10 of the adaptive visual emotion recognition system includes an acquisition module 100, a training module 200, a noise adding module 300, and an adaptive adjustment module 400.
[0069] Specifically, the acquisition module 100 is configured to acquire first multi-modal data based on environmental information of a surrounding environment.
[0070] The training module 200 is configured to generate an adversarial sample according to the first multi-modal data, to adversarially train the adaptive visual emotion recognition system with the adversarial sample, and to obtain a trained adaptive visual emotion recognition system.
[0071] The noise adding module 300 adds the generated noise to the first multi-modal data to obtain second multi-modal data.
[0072] The adaptive adjustment module 400 inputs the second multi-modal data into the trained adaptive visual emotion recognition system for emotion recognition, and adaptively adjusts at least one system parameter of the trained adaptive visual emotion recognition system according to the recognition result, to robustly enhance the trained adaptive visual emotion recognition system, so that the robustly enhanced adaptive visual emotion recognition system meets a preset adaptability condition.
[0073] Optionally, in an embodiment of the present application, the acquisition module 100 includes: acquiring visual data of at least one of facial expressions, behaviors, and visualized text from the surrounding environment, and determining the first multi-modal data.
[0074] Optionally, in an embodiment of the present application, the training module 200 includes a creating unit and a training unit.
[0075] The creating unit is configured to create a multi-modal perturbation sample containing visual data of at least one of facial expressions, behaviors, and visualized text by using a preset adversarial network.
[0076] The training unit is configured to add the multi-modal perturbation sample into a training data set for adversarial training.
[0077] Optionally, in an embodiment of the present application, the adaptive adjustment module 400 comprises an obtaining unit and an adjustment unit.
[0078] The obtaining unit is configured to obtain a deviation value according to the recognition result.
[0079] The adjustment unit is configured to adjust at least one system parameter according to the deviation value to enhance the robustness of the adaptive visual emotion recognition system after training.
[0080] It should be noted that the foregoing explanation of the embodiment of the method for enhancing the robustness of the adaptive visual emotion recognition system is also applicable to the embodiment of the device for enhancing the robustness of the adaptive visual emotion recognition system, which will not be described herein again.
[0081] The device for enhancing the robustness of the adaptive visual emotion recognition system according to the embodiment of the present application can generate adversarial samples by obtaining the first multi-modal data of the surrounding environment, perform adversarial training on the adaptive visual emotion recognition system, and enhance the resistance of the system to the adversarial samples. The system parameters are adjusted according to the recognition result after adding noise to the data for emotion recognition, so that the system optimizes itself according to the actual situation, thereby improving the adaptability and robustness of the system in a complex and changeable environment, effectively resisting various attacks, avoiding the leakage of sensitive information of users, and improving the accuracy and reliability of emotion recognition.
[0082] Figure 5 The structure schematic diagram of the electronic device provided by the embodiment of the present application is shown in FIG. 5. The electronic device can comprise:
[0083] The memory 501, the processor 502, and the computer program stored in the memory 501 and executable on the processor 502.
[0084] The processor 502 implements the method for enhancing the robustness of the adaptive visual emotion recognition system provided in the foregoing embodiments when executing the program.
[0085] Further, the electronic device further comprises:
[0086] The communication interface 503 is configured to communicate between the memory 501 and the processor 502.
[0087] The memory 501 is configured to store the computer program executable on the processor 502.
[0088] The memory 501 can include a high-speed RAM memory and can also include a non-volatile memory, such as at least one disk memory.
[0089] If the memory 501, the processor 502 and the communication interface 503 are implemented independently, the communication interface 503, the memory 501 and the processor 502 can be connected to each other through a bus and complete communication between each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For convenience of representation, Figure 5 Only one thick line is used in the figure to represent that there is only one bus or only one type of bus.
[0090] Optionally, in a specific implementation, if the memory 501, the processor 502 and the communication interface 503 are integrated on a chip, the memory 501, the processor 502 and the communication interface 503 can complete communication between each other through an internal interface.
[0091] The processor 502 can be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement one or more embodiments of the present application.
[0092] The embodiment of the present application further provides a computer readable storage medium, which stores a computer program, and the program is executed by a processor to implement the robustness enhancement method based on the adaptive visual emotion recognition system as above.
[0093] The embodiment of the present application further provides a computer program product, which is executed to implement the robustness enhancement method based on the adaptive visual emotion recognition system as above.
[0094] In the description of the application, reference to "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" means that a particular feature, structure, material, or characteristic being described is included in at least one embodiment or example of the application. The appearances of the phrase in various places in the specification are not necessarily all referring to the same embodiment or example. Furthermore, the described specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples. In addition, the usage of "N" means at least two, for example, two, three or the like, unless explicitly stated otherwise.
[0095] Furthermore, the terms "first", "second", or the like, are used merely as a designation of certain elements or features of the application, and do not imply or connote relative importance or a specific order of precedence. Thus, features defined with "first", "second", etc. can include at least one of the features, either explicitly or implicitly.
[0096] Any process or method descriptions or blocks in flow charts or otherwise described herein represent embodiments of modules, segments, or portions of code which include one or more executable instructions for implementing specific logic functions or steps, and alternate implementations are possible. In some embodiments, the processes or methods described in flow charts or otherwise described herein are not necessarily performed in the order shown or discussed, including, for example, performing or depending from other operations or stages, in parallel, in reverse order, or in some other suitable manner. Blocks can also be skipped or performed more than once depending on the logic of the method or process.
[0097] The logic and / or steps represented in the flowcharts and / or described herein, for example, can be considered as a sequence of executable instructions stored in a computer readable medium, which can be executed by an instruction execution system, apparatus or device, such as a computer-based system, a processor-based system, or other system that can fetch the instructions from the instruction execution system, apparatus or device and execute the instructions, or a combination of the above. For the purposes of this specification, a "computer readable medium" can be any apparatus that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus or device. The computer readable medium can be a computer readable storage medium or a computer readable signal medium. The computer readable storage medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or a propagation medium. The computer readable signal medium can include, but is not limited to, a computer readable medium that facilitates transfer of the program from one place to another. A specific example of a computer readable medium is a non-transitory computer-readable storage medium. A specific example of a computer readable signal medium is a source or destination of the computer readable medium. Another specific example of a computer readable signal medium is a computer readable signal travelling through space. Thus, a computer readable medium can take many forms of hardware to carry out the program for use by or in connection with the instruction execution system, apparatus or device.
[0098] It should be understood that aspects of the application can be implemented in hardware, software, firmware or a combination thereof. In the above embodiments, the N steps or methods can be implemented in software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented in hardware and in another embodiment, the hardware can be implemented using any or a combination of the following technologies, which are each well known in the art: a discrete logic circuit(s) having logic gates for implementing logic functions upon an application of data signals, an application specific integrated circuit having appropriate combinational logic gates, a programmable gate array(s) (PGA), a field programmable gate array (FPGA), etc.
[0099] Those of skill in the art would understand that the steps of the methods carried out above can be carried out by program instructions executed by relevant hardware, and the program can be stored in a computer readable storage medium, and when the program is executed, it includes one or a combination of the steps of the method embodiments.
[0100] In addition, each of the functional units in the various embodiments of the present application can be integrated in one processing module, or each of the units can be physically present separately, or two or more units can be integrated in one module. The integrated module can be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer readable storage medium.
[0101] The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it should be understood that the above embodiments are exemplary and should not be construed as limiting the present application, and those skilled in the art can make changes, modifications, replacements and variations to the above embodiments within the scope of the present application.
Claims
1. A robustness enhancement method based on an adaptive visual emotion recognition system, characterized in that: The following steps are involved: Acquiring first multimodal data based on environmental information of the surrounding environment; generating adversarial samples according to the first multimodal data, and performing adversarial training on the adaptive visual emotion recognition system using the adversarial samples to obtain a trained adaptive visual emotion recognition system; adding generated noise to the first multimodal data to obtain second multimodal data; The second multimodal data is input into the trained adaptive visual emotion recognition system for emotion recognition, and at least one system parameter of the trained adaptive visual emotion recognition system is adaptively adjusted according to the recognition result to enhance the robustness of the trained adaptive visual emotion recognition system so that the adaptive visual emotion recognition system with enhanced robustness meets the preset adaptability condition.
2. The method according to claim 1, characterized in that The acquiring of first multimodal data based on environmental information of the surrounding environment includes: At least one visual data of facial expression, behavior and visual text is obtained from the surrounding environment to determine the first multimodal data.
3. The method according to claim 2, characterized in that Generating an adversarial sample based on the first multimodal data, and performing adversarial training on the adaptive visual emotion recognition system using the adversarial sample: Using a preset adversarial network to create a multimodal perturbation sample containing at least one visual data of facial expression, behavior, and visual text; The multimodal perturbation samples are added to a training dataset to perform the adversarial training.
4. The method according to claim 1, wherein Adaptively adjusting at least one system parameter of the trained adaptive visual emotion recognition system according to the recognition result to enhance the robustness of the trained adaptive visual emotion recognition system so that the robustness-enhanced adaptive visual emotion recognition system meets a preset adaptability condition, includes: Obtaining a deviation value according to the recognition result; The at least one system parameter is adjusted according to the deviation value to enhance the robustness of the trained adaptive visual emotion recognition system.
5. A robustness enhancement device based on an adaptive visual emotion recognition system, characterized in that: include: an acquisition module, configured to acquire first multimodal data based on environmental information of the surrounding environment; a training module, configured to generate adversarial samples based on the first multimodal data, and perform adversarial training on the adaptive visual emotion recognition system using the adversarial samples to obtain a trained adaptive visual emotion recognition system; adding a noise module to add generated noise to the first multimodal data to obtain second multimodal data; An adaptive adjustment module inputs the second multimodal data into the trained adaptive visual emotion recognition system for emotion recognition, and adaptively adjusts at least one system parameter of the trained adaptive visual emotion recognition system according to the recognition result, so as to enhance the robustness of the trained adaptive visual emotion recognition system, so that the adaptive visual emotion recognition system with enhanced robustness meets a preset adaptability condition.
6. The device according to claim 5, characterized in that The acquisition module includes: At least one visual data of facial expression, behavior and visual text is obtained from the surrounding environment to determine the first multimodal data.
7. The device according to claim 5, characterized in that The training module includes: A creation unit, configured to create a multimodal perturbation sample comprising at least one visual data of facial expressions, behaviors, and visual texts using a preset adversarial network; A training unit is used to add the multimodal perturbation samples to a training data set to perform the adversarial training.
8. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the robustness enhancement method based on the adaptive visual emotion recognition system according to any one of claims 1 to 4.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the robustness enhancement method based on the adaptive visual emotion recognition system according to any one of claims 1 to 4.
10. A computer program product comprising a computer program, characterized in that The computer program is executed to implement the robustness enhancement method based on the adaptive visual emotion recognition system according to any one of claims 1 to 4.
Citation Information
Patent Citations
Processing method and device for visual model
CN113792791A
Method, device and equipment for performing student facial expression recognition based on adversarial training network, and storage medium
CN114333024A
Multi-modal emotion recognition method based on spiking neural network and attention mechanism
CN118861805A
Noise enhancement multi-mode emotion recognition method based on auxiliary network
CN119903393A