Data perception single-mode sampling method based on reinforcement learning and aggressive expression detection system
Through the data-aware single-modal sampling method based on reinforcement learning, the data input volume of the multimodal model is dynamically adjusted, the problem of modal imbalance is solved, the model performance is improved, and especially in tasks such as aggressive expression detection.
Patent Information
- Application Number
- CN202510206134.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-30
AI Technical Summary
During the training process, multimodal models are prone to modal imbalance problems, resulting in the model learning skew, ignoring the learning of weak modes, and ultimately failing to achieve optimal performance.
A single-modal sampling method based on reinforcement learning is proposed. By monitoring the cumulative modal difference score index of modal learning, the data input amount of each modal is dynamically adjusted, and the learning of the modal is balanced.
By dynamically adjusting the data input volume, the problem of modal imbalance is effectively alleviated and the performance of multimodal models is improved, especially in tasks such as aggressive expression detection and false news detection.
Smart Images

Figure CN120067862A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning, and particularly to a data-aware single-modal sampling method based on reinforcement learning and an aggressive expression detection system. Background Art
[0002] Multimodal models aim to train deep learning models through multimodal data to achieve higher accuracy than traditional single-modal models. Multimodality provides strong support for practical application scenarios such as aggressive expression detection and fake news detection. However, during the training process of multimodal models, a modality imbalance problem will occur. The main reason lies in the heterogeneity of modality data. Strong modalities are trained quickly, while weak modalities are trained slowly, resulting in a tilt in model learning, ignoring the learning of weak modalities, and ultimately failing to achieve the optimal performance of the model. Traditional rebalancing methods only focus on balancing the training speeds between different modalities by modifying the gradient magnitudes, such as increasing the gradient of weak modalities and decreasing the gradient of strong modalities, without considering solving the imbalance at the data input level of multimodal models. Therefore, a data-aware single-modal sampling method based on reinforcement learning is proposed to alleviate and improve the modality imbalance problem, thereby enhancing the model performance. Summary of the Invention
[0003] The purpose of the present invention is to provide a data-aware single-modal sampling method based on reinforcement learning and an aggressive expression detection system. The present invention proposes a cumulative modality difference score index for monitoring modality learning, and uses reinforcement learning to perceive the learning situation of modalities, dynamically change the data input volume of each modality, balance the learning of modalities, and improve the performance of multimodal models.
[0004] The technical solution for achieving the purpose of the present invention is as follows: In the first aspect, a data-aware single-modal sampling method based on reinforcement learning includes the following steps:
[0005] Step 1: Process the original samples into an original sample sequence. The original samples are aggressive expression paired graphic-text multimodal data entities, and construct a multimodal deep learning model;
[0006] Step 2: Input the original sequence into the multimodal deep learning model to obtain the aggressive expression prediction result and calculate the cumulative modality difference score;
[0007] Step 3: Construct a decision-making system for the reinforcement learning policy network, and define the agent, environment, state, action, and reward function;
[0008] Step 4: Input the cumulative modality difference score into the policy network to obtain the sampling data volume of each modality of the next-round aggressive expression data, and update the policy network according to the reward function;
[0009] Step 5: Calculate the loss based on the prediction results of the aggressive expression of the multimodal deep learning model, and update the parameters to train the model.
[0010] In a second aspect, the present invention also provides an aggressive expression detection system for implementing the data-aware unimodal sampling method based on reinforcement learning described in the first aspect. The system includes:
[0011] The first module processes the original samples into an original sample sequence. The original samples are aggressive expression paired graphic and text multimodal data entities, and constructs a multimodal deep learning model;
[0012] The second module inputs the original sequence into the multimodal deep learning model to obtain the prediction results of the aggressive expression and calculates the cumulative modal difference score;
[0013] The third module constructs a decision-making system for the reinforcement learning policy network, and defines the agent, environment, state, action, and reward function;
[0014] The fourth module inputs the cumulative modal difference score into the policy network to obtain the sampling data volume of each modality of the aggressive expression data in the next round, and updates the policy network according to the reward function;
[0015] The fifth module finally calculates the loss based on the prediction results of the aggressive expression of the multimodal deep learning model, and updates the parameters to train the model.
[0016] In a third aspect, the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps of the method described in the first aspect.
[0017] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps of the method described in the first aspect.
[0018] In a fifth aspect, the present invention provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the steps of the method described in the first aspect.
[0019] Compared with the prior art, the significant advantages of the present invention are as follows: The present invention proposes a novel multimodal model training method. Considering the unbalanced perception modalities from the perspective of model data input, a cumulative modal difference score is designed to monitor the learning state of the modalities, and reinforcement learning is used to perceive the learning process and map it into the data input volume of each modality in real time. It supplements the data input perspective ignored in the traditional multimodal rebalancing learning method. This system can be applied to many scenarios such as aggressive expression detection and fake news detection. Description of the Drawings
[0020] Figure 1 This is the overall flowchart of the data-aware unimodal sampling method based on reinforcement learning of the present invention.
[0021] Figure 2 This is the network framework diagram of the data-aware unimodal sampling method based on reinforcement learning.
[0022] Figure 3 This is the sub-flowchart of the cumulative modal difference score calculation step.
[0023] Figure 4 This is the sub-flowchart of the reinforcement learning policy network construction step. Specific implementation manner
[0024] Combined with Figures 1 to 4 , the present invention provides a data-aware unimodal sampling method based on reinforcement learning, including the following steps:
[0025] (1) Process the original samples into an original sample sequence. The original samples are aggressive expression paired graphic-text multimodal data entities, and construct a multimodal deep learning model.
[0026] Perform initialization of paired graphic-text or paired audio-video, and the specific form is:
[0027]
[0028] Define the size of the dataset D as n, and it contains m modalities; X (j) represents all the data of the j-th modality, represents the i-th data of the j-th modality; define y i ∈{0,1} c to represent the label of each data, where c represents the total number of categories.
[0029] Constructing the multimodal deep learning model includes constructing a backbone network and a classification network. Use the BERT model based on transformer and the Vision Transformer model as the backbone network for processing paired graphic-text data. The classification network is composed of the simplest linear layer and the non-linear activation function ReLU. Denote the backbone network and the classification network of the j-th modality as and g (j) .
[0030] (2) Input the original sequence into the multimodal deep learning model to obtain the prediction result of the aggressive expression and calculate the cumulative modal difference score.
[0031] ① Assume that in the t-th iteration, a mini-batch of data B (j) is obtained through random sampling, where n b represents the batch size. The specific form is:
[0032]
[0033] ② Input this small batch of data into the multi-modal deep learning model to obtain the prediction result. represents the feature of the j-th modality of the i-th sample, Then represents its prediction probability, and softmax represents the normalization function. The specific forms of its features and prediction probabilities are as follows:
[0034]
[0035]
[0036] ③ According to the prediction result of the multi-modal deep learning model, it is expected to define the modality difference score to effectively reflect the learning status of each modality and monitor the learning process of each modality. Therefore, this cumulative modality difference score should be positively correlated with the prediction accuracy of the model. Thus, the modality difference score is defined, where represents the modality difference score:
[0037]
[0038] Due to the instability of random batch sampling, the modality difference score is further generalized to the cumulative modality difference score
[0039]
[0040] (3) Construct the decision-making system of the reinforcement learning policy network, and define the necessary components of the agent, environment, state, action, and reward function.
[0041] ① Construct the agent of the reinforcement learning policy network. Similar to constructing the classification network of the multi-modal deep learning model, it is composed of the simplest linear layer and the non-linear activation function ReLU, ψ ω represents the reinforcement learning policy network, and the specific form is as follows, where FC represents the linear layer, ReLU represents the ReLU activation function, and Dim represents the dimension size of the input features:
[0042] ψ ω = {FC(Dim×256) → ReLU → FC(Dim×64) → ReLU → FC(64×Dim)}
[0043] ② Construct the environment of the reinforcement learning policy network. Define the training process of the entire multi-modal deep learning model as the environment of this reinforcement learning. The agent needs to perceive the learning status of each modality in the current modality and make corresponding actions to change the environment.
[0044] ③Construct the state of the reinforcement learning policy network. Define the state space S based on the cumulative modal difference scores defined above. In the t-th iteration, the state of the current agent can be defined as an m-dimensional vector in the following specific form: Represents the cumulative modal difference score of the i-th modality:
[0045]
[0046] ④Construct the action of the reinforcement learning policy network. Define the action space Indicates that all vectors in the action space A are positive integer vectors of dimension m. In the t-th iteration, the action vector given by the current agent is in the following specific form:
[0047]
[0048] Where Represents the next sampling amount given by the reinforcement learning model for the j-th modality.
[0049] ⑤Construct the reward function of the reinforcement learning policy network. Given a state vector, the agent in the policy network needs to generate an action vector. The specific form is as follows:
[0050]
[0051] Where N B Represents the preset maximum data volume per batch, Represents the output of the reinforcement learning policy network, and round represents the floor function. Further define the reward function as follows:
[0052]
[0053] Where Represents the indicator function, r represents the reward function, and log represents the natural logarithm.
[0054] (4) Input the cumulative modal difference scores into the policy network to obtain the sampling data volume of each modality for the next round of aggressive expression data, and update the policy network according to the reward function.
[0055] ①Input the cumulative modal difference scores into the policy network to obtain the sampling data volume for the next round, whose form is the same as that of the action vector generated in (3).
[0056] ②According to the reward function, the gradient of the agent parameter ω can be calculated, and the policy network is continuously updated based on this. The specific form is as follows:
[0057]
[0058] (5) Calculate the loss based on the prediction results of the aggressive expression of the multi-modal deep learning model, and update the parameters to train the model.
[0059] Taking the cross-entropy loss as the loss function, the empirical risk minimization of the multi-modal deep learning model is expressed as follows, where L represents the loss function of the overall data set:
[0060]
[0061] Based on the same inventive concept, the present invention also proposes a data-aware single-modal sampling aggressive expression detection system based on reinforcement learning, which is used to implement the above-mentioned reinforcement learning data-aware single-modal sampling method and determine whether the expression content on the network contains aggression. The system includes:
[0062] The first module processes the original sample into an original sample sequence. The original sample is an aggressive expression paired graphic and text multi-modal data entity, and constructs a multi-modal deep learning model;
[0063] The second module inputs the original sequence into the multi-modal deep learning model to obtain the prediction results of the aggressive expression and calculates the cumulative modal difference score;
[0064] The third module constructs a decision-making system for the reinforcement learning policy network, and defines the necessary components of the agent, environment, state, action, and reward function;
[0065] The fourth module inputs the cumulative modal difference score into the policy network to obtain the sampling data volume of each modality of the next round of aggressive expression data, and updates the policy network according to the reward function;
[0066] The fifth module finally calculates the loss based on the prediction results of the aggressive expression of the multi-modal deep learning model, and updates the parameters to train the model.
[0067] The specific implementation manners of the first to fifth modules are the same as the foregoing method steps, and will not be elaborated here.
[0068] Based on the multi-modal deep learning model, the present invention is committed to alleviating the modal imbalance problem in the training process. Considering from the data input level of the model, the cumulative modal difference score is proposed to monitor the learning situation of each modality, and the reinforcement learning network is used to perceive the learning process of the modality and map it to the sampling volume of the next round of data in real time, dynamically modifying the input data volume of each modality. In the multi-modal classification task, it has excellent performance and can be further applied to tasks such as aggressive expression detection and fake news detection.
[0069] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A data-aware unimodal sampling method based on reinforcement learning, characterized in that: The steps include: Step 1: Processing the original samples into original sample sequences, where the original samples are paired image-text multimodal data entities of offensive expressions, and building a multimodal deep learning model; Step 2, input the original sequence into the multimodal deep learning model to obtain the aggressive expression prediction results and calculate the cumulative modality difference score; Step 3: Build a decision-making system for the reinforcement learning strategy network, define the agent, environment, state, action, and reward function; Step 4: Input the accumulated modality difference scores into the policy network to obtain the amount of sampled data for each modality of the next round of aggressive expression data, and update the policy network according to the reward function; Step 5: Calculate the loss based on the aggressive expression prediction results of the multimodal deep learning model and update the parameters to train the model.
2. The data-aware unimodal sampling method based on reinforcement learning according to claim 1, characterized in that: In step 1, obtaining the original sample sequence and constructing a multimodal deep learning model includes the following steps: Initialize the multimodal data entities of offensive expression pairs of images and texts. The specific form is: Define the data set D to be of size n and contain m modes; X (j) represents the entire data of the jth mode, Represents the i-th data of the j-th mode; define y i ∈{0,1} c To represent the label of each data, where c represents the total number of categories; Building a multimodal deep learning model includes building a backbone network and a classification network; using the transformer-based BERT model and the Vision Transformer model as the backbone network for processing paired image and text data; the classification network is composed of the simplest linear layer and the nonlinear activation function ReLU; the backbone network and classification network of the jth modality are respectively denoted as and g (j) .
3. The data-aware unimodal sampling method based on reinforcement learning according to claim 2, characterized in that: In step 2, the original sequence is input into the multimodal deep learning model to obtain the aggressive expression prediction result and calculate the cumulative modality difference score, which specifically includes: ① Assume that in the tth iteration, a small batch of data B is obtained by random sampling (j) , where n b Indicates the batch size; the specific form is: ② Input the small batch data into the multimodal deep learning model to obtain the prediction results; represents the characteristics of the jth mode of the ith sample, It represents its predicted probability, spftmax represents the normalized function; its characteristics and predicted probability are in the following specific forms: ③ According to the prediction results of the multimodal deep learning model, it is expected to define the modal difference score to effectively feedback the learning status of each modality and monitor the learning process of each modality; therefore, the cumulative modal difference score should be positively correlated with the prediction accuracy of the model, so the modal difference score is defined, where Represents the modal difference score: Due to the instability of random batch sampling, the modal difference score is generalized to the cumulative modal difference score 4. The data-aware unimodal sampling method based on reinforcement learning according to claim 3, characterized in that: In step 3, a decision system of a reinforcement learning strategy network is constructed to define the agent, environment, state, action, and reward function, specifically including: ① Construct an intelligent agent of the reinforcement learning strategy network, which is composed of a linear layer and a nonlinear activation function ReLU, ψ ω Represents a reinforcement learning policy network, which is in the following form, where FC represents a linear layer, ReLU represents a ReLU activation function, and Dim represents the dimension size of the input feature: ψ ω ={FC(Dim×256)→ReLU→FC(Dim×64)→ReLU→FC(64×Dim)} ② Construct an environment for the reinforcement learning strategy network and define the entire multimodal deep learning model training process as a reinforcement learning environment. The agent needs to perceive the learning status of each modality in the current modality and take corresponding actions to change the environment. ③ Construct the state of the reinforcement learning strategy network; define the state space S based on the cumulative modal difference score defined above; in the tth iteration, the state of the current agent is defined as an m-dimensional vector, which is as follows: represents the cumulative modal difference score of the i-th mode: ④Build the action of reinforcement learning strategy network; define the action space Indicates that all vectors in the action space A are positive integer vectors of dimension m; in the tth iteration, the action vector given by the current agent is in the following form: in Represents the next sampling amount given by the reinforcement learning model for the jth mode; ⑤ Construct the reward function of the reinforcement learning strategy network; given a state vector, the agent in the strategy network needs to generate an action vector; the specific form is as follows: Where N B Indicates the preset maximum amount of data in a single batch. represents the output of the reinforcement learning policy network, and round represents the floor function; the reward function is defined as follows: in represents the indicator function, r represents the reward function, and log represents the logarithm.
5. The data-aware unimodal sampling method based on reinforcement learning according to claim 4, characterized in that: In step 4, the accumulated modality difference scores are input into the policy network to obtain the sampled data volume of each modality of the next round of aggressive expression data, and the policy network is updated according to the reward function, including the following steps: ① Input the accumulated modal difference score into the policy network to obtain the amount of sampling data for the next round, which is in the same form as the action vector generated in step 3; ② Calculate the gradient of the agent parameter ω according to the reward function, and use it to continuously update the policy network. The specific form is as follows, where E represents the expectation:
6. The data-aware unimodal sampling method based on reinforcement learning according to claim 5, characterized in that: In step 5, the loss is calculated based on the aggressive expression prediction result of the multimodal deep learning model, and the parameters are updated to train the model. The specific form is as follows: Taking cross entropy loss as the loss function, the empirical risk minimization of the multimodal deep learning model is expressed as follows, where L represents the loss function of the entire data set:
7. An aggressive expression detection system based on reinforcement learning, characterized in that: For implementing the method described in any one of claims 1 to 6, the system comprises: In the first module, the original samples are processed into original sample sequences, paired image-text multimodal data entities of the original sample's aggressive expressions, and a multimodal deep learning model is constructed; In the second module, the original sequence is input into the multimodal deep learning model to obtain the prediction results of aggressive expression and calculate the cumulative modality difference score; The third module builds the decision-making system of the reinforcement learning strategy network and defines the agent, environment, state, action, and reward function; The fourth module inputs the accumulated modality difference scores into the policy network to obtain the sampling data volume of each modality of the next round of aggressive expression data, and updates the policy network according to the reward function; In the fifth module, the loss is finally calculated based on the aggressive expression prediction results of the multimodal deep learning model, and the parameters are updated to train the model.
8. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the method described in any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method described in any one of claims 1 to 6 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 6 are implemented.
Citation Information
Cited By
Evidence deep learning method for double-layer dynamic uncertainty calibration based on meta-strategy
CN121031725A
AI data lake-based model optimization method, device, medium, and apparatus
CN122549629A