Causal relationship-based multi-modal impression recognition method and device, equipment and medium
By combining causal models, BLSTM networks, and attention networks, the problem of low impression recognition accuracy in existing technologies is solved. This enables the capture and cross-domain fusion of causal relationships between speakers and listeners, thereby improving the accuracy of impression recognition.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- PING AN TECH (SHENZHEN) CO LTD
- Filing Date
- 2023-06-15
- Publication Date
- 2026-05-01
AI Technical Summary
The accuracy of impression recognition in existing technologies is low, mainly because the causal relationship between the speaker and the listener is ignored, resulting in inaccurate recognition results.
By acquiring the speaker's first multimodal data and the listener's second multimodal data, causal identification is performed using a pre-defined causal model. By combining BLSTM networks, attention networks, and fully connected networks, cross-domain fusion and impression recognition are achieved.
It improves the accuracy of impression recognition, better captures hidden features and causal relationships between speakers and listeners, and improves the performance of single-sided impression recognition.
Smart Images

Figure CN116665306B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biometrics, and in particular to a multimodal impression recognition method, apparatus, device, and medium based on causal relationships. Background Technology
[0002] With the continuous development of technology, biometric technology is widely used in various fields. Among them, impression recognition, as a branch of emotion recognition, requires the identification of the behavior of the speaker or listener to determine the impression. Moreover, impression recognition is very important in some scenarios (such as bank loan applications, job interviews, speeches, insurance applications, or product recommendations) where a deductive approach is needed to achieve a good impression.
[0003] In existing technologies, impression recognition methods primarily rely on observing the speaker's audio data, facial expressions, and body language during interactions to determine the impression the speaker makes on the listener. In other words, current impression recognition depends on how the listener interprets the speaker's expressions and behavior. For example, in loan application scenarios, the impression of the borrower is crucial. Similarly, in credit card issuance scenarios, impression recognition can more accurately predict risk. However, most current research identifies impressions using only the behavior of a single speaker or listener, neglecting the fact that impression recognition is formed jointly by the behaviors of both, leading to lower accuracy rates. Summary of the Invention
[0004] This invention provides a multimodal impression recognition method, apparatus, device, and medium based on causal relationships to solve the problem of low impression recognition accuracy in the prior art.
[0005] A multimodal impression recognition method based on causal relationships includes:
[0006] Acquire the speaker's first multimodal data and the listener's second multimodal data;
[0007] A first feature is obtained by performing causal identification on the first multimodal data and the second multimodal data using a preset causal model;
[0008] The second feature is obtained by extracting features from the second multimodal data using a preset first BLSTM network;
[0009] The first fused feature is obtained by cross-domain fusion of the first feature and the second feature through a preset attention network;
[0010] Impression recognition results are obtained by performing impression recognition on the first fused features through a preset fully connected network.
[0011] A causal-based multimodal impression recognition device includes:
[0012] The data acquisition module is used to acquire the speaker's first multimodal data and the listener's second multimodal data;
[0013] The causal identification module is used to perform causal identification on the first multimodal data and the second multimodal data through a preset causal model to obtain a first feature;
[0014] The feature extraction module is used to extract features from the second multimodal data through a preset first BLSTM network to obtain the second features;
[0015] A cross-domain fusion module is used to perform cross-domain fusion of the first feature and the second feature through a preset attention network to obtain a first fused feature;
[0016] The recognition and prediction module is used to perform impression recognition on the first fused features through a preset fully connected network to obtain the first impression recognition result.
[0017] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the above-described causal-based multimodal impression recognition method.
[0018] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described causal-based multimodal impression recognition method.
[0019] This invention provides a multimodal impression recognition method, apparatus, device, and medium based on causal relationships. The method acquires first multimodal data of the speaker and second multimodal data of the listener, and performs causal identification on the first and second multimodal data using a preset causal model. This achieves the acquisition of a first feature, thereby uncovering hidden correlation features between the two and capturing the causal relationship. A preset attention network is used to perform cross-domain fusion of the first and second features, achieving cross-domain fusion of relevant information from the speaker and listener, which can better extract hidden correlation information. A preset fully connected network is used to perform impression recognition on the first fused features, achieving the acquisition of the first impression recognition result. Based on the multimodal data of both, the listener's impression of the speaker is identified, improving the accuracy of impression recognition and enhancing the performance of impression recognition from a single perspective. Attached Figure Description
[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a schematic diagram illustrating the application environment of a multimodal impression recognition method based on causal relationships in one embodiment of the present invention;
[0022] Figure 2 This is a flowchart of a multimodal impression recognition method based on causal relationships in one embodiment of the present invention;
[0023] Figure 3 This is a flowchart of step S20 of a multimodal impression recognition method based on causal relationships in an embodiment of the present invention;
[0024] Figure 4 This is a flowchart of step S40 of a multimodal impression recognition method based on causal relationships in an embodiment of the present invention;
[0025] Figure 5 This is a schematic diagram of a multimodal impression recognition device based on causal relationships in one embodiment of the present invention;
[0026] Figure 6 This is a schematic diagram of a computer device according to an embodiment of the present invention. Detailed Implementation
[0027] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0028] The causal-based multimodal impression recognition method provided in this invention can be applied to, for example... Figure 1 The application environment is shown. Specifically, this causal-based multimodal impression recognition method is applied in a causal-based multimodal impression recognition device, which includes, as shown in the example, [details of the example]. Figure 1The client and server shown communicate over a network to address the low accuracy of image recognition in existing technologies. The server can be a standalone server or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The client, also known as the user terminal, refers to the program that provides classification services to customers, corresponding to the server. The client can be installed on, but is not limited to, various computers, laptops, smartphones, tablets, and portable wearable devices.
[0029] In one embodiment, such as Figure 2 As shown, a multimodal impression recognition method based on causal relationships is provided, which can be applied to... Figure 1 Taking the server in the example, the following steps are included:
[0030] S10: Obtain the speaker's first multimodal data and the listener's second multimodal data.
[0031] Understandably, the first type of multimodal data includes the speaker's audio data, facial video, eye gaze data, and physiological signals. For example, during product recommendation, listeners use the content of an employee's explanation, facial expressions, eye gaze data, and heart rate (or EEG, ECG, and blood volume pulse) to confirm the employee's engagement; during insurance purchase, the user's speech, facial expressions, eye gaze data, and heart rate confirm the user's purchase intention; and in Ping An Bank's business transaction scenarios, the salesperson's tone of voice, facial expressions, or eye gaze data are used; in identity authentication scenarios, opening accounts at Ping An Bank and Ping An Securities, and transferring large sums of money, require security verification of the operator, which can be confirmed through the user's tone of voice, facial expressions, or eye gaze data to confirm their willingness. The second type of multimodal data includes the listener's audio data, facial video, eye gaze data, and physiological signals. For example, during product recommendations, the audience's discussion audio, facial expressions, eye gaze data, and physiological signals related to an employee's presentation determine their level of acceptance of the content. In insurance purchases, the salesperson's speech, facial expressions, eye gaze data, and heart rate confirm whether the price can be adjusted. Both primary and secondary multimodal data can be collected from different databases or sent from the client to the server. In some scenarios, there may be one speaker and multiple listeners.
[0032] S20: Perform causal identification on the first multimodal data and the second multimodal data using a preset causal model to obtain a first feature.
[0033] Understandably, the presupposed causal model is based on projection-weighted canonical correlation analysis (PWCCA) and causal gating. The first feature is the speaker's characteristics.
[0034] Specifically, the first and second multimodal data are input together into a preset causal model. The preset causal model then performs causal identification on the first and second multimodal data. First, the first and second multimodal data are segmented into multiple data segments, preferably 450 segments. Then, the corresponding segments of the first and second multimodal data are combined into segmented data pairs, i.e., one segment of first multimodal data and one segment of second multimodal data form a segmented data pair. Causal relationships are calculated across all segmented data pairs using projection-weighted canonical correlation analysis (PWCCA) to identify causal relationships and correlation coefficients (association values) for each pair. Directions with a higher proportion of input are given higher weights. Next, causal gating is used to assign the calculated PWCCA causal correlation values to each speaker's segmented data as weights. Finally, by weighting all the speaker's segmented data and their corresponding weights, the speaker's first feature is obtained. For example, in the financial sector, during loan processing, salespeople observe changes in the applicant's speech, such as facial expressions like joy or a smile. Or, in insurance customer service scenarios, agents need to determine a customer's interest based on their tone of voice to decide whether to continue the explanation.
[0035] S30, the second feature is obtained by extracting features from the second multimodal data through a preset first BLSTM network.
[0036] Understandably, the second feature is the listener feature extracted from the second multimodal data.
[0037] Specifically, after obtaining the second multimodal data, the various data features in the second multimodal data are concatenated, and the concatenated data features are input into a preset first BLSTM network. The preset first BLSTM network extracts features from the concatenated data features; that is, it extracts features from the concatenated data features using two LSTM networks in opposite directions within the preset first BLSTM network, thus obtaining the second feature corresponding to the listener's second modal data. Understandably, when there are multiple listeners, the preset first BLSTM network extracts features from the data features of each listener's second multimodal data separately, obtaining the second feature corresponding to each second multimodal data point. For example, in a technical personnel interview process, multiple interviewers may interview the same candidate; therefore, feature extraction is needed for each interviewer's multimodal data. Similarly, in some risk detection processes, multiple staff members may detect a user's voice content, facial expressions, and eye contact data; therefore, feature extraction is needed for each staff member's multimodal data.
[0038] S40: The first feature and the second feature are fused across domains using a preset attention network to obtain the first fused feature.
[0039] Understandably, the first fused feature is obtained by fusing the features of the speaker and the listener through a preset attention network. The preset attention network includes a first preset number of intra-attention networks and a second preset number of inter-attention networks. The first preset number and the second preset number are the same.
[0040] Specifically, the first and second features are input into a pre-defined attention network. This network performs cross-domain fusion of the first and second features, treating the speaker's and listener's features as different domains. The network then extracts the correlation between features from different domains and extracts individual features from each domain, resulting in a predetermined number of equal-length encoded features. These equal-length encoded features are then fused together, either by concatenating them or by fusing them with different weights, to obtain the first fused feature. When multiple listeners are present, the pre-defined attention network performs cross-domain fusion of the first feature and all second features separately, yielding the first fused feature corresponding to each second feature. For example, in some risk detection processes, multiple staff members may be detecting a user's voice content, facial expressions, and eye contact data. Therefore, it is necessary to fuse the user's features with the features of each staff member and identify the user's impression based on all the fused features.
[0041] S50: The first fused feature is subjected to impression recognition through a preset fully connected network to obtain the first impression recognition result.
[0042] Understandably, a fully connected network is presumably composed of fully connected layers. The first impression recognition result is obtained by performing impression recognition based on first and second multimodal data. Examples include competence, enthusiasm, good temper, trustworthiness, tolerance, friendliness, and sincerity.
[0043] Furthermore, all the first fused features are input into a pre-defined fully connected network. This network performs impression recognition on the first fused features, specifically through the fused features of the speaker and listener. This involves a non-linear transformation of the first fused features using fully connected layers and calculations performed on the features using hidden units in hidden layers. The calculation results are then classified using a function in the output layer and activated by a ReLU function to obtain the first impression recognition result. This impression recognition result includes at least one impression and its corresponding probability. For example, in an insurance customer service scenario, the customer service representative determines the customer's impression based on the user's tone of voice and facial expressions, such as 95% good temper and 92% tolerant and friendly, and can then continue explaining to the user.
[0044] This invention discloses a causal-based multimodal impression recognition method. This method acquires first multimodal data of the speaker and second multimodal data of the listener, and performs causal identification on the first and second multimodal data using a preset causal model. This achieves the acquisition of a first feature, captures the correlation between the speaker's and listener's behaviors, and further mines hidden related features between them, thus capturing the causal relationship between them. A preset attention network is used to perform cross-domain fusion of the first and second features, achieving cross-domain fusion of relevant information between the speaker and listener, which can better extract hidden related information between them. A preset fully connected network is used to perform impression recognition on the first fused features, acquiring the first impression recognition result. Then, based on the multimodal data of both parties, the listener's impression of the speaker is identified, improving the accuracy of impression recognition and enhancing the performance of impression recognition from a single aspect.
[0045] In one embodiment, step S20, namely, performing causal identification on the first multimodal data and the second multimodal data using a preset causal model to obtain a first feature, includes:
[0046] S201, the first multimodal data and the second multimodal data are segmented through the segmentation layer in the preset causal model to obtain the first segmented data and the second segmented data.
[0047] Understandably, segmented data is obtained by segmenting multimodal data.
[0048] Specifically, after obtaining the first and second multimodal data, the data features in the first and second multimodal data are fused to obtain the first and second data features. The fused first and second data features are then input into a preset causal model. The preset causal model performs causal identification on the first and second data features, that is, it segments the first and second data features using a segmentation layer. First, the first and second data features are aligned, and then the aligned first and second data features are segmented together to obtain at least one first segment data corresponding to the first data feature, and at least one second segment data corresponding to the second data feature. For example, in an insurance claims scenario, a 5-minute multimodal data segment between staff and customers is divided into 450 segments.
[0049] S202, the first segmented data and the second segmented data are analyzed for correlation through the PWCCA layer in the preset causal model to obtain the correlation value.
[0050] Understandably, the correlation value is the degree of correlation between the segmented data pairs (the first segment and the second segment).
[0051] Furthermore, the PWCCA layer in the pre-defined causal model is used to perform correlation analysis on the first and second segment data, that is, to associate a segment of first segment data with a segment of second segment data to obtain segment data pairs. The PWCCA layer in the pre-defined causal model is then used to identify the causal relationship between the first and second segment data in each segment data pair through projection-weighted canonical correlation analysis. In this case, PWCCA assigns higher weights to directions with a higher proportion of input, thereby finding data in all segment data pairs that are highly correlated between the speaker's first segment data and the listener's second segment data. The correlation between highly correlated segment data pairs is then calculated, that is, the correlation coefficient is calculated, and the segment data pairs are weighted according to the correlation coefficient to obtain the correlation value between the first and second segment data.
[0052] Furthermore, in one embodiment, X = [a1, a2, ..., a...] p ] T Y = [b1, b2, ..., b q ] T a i and b i Let X and Y be the eigenvectors of the i-th features of X and Y, respectively. Calculate the eigenvariance λ of each feature.i =var(a i ), μ i =var(b i ). Utilizing λ i and μ i We perform weighted processing on the feature datasets X and Y to obtain the processed feature dataset. and Depend on and Calculate the covariance matrix S XX and S YY and the cross-covariance matrix S XY Solve the objective function. This yields the projection vector set, from which d feature projection vectors are selected. Then, α is calculated. i and β i The corresponding correlation coefficient ρ i Based on the correlation coefficient ρ i right and By weighting the vectors and combining them into the weighted projection vector set, the associated values can be obtained.
[0053] S203, the associated values are allocated through the causal relationship gating in the preset causal model to obtain the weight values corresponding to each of the first segment data.
[0054] S204, Based on all the first segment data and all the weight values, determine the speaker's target features;
[0055] S205, the target features are extracted by a preset second BLSTM network to obtain the first feature.
[0056] Understandably, the weight values are the association values assigned to the first segment of data. The target features are the speaker's features.
[0057] Specifically, all correlation values, first segment data, and second segment data are input into a causal relationship gating system. The correlation values are assigned using the causal relationship gating system within the preset causal model. That is, the causal relationship gating system assigns the calculated PWCCA causal relationship correlation values to the first segment data of each speaker as weights, thus obtaining the weight values corresponding to each first segment data. Based on all first segment data and all corresponding weight values, the speaker's features are extracted to obtain the speaker's target features. Furthermore, the target features are input into a preset second BLSTM network. The preset second BLSTM network extracts features from the target features; that is, the forward and backward LSTM networks within the preset second BLSTM network extract the hidden state features from the target features, thus obtaining the speaker's first features. The preset second BLSTM network has the same structure as the preset first BLSTM network, and the specific process is the same as the processing of the preset first BLSTM network. Please refer to steps S301 to S303 below, which will not be repeated here.
[0058] This invention, through dividing the first and second multimodal data into segments and calculating the correlation coefficient between the segmented data pairs, achieves the calculation of the association value between the segmented data pairs. By assigning the association value through causal gating, the weight value of the first segmented data is obtained, thereby achieving the acquisition of the speaker's target features. Feature extraction of the first feature is achieved by using a pre-set second BLSTM network to extract features from the first feature, realizing the extraction of the first feature in the hidden state.
[0059] In one embodiment, step S30, namely, extracting features from the second multimodal data using a preset first BLSTM network to obtain second features, includes:
[0060] S301, the positive features are obtained by extracting features from the second multimodal data through the forward LSTM network.
[0061] S302, the reverse LSTM network is used to extract features from the second multimodal data to obtain reverse features.
[0062] S303, the positive feature and the negative feature are fused to obtain the second feature.
[0063] The first BLSTM network is pre-defined to include a forward LSTM network and a reverse LSTM network.
[0064] Understandably, positive features are obtained by the forward-order long short-term memory network extracting features from the second multimodal data. Negative features are obtained by the reverse-order long short-term memory network extracting features from the second multimodal data.
[0065] Specifically, after obtaining the second multimodal data, it is input into a preset first BLSTM network. The preset first BLSTM network extracts features from the second multimodal data; that is, it extracts features from the second multimodal data using both the forward and backward LSTM networks within the preset first BLSTM network. Specifically, the forget gate in either the forward or backward LSTM network selects features from the hidden features of the previous time step and the features of the currently input second multimodal data, outputting a value between 0 and 1. Based on this value, the features of the second multimodal data are selectively forgotten. The sigmoid layer in the input gate determines the information that needs updating from the hidden features of the previous time step and the features of the currently input second multimodal data, and a new candidate feature is created through a tanh layer. The two features are multiplied to obtain a feature result. The selected and retained features of the second multimodal data at the current time step are multiplied with the output features from the previous time step, and the feature result is added to obtain the updated features. Finally, the sigmoid layer in the output gate determines the information to be output, and the updated features are normalized by the tanh layer to obtain normalized features. The normalized features are then multiplied by the output information to obtain either positive or negative features. Finally, the positive and negative features are concatenated to obtain the listener's second feature. The negative feature is obtained based on a reverse LSTM network.
[0066] Furthermore, in one embodiment, the forward LSTML is sequentially input with "I", "purchase", and "insurance" to obtain three vectors {hL0, hL1, hL2}. The reverse LSTMR is sequentially input with "insurance", "purchase", and "I" to obtain three vectors {hR0, hR1, hR2}. Finally, the forward and reverse latent vectors are concatenated to obtain {[h... L0 h R2 ], [h L1 h R1 ], [h L2 h R2 ]}, that is, {h0, h1, h2}. Understanding this statement from left to right, we can obtain h at each time step. Li Hidden layer output, and h can be obtained at each time step. Ri The hidden layer output is concatenated by BLSTM, which concatenates the forward and backward hidden layer outputs at each time step. Li h Ri ] represents the hidden layer output of the features at the current time.
[0067] This invention embodiment extracts features from the second multimodal data by using a preset first BLSTM network, thereby extracting hidden state information from the second multimodal data, extracting positive and negative features, and ultimately extracting the second feature, thus improving the accuracy of subsequent impression recognition.
[0068] In one embodiment, step S40, namely, performing cross-domain fusion of the first feature and the second feature through a preset attention network to obtain a first fused feature, includes:
[0069] S401, the first feature is input into the first attention network to obtain the first encoded feature;
[0070] S402, input the first feature and the second feature into the first attention network to obtain the second encoded feature;
[0071] S403, input the first feature and the second feature into the second attention network to obtain the third encoded feature;
[0072] S404, The second feature is input into the second attention network to obtain the fourth encoded feature;
[0073] S405, the first coding feature, the second coding feature, the third coding feature and the fourth coding feature are fused to obtain the first fused feature.
[0074] The pre-defined attention network includes an intra-attention network and an inter-attention network.
[0075] Specifically, the first and second features are input into a preset attention network. The preset attention network encodes the first and second features, specifically by performing attention processing on the first feature within the first attention network. This involves calculating the Q, K, and V vectors of the first feature using an attention mechanism. Specifically, the correlation score between the Q and K vectors of the first feature is calculated using the dot product method, i.e., calculating the dot product between each feature in Q and each feature in K, and then normalizing the correlation score between the Q and K vectors. Then, the softmax function is used to convert the scores between the vectors into a probability distribution between [0, 1]. Based on this probability distribution, the corresponding Values vector is multiplied to obtain the first encoded feature. Similarly, the process of processing the second feature through the second attention network is similar to the above process; please refer to the process of processing the first feature through the first attention network, which will not be repeated here.
[0076] Furthermore, the first and second features are input into a first inter-attention network. This network encodes the first and second features by dividing the first feature into K and V vectors, and the second feature into a Q vector. Specifically, the correlation score between the Q vector of the second feature and the K vector of the first feature is calculated using the dot product method. This involves calculating the dot product between each feature in Q and each feature in K, and then normalizing the correlation score between the Q and K vectors. Then, the softmax function is used to convert the scores between the vectors into a probability distribution between [0, 1]. Based on this probability distribution, the corresponding Values vector of the first feature is multiplied to obtain the second encoded feature. For example, the speaker's first feature is divided into two sub-features, designated as K and V vectors. The listener's second feature is designated as a Q vector. The Q, K, and V vectors are then calculated using a pre-defined inter-attention network to obtain the second encoded feature.
[0077] Similarly, the first and second features are input into the second attentional inner network. The second attentional inner network encodes the first and second features by dividing the second feature into K and V vectors, and the first feature into Q vector. This involves using the dot product method to calculate the correlation score between the Q vector of the first feature and the K vector of the second feature; that is, calculating the dot product between each feature in Q and each feature in K, and normalizing the correlation score between the Q and K vectors. Then, the softmax function is used to convert the scores between the vectors into a probability distribution between [0, 1]. Based on this probability distribution, the corresponding second feature's Values vector is multiplied to obtain the third encoded feature. For example, the speaker's second feature is divided into two sub-features, designated as K and V vectors. The listener's first feature is designated as the Q vector. The Q, K, and V vectors are then calculated using a pre-defined attentional inner network to obtain the third encoded feature.
[0078] This invention employs an intra-attention network to encode a first feature and an inter-attention network to encode a second feature, thereby acquiring both the first and fourth encoded features. Furthermore, by performing cross-domain fusion of the first and second features through the intra-attention network and the inter-attention network, the second and third encoded features are acquired, leading to better extraction of hidden information and improved accuracy in impression recognition.
[0079] In one embodiment, after step S30, that is, after extracting features from the second multimodal data using a preset first BLSTM network to obtain the second features, the process includes:
[0080] S601, Obtain the listener's identity information and encode the identity information using a preset monitoring model to obtain embedded features.
[0081] Understandably, the identity information refers to the listener's identity, such as an engineer, a regular employee, or a teacher. The embedded features are obtained by encoding the identity information.
[0082] Specifically, the identity information of all listeners is retrieved from the database and input into a preset monitoring model. The model encodes this identity information, first converting it into a vector using a one-hot encoder, thus obtaining the vector embedding. Then, the vector embedding is encoded using a linear layer within the model, calculating the embedding through hidden units to obtain the vector features. Finally, the vector features are upsized, converting their dimension to match the second feature dimension, resulting in the embedded features. For example, in an insurance interview, listeners are typically department leaders or key personnel. Similarly, in a bank loan application process, listeners are typically individuals from various fields.
[0083] S602, the second feature and the embedded feature are concatenated to obtain the concatenated feature;
[0084] S603, the first feature and the spliced feature are fused across domains through the preset attention network to obtain the second fused feature;
[0085] S604, the second fused feature is subjected to impression recognition through the preset fully connected network to obtain the second impression recognition result.
[0086] Understandably, the concatenated feature is obtained by concatenating the second feature and the embedded feature. The second fused feature is obtained by cross-domain fusion of the first feature and the concatenated feature.
[0087] Specifically, after obtaining the embedded features, the second feature and the embedded features are concatenated, that is, the second feature of the listener's multimodal data and the embedded features of the listener's identity information are concatenated together to obtain the concatenated feature. The first feature and the concatenated feature are then input into a preset attention network. The preset attention network performs cross-domain fusion on the first feature and the concatenated feature; that is, the speaker's and listener's information are treated as different domains, and the information between different domains is fused to obtain the second fused feature. Finally, the second fused feature is input into a preset fully connected network. The preset fully connected network performs impression recognition on the second fused feature; that is, the fully connected layer predicts the impression in the second fused feature to obtain the second impression recognition result. The specific process is the same as the specific process of steps S40 to S50 above, and will not be repeated here.
[0088] This invention, through the introduction of identity information, addresses annotation biases introduced when evaluating speaker recordings from different listeners. By concatenating the embedded features of the listener's identity information with the listener's secondary features, the concatenated features are obtained, thereby enabling the acquisition of the secondary impression recognition result and improving the accuracy of impression recognition.
[0089] In one embodiment, before step S20, that is, before performing causal identification on the first multimodal data and the second multimodal data using a preset causal model, the following steps are included:
[0090] S206, Obtain a sample training dataset, wherein the sample training dataset includes at least one sample training data and sample labels corresponding to the sample training data.
[0091] Understandingly, sample training data can be multimodal data such as audio, facial expressions, body language, and physiological signals of speakers and listeners. Examples include video recordings of historical claims scenarios in the insurance industry, or video recordings of past employee interviews in the banking industry. Each sample training data point is associated with a sample label, which characterizes the speaker's true features within the sample training data. These true features can be obtained through manual or other methods of analyzing the sample training data. Sample training data and sample labels can be collected from different databases or can be pre-prepared data sent from the client to a database. A sample training dataset is then constructed based on all the acquired sample training data and the corresponding sample labels.
[0092] S207, perform causal identification on the sample training data using a preset training model to obtain predicted labels.
[0093] Understandably, the predicted labels are used to characterize the speaker's features obtained by causal identification of the sample training data.
[0094] Specifically, all sample training data and sample labels are input into a pre-set training model. The model then performs causal identification on the training data, first dividing it into segments using a segmentation layer. Next, causal relationships are calculated using each segment feature. The PWCCA layer in the pre-set training model performs correlation analysis between these segments, specifically calculating the correlation between speakers and listeners in the training data and assigning higher weights to directions with higher input proportions, thus obtaining correlation values. Causal relationship gating assigns the calculated PWCCA correlation values to each speaker's segment feature as weights. By combining all speaker segment features and all weights, a causal-weighted speaker feature is formed, yielding the predicted label. The specific process is the same as the causal identification process of the pre-set causal model described above and will not be repeated here.
[0095] S208, determine the prediction loss value of the preset training model based on the prediction label and the sample label corresponding to the same sample training data.
[0096] Understandably, the prediction loss is generated during the prediction process using the training data.
[0097] Specifically, after obtaining the predicted labels, all predicted labels corresponding to the sample training data are arranged according to the order of the sample training data in the sample training dataset. Then, the predicted labels associated with the sample training data are compared with the sample labels of the sample training data with the same sequence. That is, according to the sample training data, the sample label corresponding to the first sample training data is compared with the predicted label corresponding to the first sample training data. The loss value between the sample label and the predicted label is determined by the loss function. This process continues until all sample labels and predicted labels have been compared, and then the predicted loss value of the preset training model can be obtained.
[0098] S209, when the predicted loss value reaches the preset convergence condition, the preset training model after convergence is recorded as the preset causal model.
[0099] Understandably, the convergence condition can be either the predicted loss value being less than a set threshold, or the predicted loss value being very small after 500 calculations and no longer decreasing, at which point training can stop.
[0100] Specifically, after obtaining the predicted loss value, if the predicted loss value does not reach the preset convergence condition, the initial parameters of the preset training model are adjusted based on the predicted loss value. All sample training data are then re-input into the preset training model with adjusted initial parameters, and iterative training is performed to obtain the predicted loss value corresponding to the preset training model with adjusted initial parameters. Then, if the predicted loss value does not reach the preset convergence condition, the initial parameters of the preset training model are adjusted again based on the predicted loss value, so that the predicted loss value of the preset training model with adjusted initial parameters reaches the preset convergence condition. In this way, the accuracy of the preset training model increases, and the obtained quality inspection results continuously approach the correct results, until the predicted loss value of the preset training model reaches the preset convergence condition. At this point, the converged preset training model is determined as the preset causal model.
[0101] This invention iteratively trains a pre-defined training model using a large amount of sample training data and calculates the overall loss value of the pre-defined training model by comparing the loss function, thereby determining the predicted loss value of the pre-defined training model. The initial parameters of the pre-defined training model are adjusted based on the predicted loss value until the model converges, thus training the pre-defined causal model and ensuring its high accuracy.
[0102] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0103] In one embodiment, a causal-based multimodal impression recognition device is provided, which corresponds one-to-one with the causal-based multimodal impression recognition method described in the above embodiments. For example... Figure 5 As shown, the causal-based multimodal impression recognition device includes a data acquisition module 11, a causal recognition module 12, a feature extraction module 13, a cross-domain fusion module 14, and an impression recognition module 15. Detailed descriptions of each functional module are as follows:
[0104] Data acquisition module 11 is used to acquire the speaker's first multimodal data and the listener's second multimodal data;
[0105] The causal identification module 12 is used to perform causal identification on the first multimodal data and the second multimodal data through a preset causal model to obtain a first feature;
[0106] Feature extraction module 13 is used to extract features from the second multimodal data through a preset first BLSTM network to obtain second features;
[0107] The cross-domain fusion module 14 is used to perform cross-domain fusion of the first feature and the second feature through a preset attention network to obtain a first fused feature;
[0108] The impression recognition module 15 is used to perform impression recognition on the first fused feature through a preset fully connected network to obtain the first impression recognition result.
[0109] In one embodiment, the causality identification module 12 includes:
[0110] The segmentation unit is used to segment the first multimodal data and the second multimodal data through the segmentation layer in the preset causal model to obtain the first segmented data and the second segmented data.
[0111] The analysis unit is used to perform correlation analysis on the first segment data and the second segment data through the PWCCA layer in the preset causal model to obtain correlation values;
[0112] The allocation unit is used to allocate the associated values through the causal relationship gating in the preset causal model to obtain the weight values corresponding to each of the first segment data.
[0113] The target feature unit is used to determine the speaker's target features based on all the first segment data and all the weight values;
[0114] The feature extraction unit is used to extract features from the target features through a preset second BLSTM network to obtain the first feature.
[0115] In one embodiment, the feature extraction module 13, wherein the preset first BLSTM network includes a forward LSTM network and a reverse LSTM network, comprising:
[0116] A positive feature unit is used to extract features from the second multimodal data through the positive LSTM network to obtain positive features;
[0117] The inverse feature unit is used to extract features from the second multimodal data through the inverse LSTM network to obtain inverse features;
[0118] The second feature unit is used to fuse the positive feature and the negative feature to obtain the second feature.
[0119] In one embodiment, the cross-domain fusion module 14 includes a preset attention network comprising an intra-attention network and an inter-attention network, comprising:
[0120] The first unit is used to input the first feature into the first attention network to obtain the first encoded feature;
[0121] The second unit is used to input the first feature and the second feature into the first attention network to obtain the second encoded feature;
[0122] The third unit is used to input the first feature and the second feature into the second attention network to obtain the third encoded feature;
[0123] The fourth unit is used to input the second feature into the second attention network to obtain the fourth encoded feature;
[0124] The feature fusion unit is used to fuse the first coding feature, the second coding feature, the third coding feature and the fourth coding feature to obtain the first fused feature.
[0125] In one embodiment, the causal identification module 12 further includes:
[0126] A sample acquisition unit is used to acquire a sample training dataset, wherein the sample training dataset includes at least one sample training data and sample labels corresponding to the sample training data;
[0127] The causal identification unit is used to perform causal identification on the sample training data through a preset training model to obtain a predicted label;
[0128] The loss prediction unit is used to determine the predicted loss value of the preset training model based on the predicted label and the sample label corresponding to the same sample training data;
[0129] The model convergence unit is used to record the converged preset training model as a preset causal model when the predicted loss value reaches the preset convergence condition.
[0130] In one embodiment, the causal-based multimodal impression recognition device further includes:
[0131] An embedded feature unit is used to obtain the listener's identity information and encode the identity information through a preset listening model to obtain embedded features;
[0132] A feature splicing unit is used to splice the second feature and the embedded feature to obtain a spliced feature;
[0133] The second fusion feature unit is used to perform cross-domain fusion of the first feature and the spliced feature through the preset attention network to obtain the second fusion feature;
[0134] The second impression recognition unit is used to perform impression recognition on the second fused feature through the preset fully connected network to obtain the second impression recognition result.
[0135] Specific limitations regarding the causal-based multimodal impression recognition device can be found in the limitations of the causal-based multimodal impression recognition method described above, and will not be repeated here. Each module in the aforementioned causal-based multimodal impression recognition device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.
[0136] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 6As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system, computer programs, and the database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database stores data used in the causal-based multimodal impression recognition method described in the above embodiments. The network interface communicates with external terminals via a network connection. When the computer program is executed by the processor, it implements a causal-based multimodal impression recognition method.
[0137] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described causal-based multimodal impression recognition method.
[0138] In one embodiment, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described causal-based multimodal impression recognition method.
[0139] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.
[0140] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0141] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A multimodal impression recognition method based on causal relationships, characterized in that, include: Acquire the speaker's first multimodal data and the listener's second multimodal data; A first feature is obtained by performing causal identification on the first multimodal data and the second multimodal data using a preset causal model; The second feature is obtained by extracting features from the second multimodal data using a preset first BLSTM network; The first fused feature is obtained by cross-domain fusion of the first feature and the second feature through a preset attention network; Impression recognition results are obtained by performing impression recognition on the first fused features through a preset fully connected network; The step of performing causal identification on the first multimodal data and the second multimodal data using a preset causal model to obtain a first feature includes: The first multimodal data and the second multimodal data are segmented by the segmentation layer in the preset causal model to obtain the first segmented data and the second segmented data. The first segment data and the second segment data are correlated using the PWCCA layer in the preset causal model to obtain correlation values; The associated values are assigned by the causal relationship gating in the preset causal model to obtain the weight values corresponding to each of the first segment data. Based on all the first segment data and all the weight values, determine the speaker's target features; The target features are extracted using a pre-set second BLSTM network to obtain the first feature; The preset attention network includes an intra-attention network and an inter-attention network; The step of fusing the first feature and the second feature across domains through a preset attention network to obtain the first fused feature includes: The first feature is input into the first attention network to obtain the first encoded feature; The first feature and the second feature are input into the first attention network to obtain the second encoded feature; The first feature and the second feature are input into the second attention network to obtain the third encoded feature; The second feature is input into the second attention network to obtain the fourth encoded feature; The first coding feature, the second coding feature, the third coding feature, and the fourth coding feature are fused to obtain the first fused feature.
2. The multimodal impression recognition method based on causal relationships as described in claim 1, characterized in that, The preset first BLSTM network includes a forward LSTM network and a reverse LSTM network; The step of extracting features from the second multimodal data using a preset first BLSTM network to obtain the second features includes: The positive features are obtained by extracting features from the second multimodal data using the positive LSTM network. The inverse LSTM network is used to extract features from the second multimodal data to obtain inverse features; The positive feature and the negative feature are fused to obtain the second feature.
3. The multimodal impression recognition method based on causal relationships as described in claim 1, characterized in that, After extracting features from the second multimodal data using a preset first BLSTM network to obtain the second features, the method further includes: The listener's identity information is obtained, and the identity information is encoded through a preset listening model to obtain embedded features; The second feature and the embedded feature are concatenated to obtain the concatenated feature; The first feature and the spliced feature are fused across domains using the preset attention network to obtain the second fused feature; The second fused feature is subjected to impression recognition through the preset fully connected network to obtain the second impression recognition result.
4. The multimodal impression recognition method based on causal relationships as described in claim 1, characterized in that, Before performing causal identification on the first multimodal data and the second multimodal data using a preset causal model, the following steps are included: Obtain a sample training dataset, which includes at least one sample training data and sample labels corresponding to the sample training data; Causal identification is performed on the sample training data using a preset training model to obtain predicted labels; The prediction loss value of the preset training model is determined based on the prediction label and the sample label corresponding to the same sample training data. When the predicted loss value reaches the preset convergence condition, the preset training model after convergence is recorded as the preset causal model.
5. A multimodal impression recognition device based on causal relationships, characterized in that, include: The data acquisition module is used to acquire the speaker's first multimodal data and the listener's second multimodal data; The causal identification module is used to perform causal identification on the first multimodal data and the second multimodal data through a preset causal model to obtain a first feature; The feature extraction module is used to extract features from the second multimodal data through a preset first BLSTM network to obtain the second features; A cross-domain fusion module is used to perform cross-domain fusion of the first feature and the second feature through a preset attention network to obtain a first fused feature; An impression recognition module is used to perform impression recognition on the first fused feature through a preset fully connected network to obtain a first impression recognition result.
6. The multimodal impression recognition device based on causal relationships as described in claim 5, characterized in that, The device further includes: An embedded feature unit is used to obtain the listener's identity information and encode the identity information through a preset listening model to obtain embedded features; The feature splicing unit is used to splice the second feature and the embedded feature to obtain the spliced feature; The second fusion feature unit is used to perform cross-domain fusion of the first feature and the spliced feature through the preset attention network to obtain the second fusion feature; The second impression recognition unit is used to perform impression recognition on the second fused feature through the preset fully connected network to obtain the second impression recognition result.
7. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the causal-based multimodal impression recognition method as described in any one of claims 1 to 5.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the causal-based multimodal impression recognition method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Multi-modal emotion recognition method and system based on subspace sparse feature fusion
CN111931795A
Multi-modal emotion recognition method and system based on attention mechanism and GMN
CN113095357A