Vehicle cabin control method, device and equipment, storage medium and product
By constructing a set of positive and negative samples and training a model in the vehicle cabin, and using multimodal emotional information for emotion recognition, the problem of insufficient accuracy in emotion recognition in existing technologies is solved, achieving more efficient emotion recognition and intelligent control, and improving the user experience.
Patent Information
- Application Number
- CN202511174418.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2025-11-28
AI Technical Summary
Existing multimodal sentiment analysis algorithms struggle to fully capture the diverse correlations between cross-modal representations in vehicle cockpits, resulting in insufficient accuracy and reliability in sentiment recognition. Furthermore, the limited dataset size fails to meet users' demands for intelligent and emotional vehicle cockpits.
By using a trained emotion recognition model, multimodal emotion information is utilized for emotion recognition. A set of positive and negative samples is constructed, and single-modal representation sequences are added to train the model. This avoids introducing additional network parameters, improves the accuracy and efficiency of emotion recognition, and controls the in-cabin equipment based on the emotion score.
It improves the accuracy and reliability of emotion recognition, meets the high real-time requirements in smart cockpit scenarios, enhances the intelligence and emotionality of vehicle cockpit control, and improves the user experience.
Smart Images

Figure CN121028536A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of automobiles, in particular to a vehicle cabin control method, device, equipment, storage medium and product. BACKGROUND
[0002] With the continuous development of the automobile industry, the intelligent level of the vehicle cabin gradually becomes one of the important standards for measuring the quality of the automobile. Passenger experience is one of the important dimensions for evaluating the intelligent cabin of the vehicle, and accurately perceiving the emotional state of the passenger is of great significance for providing personalized intelligent services and improving comfort.
[0003] Traditional single-modal emotion analysis methods (such as only based on voice or expression) cannot fully utilize the complementarity of multi-modal information, are easily affected by the limitations of noise or single data source, and result in insufficient recognition accuracy.
[0004] However, existing multi-modal emotion analysis algorithms mostly focus on the design of the fusion module, ignoring the optimization and design of the representation learning, making it difficult to fully capture various associated information between cross-modal representations, thereby affecting the accuracy and reliability of emotion recognition. In addition, the existing multi-modal emotion analysis dataset is limited in size, and only uses the emotion labels of the dataset for multi-modal emotion analysis tasks, which fails to fully and sufficiently mine the information in the multi-modal data, further reducing the accuracy of perceiving passenger emotions, making it difficult for the vehicle cabin system to accurately regulate based on passenger emotions, and further leading to the inability to meet the user's demand for the intelligence and emotionalization of the vehicle cabin. SUMMARY
[0005] The present application provides a vehicle cabin control method, device, equipment, storage medium and product to improve the accuracy of recognizing user emotions and meet the user's demand for the intelligence and emotionalization of the vehicle cabin.
[0006] In a first aspect, an embodiment of the present application provides a vehicle cabin control method, which comprises:
[0007] Obtaining multi-modal emotion information of a user in a cabin;
[0008] Performing emotion recognition based on the multi-modal emotion information through a trained emotion recognition model to obtain an emotion score of the user; the emotion recognition model is obtained through model training of a sample training set and a positive and negative sample set, and the positive and negative sample pairs contained in the positive and negative sample set are respectively generated by constructing sample pairs according to the hierarchical single-modal representation sequence of the sample data in the sample training set;
[0009] Controlling the devices in the cabin according to the control mode corresponding to the emotion score.
[0010] In a second aspect, embodiments of the present invention provide a vehicle cockpit control device, the device comprising:
[0011] The information acquisition module is used to acquire multimodal emotional information of users in the cockpit;
[0012] The scoring module is used to perform emotion recognition based on the multimodal emotion information using a trained emotion recognition model to obtain the user's emotion score; the emotion recognition model is obtained by training the model through a sample training set and a set of positive and negative samples, and the positive and negative sample pairs contained in the set of positive and negative samples are generated by constructing sample pairs hierarchically from the single-modal representation sequence of the sample data in the sample training set;
[0013] The control module is used to control the equipment in the cockpit according to the control method corresponding to the emotion score.
[0014] Thirdly, embodiments of the present invention provide a vehicle, the vehicle comprising:
[0015] At least one processor;
[0016] and a memory communicatively connected to the at least one processor;
[0017] The memory stores a computer program that can be executed by the at least one processor, which is then executed by the at least one processor to enable the at least one processor to perform the vehicle cockpit control method according to any embodiment of the present invention.
[0018] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing computer instructions, which are used to cause a processor to execute and implement the vehicle cockpit control method described in any embodiment of the present invention.
[0019] Fifthly, embodiments of the present invention also provide a computer program product, the computer program product including a computer program, which, when executed by a processor, implements the vehicle cockpit control method according to any embodiment of the present invention.
[0020] The technical solution of this invention involves acquiring multimodal emotional information of users in the cockpit; performing emotional recognition based on the multimodal emotional information using a trained emotional recognition model to obtain the user's emotional score. The emotional recognition model is obtained through model training using a sample training set and a set of positive and negative samples. The positive and negative sample pairs contained in the set of positive and negative samples are generated by constructing sample pairs hierarchically from the single-modal representation sequence of the sample data in the sample training set; and controlling the equipment in the cockpit according to the control method corresponding to the emotional score. This method trains the model by adding a set of positive and negative samples constructed from the single-modal representation sequences of the sample data, rather than introducing additional network parameters or structures into the emotion recognition model. This not only allows the trained emotion recognition model to fully learn the distributed representations, improving the accuracy of emotion recognition for users in the cockpit, but also improves the efficiency of emotion recognition, meeting the high real-time requirements of emotion recognition in smart cockpit scenarios. Furthermore, by constructing rich pairs of positive and negative samples, more comprehensive information mining and utilization of the limited training set is achieved, further improving the accuracy and reliability of emotion recognition. Based on this, the control methods corresponding to the emotion scores are used to control the equipment in the cockpit, enhancing the matching degree between vehicle cockpit control and user emotional states, improving the intelligence and emotionality of vehicle cockpit control, and ultimately improving the user experience.
[0021] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 A flowchart of a vehicle cockpit control method provided in an embodiment of the present invention;
[0024] Figure 2 This is a schematic diagram of the framework for training an initial recognition model in a vehicle cockpit control method provided in an embodiment of the present invention.
[0025] Figure 3 This is a schematic diagram of the structure of a vehicle cockpit control device provided in an embodiment of the present invention;
[0026] Figure 4A schematic diagram of the structure of a vehicle that can be used to implement an embodiment of the present invention is shown. Detailed Implementation
[0027] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0028] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0029] It's worth noting that the level of intelligence in a vehicle's cockpit is gradually becoming one of the important standards for measuring automotive quality. Passenger experience, as a key dimension for evaluating a vehicle's intelligent cockpit, is becoming increasingly important. Accurately sensing passengers' emotional state enables the provision of personalized and intelligent services, thereby significantly improving passenger comfort and satisfaction.
[0030] However, existing multimodal sentiment analysis techniques have many shortcomings when applied to the perception of emotions among passengers in vehicle cabins. Current multimodal sentiment analysis algorithms primarily focus on the design of the fusion module, neglecting the optimization and design of representation learning. They simply extract single-modal representations and directly input them into the fusion module to obtain multimodal representations, resulting in insufficient representation learning and difficulty in fully capturing the various correlations between cross-modal representations, thus affecting the accuracy and reliability of emotion state recognition.
[0031] Furthermore, the existing datasets for multimodal sentiment analysis are limited in size. Using only the sentiment labels of the datasets for multimodal sentiment analysis tasks fails to fully and comprehensively mine the information in the multimodal data, further reducing the accuracy of perceiving passenger emotions. This makes it difficult for the vehicle cabin system to accurately adjust based on passenger emotions, thus failing to meet users' needs for intelligent, personalized, and emotional vehicle cabins.
[0032] Based on this, embodiments of the present invention provide a vehicle cockpit control method. Figure 1 The flowchart illustrates a vehicle cockpit control method provided in an embodiment of the present invention. This embodiment is applicable to scenarios where cockpit equipment is intelligently controlled based on the emotional state of the user in the cockpit. The method can be executed by a vehicle cockpit control device, which can be implemented in software and / or hardware, and optionally, through a vehicle.
[0033] like Figure 1 As shown, the vehicle cockpit control method provided in this embodiment of the invention may specifically include:
[0034] S101. Obtain multimodal emotional information of users in the cockpit.
[0035] Multimodal emotional information can be understood as information in multiple modalities that can reflect a user's emotional state. Multimodalities can include text modality, voice modality, visual modality, and physiological modality.
[0036] In this embodiment, the user's voice intonation information can be obtained through a microphone array or voice interaction device in the cockpit. After obtaining the voice intonation information, it can be divided into multiple voice segments, and a voice sequence can be constructed based on each voice segment. The voice sequence is used as voice modal emotional information. Alternatively, the voice intonation information or voice modal emotional information can be automatically recognized through a speech recognition algorithm or speech recognition model to obtain the user's text expression information. By dividing the characters in the text expression information, a text sequence can be constructed based on each character, and the text sequence is used as text modal emotional information.
[0037] In this embodiment, the user's visual emotional information, especially visual facial expression information, can also be acquired through in-cabin image acquisition devices, such as cameras. After acquiring the visual emotional information, each video frame in the visual emotional information can be extracted, and a visual sequence can be constructed based on each video frame, using the visual sequence as visual modal emotional information.
[0038] In an alternative embodiment, audio segments and video frames can be filtered to obtain audio segments and video frames that are more relevant to the user's emotional state.
[0039] In this embodiment, multimodal emotional information is constructed based on emotional information from each modality.
[0040] S102. Based on multimodal emotional information, the trained emotion recognition model is used to perform emotion recognition and obtain the user's emotion score. The emotion recognition model is obtained by training the model through a sample training set and a set of positive and negative samples. The positive and negative sample pairs contained in the set of positive and negative samples are generated by constructing sample pairs hierarchically from the single-modal representation sequence of the sample data in the sample training set.
[0041] The emotion recognition model can be understood as a model used to identify a user's emotions or emotional state. It outputs an emotion score reflecting the emotional tendency and intensity based on emotional information, including positive and negative emotional tendencies. The training set includes multiple sample data sets and an emotion score label for each sample data set. Each sample data set includes multiple sample sequences of different modalities. The positive and negative sample sets can be considered as a set of positive samples and a set of negative samples, and positive and negative sample pairs can be considered as a pair of positive samples and a pair of negative samples. The positive sample set includes at least one positive sample pair, and the negative sample set includes at least one negative sample pair.
[0042] In this embodiment, multimodal emotional information can be input into a trained emotion recognition model. The emotion recognition model extracts single-modal representations of each modality of emotional information from the multimodal emotional information, fuses the extracted representation sequences, and then classifies or regresses the fusion results to obtain the user's emotion score.
[0043] In one alternative implementation, sample data from the training set is input into an untrained emotion recognition model to extract unimodal representations, resulting in a sequence of unimodal representations of the sample data. Positive and negative sample sets can be generated by constructing sample pairs from each unimodal representation sequence hierarchically.
[0044] For example, the hierarchy may include an initial level that captures the information correlation within a modality, a middle level that captures the information correlation between modalities to alleviate the problem of modality distribution differences, and a high level that captures the information correlation between modalities within a single sample to improve the interactive capture capability of sentiment differentiation.
[0045] It is understandable that constructing sample pairs for each single-modal representation sequence at each level can generate the corresponding positive sample set and / or negative sample set for that level.
[0046] S103. Control the equipment in the cockpit according to the control method corresponding to the emotion score.
[0047] In this embodiment, the control method corresponding to the emotion score can be determined according to the preset mapping table, control rules, etc., and the control method can be used to interact with the various control modules in the cockpit to regulate the corresponding equipment in the cockpit.
[0048] For example, control methods may include adjusting the interaction strategy of the in-vehicle voice assistant; for instance, if the emotion score is lower than a preset threshold, the voice assistant's tone will be made gentler, and the frequency of system prompts will be reduced. Control methods may also include adjusting the cabin lighting; if the emotion score is lower than a preset threshold, the lighting will be adjusted to a softer light to alleviate negative emotions. Control methods may also include adjusting in-vehicle music recommendations and playback; if the emotion score is lower than a preset threshold, tracks that alleviate negative emotions will be played, which can be user-preset or system-set tracks. Control methods may further include adjusting the in-vehicle infotainment display interface; for example, displaying soothing visual content.
[0049] The above-described technical solution in this embodiment acquires multimodal emotional information of the user in the cockpit; performs emotional recognition based on the multimodal emotional information using a trained emotional recognition model to obtain the user's emotional score. The emotional recognition model is obtained through model training using a sample training set and a set of positive and negative samples. The positive and negative sample pairs contained in the set of positive and negative samples are generated by constructing sample pairs hierarchically from the single-modal representation sequence of the sample data in the sample training set; and controls the equipment in the cockpit according to the control method corresponding to the emotional score. This method trains the model by adding a set of positive and negative samples constructed from the single-modal representation sequences of the sample data, rather than introducing additional network parameters or structures into the emotion recognition model. This not only allows the trained emotion recognition model to fully learn the distributed representations, improving the accuracy of emotion recognition for users in the cockpit, but also improves the efficiency of emotion recognition, meeting the high real-time requirements of emotion recognition in smart cockpit scenarios. Furthermore, by constructing rich pairs of positive and negative samples, more comprehensive information mining and utilization of the limited training set is achieved, further improving the accuracy and reliability of emotion recognition. Based on this, the control methods corresponding to the emotion scores are used to control the equipment in the cockpit, enhancing the matching degree between vehicle cockpit control and user emotional states, improving the intelligence and emotionality of vehicle cockpit control, and ultimately improving the user experience.
[0050] As a first optional embodiment of the present invention, based on the above embodiments, the emotion recognition model can be obtained by training the model through a sample training set and a set of positive and negative samples, and is specified as follows:
[0051] a1) Obtain a sample training set and a pre-built initial recognition model. The sample training set includes multiple sample data and sentiment score labels corresponding to each sample data. Each sample data includes multiple sample sequences of different modalities. The initial recognition model includes a single-modal representation extraction sub-model and a multi-modal fusion sub-model.
[0052] The initial recognition module can be understood as a model with training requirements. The sentiment score label can be understood as a label corresponding to the sample data, reflecting the sentiment tendency and intensity of the sample data. Each sample data includes multiple sample sequences of different modalities. For example, the sample data may include at least two sample sequences from text sample sequences, speech sample sequences, and visual sample sequences corresponding to the same user in the cabin at the same time period. The text sample sequence includes multiple sample characters, the speech sample sequence includes multiple sample speech segments, and the visual sample sequence includes multiple sample video frames.
[0053] The unimodal representation extraction sub-model can be considered as a model used to extract unimodal representations for each modality of data. In this optional embodiment, each modality of data can be considered as a sample sequence for each modality. The multimodal fusion sub-model can be understood as a model used to fuse the unimodal representations and evaluate the sentiment state based on the fusion result.
[0054] In this embodiment, one sample data can correspond to one sentiment score label, or one sample data can correspond to multiple sentiment score labels, and each sentiment score label corresponds to a sample sequence in the sample data.
[0055] b1) Input the sample data into the initial recognition model to obtain the single-modal representation sequence output by the single-modal representation extraction sub-model. Each single-modal representation sequence corresponds one-to-one with each sample sequence in the sample data.
[0056] In this embodiment, each sample sequence in the sample data is input into the unimodal representation extraction sub-model included in the initial recognition model. The unimodal representation extraction sub-model performs processing such as partitioning, filtering, encoding, mapping, and / or pooling on each sample sequence to obtain the unimodal representation sequence corresponding to each sample sequence. The unimodal representation sequence may include text representation sequence, speech representation sequence, and visual representation sequence.
[0057] As one implementation, the sample data includes text sample sequences, speech sample sequences, and visual sample sequences. Accordingly, the step of inputting the sample data into the initial recognition model to obtain the single-modal representation sequences output by the single-modal representation extraction sub-model can be further optimized as follows:
[0058] b11) Input the text sample sequence into the first encoder included in the single-modal representation extraction sub-model, and encode the text sample sequence through the first encoder to obtain the first representation vector.
[0059] The first encoder can be considered as an encoder that encodes the text sample sequence under the text modality in order to extract and represent the text features of the text sample sequence, thereby obtaining the first representation vector of the text modality corresponding to the text sample sequence.
[0060] It is understandable that data from different modalities have different features and structures. In order to capture the unique information in the text modality, in this embodiment, the first encoder can choose a Transformer-based bidirectional representation encoder (BERT encoder) to extract representation information such as the sentiment expression grammar structure in the text sample sequence.
[0061] For example, the text sample sequence U is processed by the first encoder (BERT encoder). l Encode to obtain the first representation vector The method can be specifically expressed as:
[0062]
[0063] Where T represents the time dimension; d l Θ represents the dimension of the text sample sequence; l These are the parameters of the first encoder.
[0064] b12) Input the speech sample sequence and the visual sample sequence into the second encoder included in the single-modal representation extraction sub-model, respectively, and encode the speech sample sequence and the visual sample sequence through the second encoder to obtain the second representation vector and the third representation vector.
[0065] The second encoder can be considered as an encoder used to encode speech sample sequences in the speech modality and visual sample sequences in the visual modality to extract speech features from the speech sample sequences and visual features from the visual sample sequences, thereby obtaining a second representation vector for the speech modality corresponding to the speech sample sequence and a third representation vector for the visual modality corresponding to the visual sample sequence. Optionally, the second encoder can be a Transformer model. The second encoder can extract relevant representation information such as speech, intonation, and speech rate from the speech sample sequences; and it can also extract relevant representation information such as facial expressions and body language from the visual sample sequences.
[0066] For example, a second representation vector is obtained by encoding a speech sample sequence or a visual sample sequence using a second encoder. The method can be specifically expressed as:
[0067]
[0068] Where, d mdenoted as the representation dimension of the speech sample sequence or the visual sample sequence; a represents the speech sample sequence; v represents the visual sample sequence.
[0069] b13) Map each representation vector obtained through encoding according to a preset representation dimension, and perform average pooling on the representation vectors of each preset representation dimension obtained by mapping, so as to obtain each single-modal representation sequence corresponding to each sample sequence in the sample sequence data.
[0070] It is understandable that, since sample sequences of different modalities usually have different feature dimensions and data structures, the representation dimensions of the representation vectors corresponding to each sample sequence obtained through encoding may also be different.
[0071] To facilitate cross-modal fusion and improve the scalability and consistency of the model, in this embodiment, the first, second, and third representation vectors obtained by encoding can be mapped to the same preset representation dimension through a linear layer, so that the three representation vectors have consistent feature lengths. Then, average pooling can be performed on each representation vector in the time dimension to obtain the sequence-level single-modal representation corresponding to each sample sequence.
[0072] For example, the method of obtaining the mapped second representation vector by mapping the second representation vector according to a preset representation dimension through a linear layer can be specifically represented as follows:
[0073]
[0074] Among them, X m The second representation vector after mapping; l is the text sample sequence; W pro_m d is the weight matrix of the linear layer; b is the bias vector of the linear layer; d is the representation dimension of each unimodal representation sequence; l is the text sample sequence.
[0075] For example, the method of performing average pooling on the representation vectors of each preset representation dimension obtained by mapping can be specifically expressed as follows:
[0076]
[0077] Where, X′ m It is a single-modal characterization sequence.
[0078] The above-described technical solution in this embodiment obtains more accurate single-modality representation vectors by using different encoding models to extract targeted features from sample sequences of different modalities. By mapping each representation vector obtained through encoding according to a preset representation dimension and performing average pooling on the representation vectors of each preset representation dimension obtained by mapping, it provides convenience and support for subsequent fusion operations and the construction of positive and negative sample pairs.
[0079] c1) Input each of the single-modal representation sequences into the sample pair construction auxiliary model and the multimodal fusion sub-model respectively to obtain the positive and negative sample sets corresponding to each of the single-modal representation sequences output by the sample pair construction auxiliary model, and the predicted sentiment scores corresponding to the sample data output by the multimodal fusion sub-model.
[0080] Among them, the sample pair construction auxiliary model can be considered as a model used to construct corresponding positive and negative sample pairs for each single-modal representation sequence, so as to assist the initial recognition model in training through the constructed positive and negative sample pairs.
[0081] It should be noted that the sample pair auxiliary model is the model configured during the training of the initial recognition model. After obtaining the emotion recognition model, it is not necessary to configure the sample pair auxiliary model during the actual emotion recognition process.
[0082] In this embodiment, each unimodal representation sequence is input into a sample pair construction auxiliary model. Relative to each unimodal representation sequence, the sample pair construction auxiliary model constructs positive and negative samples for the unimodal representation sequence based on a preset positive-negative sample pair definition. Each positive sample and unimodal representation sequence can form a positive sample pair, and each negative sample and unimodal representation sequence can form a negative sample pair. The positive and negative sample pairs are summarized to obtain the positive and negative sample sets for each unimodal representation sequence. The positive and negative sample sets corresponding to each unimodal representation sequence can be used as the positive and negative sample sets corresponding to that unimodal representation sequence, or the positive and negative sample sets corresponding to each unimodal representation sequence can be summarized separately to obtain the common positive and negative sample set corresponding to all unimodal representation sequences.
[0083] In this embodiment, each of the single-modal representation sequences is input into a multimodal fusion sub-model. The multimodal fusion sub-model fuses the single-modal representation sequences corresponding to each sample data. Based on the fusion result, a classification model or regression model with sentiment score prediction is used to obtain the predicted sentiment score corresponding to each sample data.
[0084] As one implementation, each of the single-modal representation sequences is input into the multimodal fusion sub-model to obtain the predicted sentiment score corresponding to the sample data output by the multimodal fusion sub-model, including:
[0085] c11) The input single-modal representation sequences are concatenated through the cascaded layer of the multimodal fusion sub-model to obtain the initial representation sequence of the sample data.
[0086] The initial representation sequence can be understood as the representation sequence obtained by cascading the single-modal representation sequences of the sample data.
[0087] In this embodiment, for each sample data, the input single-modal representation sequences relative to that sample data are concatenated through a cascaded layer to obtain the initial representation sequence of the sample data. Specifically, this can be represented as follows:
[0088]
[0089] X M ∈R 3d ;
[0090] Among them, X M X is the initial representation sequence of the sample data; l X is the text representation sequence of the sample data; a X is the speech representation sequence of the sample data; v R is a visual representation sequence of the sample data. 3d Let represent a 3d-dimensional real vector space.
[0091] c12) The initial representation sequence is mapped according to a preset representation dimension through the first linear layer of the multimodal fusion sub-model, and the initial representation sequence with the preset representation dimension obtained by the mapping is processed by a nonlinear activation function.
[0092] The first linear layer can be understood as a network layer used for representation dimension transformation.
[0093] In this embodiment, the initial representation sequence can be mapped to a preset representation dimension through the first linear layer of the multimodal fusion sub-model. For example, the preset representation dimension is d. If the initial representation sequence is 3d-dimensional, then the mapped initial representation sequence will be d-dimensional.
[0094] For example, the method of mapping the initial representation sequence according to a preset representation dimension through a first linear layer, and then processing the initial representation sequence of the preset representation dimension obtained by the mapping through a nonlinear activation function, can be specifically represented as follows:
[0095] X′ M ←(X M W1+b1)∈R d ;
[0096]
[0097] Where W1 is the weight matrix of the first linear layer; b1 is the bias vector of the first linear layer; X′ M The initial representation sequence of the preset representation dimension obtained by mapping; X″ M This is the initial representation sequence after processing with a nonlinear activation function.
[0098] c13) The processed initial representation sequence is subjected to deep representation extraction through the second linear layer of the multimodal fusion sub-model to obtain a multimodal representation sequence.
[0099] The second linear layer can be understood as a network layer used for deep representation extraction.
[0100] For example, a multimodal representation sequence X″′ is obtained by performing deep representation extraction on the processed initial representation sequence through a second linear layer. M The method can be specifically expressed as:
[0101] X″′ M ←(X″ M W2+b2)∈R d
[0102] Where W2 is the weight matrix of the second linear layer; b2 is the bias vector of the second linear layer.
[0103] c14) Input the multimodal representation sequence into the regression layer of the multimodal fusion sub-model to obtain the predicted sentiment score corresponding to the sample data output by the regression layer.
[0104] The regression layer can be understood as a model that maps the input representation sequence to continuous numerical values, and the numerical value output by the regression layer can be considered as the predicted sentiment score corresponding to the sample data.
[0105] For example, the predicted sentiment score corresponds to the interval [-3, 3]. A predicted sentiment score less than 0 represents a negative sentiment tendency, and a predicted sentiment score greater than 0 represents a positive sentiment tendency. The magnitude of the predicted sentiment score can represent the intensity of the sentiment.
[0106] In this embodiment, the above-described technical solution concatenates the input single-modal representation sequences through a cascaded layer in the multimodal fusion sub-model to obtain an initial representation sequence capable of representing the multimodal features of the initial sample data. By mapping the initial representation sequence to a preset representation dimension, dimensionality reduction is achieved, making subsequent representation extraction more effective. A nonlinear activation function is introduced after the first linear layer, enabling the multimodal fusion sub-model to learn more complex feature representations. A second linear layer performs deep representation extraction on the processed initial representation sequence, thereby obtaining a more accurate and effective multimodal representation sequence. Based on the multimodal representation sequence, a regression layer is used to obtain the predicted sentiment score corresponding to the sample data, improving the accuracy of the predicted sentiment score.
[0107] d1) Determine the target loss function value based on the set of positive and negative samples, the predicted sentiment score, and the sentiment score label, combined with the set loss function.
[0108] The objective loss function can be considered as a loss function used to assist the network model in adjusting network parameters.
[0109] In this embodiment, by comparing positive and negative sample pairs in the positive and negative sample sets and combining them with the comparison loss function, the relative loss function values of the positive and negative sample sets are determined. Based on the predicted sentiment score and sentiment score label, and combined with loss functions such as mean absolute error loss function and mean squared error loss function, the loss function value corresponding to the predicted sentiment score is determined. By fusing the relative loss function values of the positive and negative sample sets and the loss function value corresponding to the predicted sentiment score, the target loss function value is obtained.
[0110] e1) Based on the target loss function value, the network parameters in the initial recognition model are back-learned and adjusted to obtain the adjusted initial recognition model. Then, the input operation of the sample data is re-executed until the training end condition is met. The initial recognition model obtained after training is determined as the emotion recognition model.
[0111] In this embodiment, the training termination condition may include the maximum number of iterations, the accuracy threshold, and the total number of sample data.
[0112] The above-described technical solution in this embodiment obtains the single-modal representation sequence corresponding to each sample sequence in the sample data through a single-modal representation extraction sub-model, and constructs a positive and negative sample set for each single-modal representation sequence based on the single-modal representation sequence through sample pair construction auxiliary model, thereby achieving more comprehensive information mining and utilization of the limited sample training set; it fuses the single-modal representation sequences of the sample data through a multi-modal fusion sub-model, and determines the predicted sentiment score corresponding to the sample data based on the fusion result, thereby improving the accuracy and reliability of the predicted sentiment score; based on the positive and negative sample set, the predicted sentiment score, and the sentiment score label, it determines the target loss function value in combination with the set loss function, and uses the target loss function to guide the initial recognition model to adjust the network parameters, so that the initial recognition model can fully learn the consistent representation, improving the representation extraction performance of the initial recognition model, thereby improving the accuracy and reliability of the initial recognition model's sentiment recognition.
[0113] As one implementation of step c1), the sample pair construction auxiliary model includes at least one of a primary-level modal intra-modal aggregation sub-model, a mid-level inter-modal alignment sub-model, and a high-level sample intra-linkage sub-model; correspondingly, inputting each of the single-modal representation sequences into the sample pair construction auxiliary model to obtain the positive and negative sample sets corresponding to each of the single-modal representation sequences output by the sample pair construction auxiliary model can be further optimized to the following steps:
[0114] Each of the single-modal representation sequences is input into the initial-level intramodal aggregation sub-model to obtain a first positive sample set and a first negative sample set output by the initial-level intramodal aggregation sub-model; each of the single-modal representation sequences is input into the intermediate-level intermodal alignment sub-model to obtain a second positive sample set and a second negative sample set output by the intermediate-level intermodal alignment sub-model; each of the single-modal representation sequences is input into the high-level sample intraconnection sub-model to obtain a third positive sample set output by the high-level sample intraconnection sub-model.
[0115] It should be noted that the core of constructing an auxiliary model using sample pairs lies in building a phased, progressively more difficult learning process. This process can include the following main levels: initial level, intermediate level, and high level. Therefore, the auxiliary model constructed using sample pairs can include an initial level intramodal aggregation sub-model, an intermediate level intermodal alignment sub-model, and a high level sample intra-connection sub-model. In each sub-model, comparative learning is used to assist in the gradual optimization of the emotion recognition task, thereby improving the initial recognition model's representation learning ability for multimodal emotion analysis tasks.
[0116] The primary-level modal aggregate sub-model can be understood as a way to capture the information correlation between single-modal representation sequences within each modality, thereby improving the intrinsic consistency of single-modal representations, especially in sample pairs with similar emotions within the same modality, and reducing distribution perturbations caused by noise or feature bias.
[0117] The mid-level intermodal alignment sub-model can be understood as capturing the correlation between unimodal representation sequences in different modalities, bringing unimodal representation sequences expressing the same emotion closer together in a unified feature space while preserving supplementary information brought about by modal differences. Compared to the low-level intramodal aggregation sub-model, the mid-level intermodal alignment sub-model constructs a richer sample set and considers both differences in emotion score labels and modal characteristics. While the mid-level intermodal alignment sub-model can reduce the distributional differences of cross-modal sample pairs, it may have some impact on the cross-modal information connections within the same sample data.
[0118] Therefore, the high-level sample intra-connection sub-model can be understood as the ability to capture the information correlation between single-modal representation sequences within sample data, maintain comprehensive information discrimination capabilities, and thus improve the ability of the initial recognition model to capture the interaction of different emotional information components in the same sample data.
[0119] In this embodiment, each single-modal representation sequence is input into the initial-level modal aggregate sub-model to obtain the first positive sample set and the first negative sample set output by the initial-level modal aggregate sub-model.
[0120] In an optional embodiment, the first positive sample set includes at least one first positive sample pair, which consists of two single-modal representation values belonging to the same modality, having the same sentiment tendency, and having a sentiment score label difference value less than the first positive sample selection threshold. The single-modal representation values are determined based on the single-modal representation sequence. The first negative sample set includes at least one first negative sample pair, which consists of two single-modal representation values belonging to the same modality, having different sentiment tendencies, and having a sentiment score label difference value greater than the first negative sample selection threshold.
[0121] In this embodiment, for each single-modal representation value, the single-modal representation value can be obtained by regularizing the feature values in the single-modal representation sequence, specifically as follows:
[0122]
[0123] Where x is the numerical value representing a single mode; x i is the i-th eigenvalue in the single-modal representation sequence; s is the total number of eigenvalues in the single-modal representation sequence.
[0124] For example, the methods for determining the first positive sample set A1(p) and the first negative sample set A1(n) can be expressed as follows:
[0125] A1(p) = {p ≠ a, |y a -y p | <w ap};
[0126] A1(n)={n≠a,|y a -y n |>w an};
[0127] Among them, y a Let y be the sentiment score label for the a-th unimodal representation sequence, which can also be considered as the anchor unimodal representation sequence; p w represents the sentiment score label for the p-th unimodal representation sequence. ap Select a threshold for the first positive sample; w an Select a threshold for the first negative sample.
[0128] It is understandable that after the sample training set is input into the initial recognition model, the data used to construct positive and negative sample pairs can be sample data belonging to the same batch for model training, which is determined in advance by dividing the sample training set.
[0129] In this embodiment, each of the single-modal representation sequences is input into the intermediate-level inter-modal alignment sub-model to obtain the second positive sample set and the second negative sample set output by the intermediate-level inter-modal alignment sub-model.
[0130] In one optional embodiment, the second positive sample set includes at least one second positive sample pair, which consists of two single-modal representation values belonging to different modalities, having the same sentiment tendency, and having a sentiment score label difference value less than the second positive sample selection threshold. The second positive sample selection threshold is the difference between the first positive sample selection threshold and the first threshold compensation value, and the first threshold compensation value is used to compensate for the first positive sample selection threshold. The second negative sample set includes at least one second negative sample pair, which consists of two single-modal representation values belonging to different modalities, having different sentiment tendencies, and having a sentiment score label difference value greater than the second negative sample selection threshold. The second negative sample selection threshold is the sum of the first negative sample selection threshold and the second threshold compensation value, and the second threshold compensation value is used to compensate for the first negative sample selection threshold.
[0131] Optionally, the first threshold compensation value can be the same as the second threshold compensation value.
[0132] It is understandable that different modalities contain their own modality-specific information, and cross-modal sample pairs will inevitably have greater distribution differences. Adding a threshold compensation value to the selection of sentiment score label difference value can preserve certain modality-specific information during the learning and optimization of the alignment sub-model between mid-level modalities.
[0133] For example, the way to determine the second positive sample set A2(p) and the second negative sample set A2(n) can be expressed as follows:
[0134] A2(p)={p≠a,|y a -y p | <w ap -m1};
[0135] A2(n)={n≠a,|y a -y n |>w an +m2};
[0136] Where m1 is the first threshold compensation value; m2 is the second threshold compensation value.
[0137] Each of the single-modal representation sequences is input into the high-level sample intralinkage sub-model to obtain the third positive sample set output by the high-level sample intralinkage sub-model.
[0138] In an optional embodiment, the third positive sample set includes at least one third positive sample pair, which consists of two single-modal representation values belonging to the same sample data but belonging to different modalities.
[0139] Understandably, in the high-level sample intra-connection sub-model, only positive sample pairs are considered, and positive sample pairs are defined as sample pairs from the same sample data but different modalities, in order to improve the ability of the initial recognition model to capture the information correlation within the same sample data.
[0140] In one alternative implementation, the first positive sample selection threshold increases with the number of iterations, and the first negative sample selection threshold decreases with the number of iterations.
[0141] Understandably, in sentiment analysis and recognition tasks, sentiment score labels are typically represented as discrete categories or continuous scores. Traditional contrastive learning methods categorize sentiment into positive and negative sentiment, using fixed positive and negative sample pairs for construction, which easily overlooks the gradual information implicit in continuous labels.
[0142] Therefore, in this embodiment, a method of gradually increasing the thresholds is adopted. The first positive sample selection threshold and the first negative sample selection threshold can be designed to change with the number of iterations during the iterative training of the initial recognition model, thereby achieving dynamic construction of positive and negative sample pairs. By gradually widening the selection range of positive samples as training progresses, more similar but not identical single-modal representation sequences are covered. Simultaneously, as the first positive sample selection threshold increases, the first negative sample selection threshold is gradually narrowed, reducing the selection range of negative samples. This strategy of gradually adjusting the thresholds allows the initial recognition model to focus on high-precision emotional feature extraction in the early training stages, and gradually learn the global relationships of emotional distribution in later stages.
[0143] The technical solution described in this embodiment constructs sample pairs through multi-level sub-models, achieving effective representation learning of multimodal sample data. Simultaneously, the progressive sample pair construction strategy fully considers the continuity of sentiment score labels in intensity, enabling the initial recognition model to balance local fine-grained learning and global information capture, fully preserving detailed information and further enhancing the modeling ability for the continuity of sentiment score labels. This improves the accuracy of sentiment recognition and overcomes the problem of existing multimodal sentiment analysis methods ignoring the differences between different sentiment intensities in their task definition, relying solely on positive and negative sentiment for sentiment recognition and module design, resulting in the loss of detailed information and reduced accuracy of sentiment recognition.
[0144] As one implementation of step d1), the positive and negative sample sets include a first positive sample set, a first negative sample set, a second positive sample set, a second negative sample set, and a third positive sample set.
[0145] Accordingly, the step of determining the target loss function value based on the set of positive and negative samples, the predicted sentiment score, and the sentiment score label, combined with a set loss function, can be further optimized into the following steps:
[0146] d11) For each single-modal representation sequence, based on the single-modal representation sequence and the positive and negative samples corresponding to the single-modal representation sequence in the first positive sample set and the first negative sample set, respectively, calculate the first function value of the first contrastive learning loss function.
[0147] The first contrastive learning loss function can be understood as a function used to assist the initial recognition model in learning the representation distribution of the single-modal representation sequence within the same modality. The first function value can be understood as the corresponding loss function value determined based on the first contrastive learning loss function.
[0148] For example, the first contrastive learning loss function L low It can be represented as:
[0149]
[0150] Where sim(·) is the similarity calculation function, used to calculate the similarity score of sample pairs; x a For the single-mode characterization value of the anchor point, x p The positive sample single-mode representation value corresponding to the anchor point single-mode representation value; x n τ1 represents the negative sample single-mode representation value corresponding to the anchor single-mode representation value; τ1 is the temperature coefficient in the first contrastive learning loss function that controls the smoothness of the sample distribution.
[0151] d12) Based on the single-modal representation sequence, and the positive and negative samples corresponding to the single-modal representation sequence in the second positive sample set and the second negative sample set, respectively, calculate the second function value of the second contrastive learning loss function. The second contrastive learning loss function is the sum of the first contrastive learning loss function and the function compensation value. The function compensation value is used to compensate the first contrastive learning loss function.
[0152] The second contrastive learning loss function can be understood as a function used to assist the initial recognition model in learning the representation distribution of single-modal representation sequences among different modalities. The second function value can be considered as the corresponding loss function value determined based on the second contrastive learning loss function.
[0153] It is understandable that different modalities contain their own modality-specific information, and cross-modal sample pairs inevitably have greater distribution differences. Therefore, in this embodiment, a function compensation value can be added to the first contrastive learning loss function to form a second contrastive learning loss function, thereby relaxing the optimization requirements of the initial recognition model and retaining certain modality-specific information during the learning and optimization process of the high-level sample intra-connection sub-model.
[0154] For example, the second contrastive learning loss function L mid It can be represented as:
[0155]
[0156] Where τ2 is the temperature coefficient in the second contrastive learning loss function, which controls the smoothness of the sample distribution; m3 is the function compensation value.
[0157] d13) Based on the single-modal representation sequence and the positive samples in the third positive sample set corresponding to the single-modal representation sequence, calculate the third function value of the second-norm loss function.
[0158] The second-norm loss function can be understood as a function used to assist the initial recognition model in learning the representation distribution of single-modal representation sequences among different modalities in the same sample data. The third function value can be considered as the corresponding loss function value determined based on the second-norm loss function.
[0159] For example, the second-order norm loss function L high It can be represented as:
[0160] L high =‖x a -x p ‖2.
[0161] d14) Based on the predicted sentiment score and the sentiment score label, calculate the fourth function value of the mean absolute error loss function.
[0162] The mean absolute error loss function can be understood as a quantified sentiment score label used to assist the initial recognition model in learning the specific representation of sentiment in the data. The fourth function value can be considered as the corresponding loss function value determined based on the second-order norm loss function.
[0163] For example, the mean absolute error loss function can be expressed as:
[0164]
[0165] Where y is the sentiment score label; To predict sentiment scores.
[0166] d15) The first function value, the second function value, the third function value, and the fourth function value are weighted and summed to obtain the target loss function.
[0167] For example, the target loss function L is obtained. loss The method can be expressed as:
[0168] L loss =λ1L task +λ2L low +λ3L mid +λ4L high ;
[0169] Where λ1, λ2, λ3 and λ4 are respectively L task L low L mid and L high The hyperparameter weight coefficients are used to balance the contributions of different sub-models to the network learning.
[0170] The above-described technical solution in this embodiment obtains the target loss function by weighted summing of the first function value, the second function value, the third function value, and the fourth function value to train the initial recognition model. This improves the initial recognition model's ability to capture the information correlation within each modality, the potential information correlation across modalities, and the interaction of different emotional information components of the same sample.
[0171] To better understand the method for training the initial recognition model in a vehicle cockpit control method provided in this embodiment of the invention, a specific example is given here. Figure 2 This is a schematic diagram of the framework for training an initial recognition model in a vehicle cockpit control method provided in an embodiment of the present invention.
[0172] like Figure 2 As shown, each sample data includes a text sample sequence, a speech sample sequence, and a visual sample sequence. The sample data is input into the unimodal representation extraction sub-model in the initial recognition model to obtain the unimodal representation sequences output by the unimodal representation extraction sub-model. Each unimodal representation sequence includes a text representation sequence, a speech representation sequence, and a visual representation sequence. Each unimodal representation sequence is then input into the initial-level modal intra-modal aggregation sub-model, the intermediate-level inter-modal alignment sub-model, and the high-level sample intra-linkage sub-model, which are used to construct auxiliary models for sample pairs. This yields the first positive sample set and the first negative sample set output by the initial-level modal intra-aggregation sub-model, the second positive sample set and the second negative sample set output by the intermediate-level inter-modal alignment sub-model, and the third positive sample set output by the high-level sample intra-linkage sub-model, respectively.
[0173] Furthermore, each unimodal representation sequence is input into a multimodal fusion sub-model, which includes a cascaded layer, a first linear layer, a second linear layer, and a regression layer. The cascaded layer concatenates the unimodal representation sequences of each input sample data to obtain an initial representation sequence for the sample data. Then, the first linear layer maps the initial representation sequence according to a preset representation dimension, and a nonlinear activation function processes the mapped initial representation sequence with the preset representation dimension to obtain a processed initial representation sequence. Next, the second linear layer performs deep representation extraction on the processed initial representation sequence to obtain a multimodal representation sequence. Finally, the regression layer obtains the predicted sentiment score corresponding to the sample data.
[0174] Figure 3 This is a schematic diagram of the structure of a vehicle cockpit control device provided in an embodiment of the present invention. Figure 3 As shown, the device includes: an information acquisition module 31, a score determination module 32, and a control module 33, wherein,
[0175] Information acquisition module 31 is used to acquire multimodal emotional information of users in the cockpit;
[0176] The scoring module 32 is used to perform emotion recognition based on the multimodal emotion information using a trained emotion recognition model to obtain the user's emotion score; the emotion recognition model is obtained by training the model through a sample training set and a set of positive and negative samples, and the positive and negative sample pairs contained in the set of positive and negative samples are generated by constructing sample pairs hierarchically from the single-modal representation sequence of the sample data in the sample training set;
[0177] The control module 33 is used to control the equipment in the cockpit according to the control method corresponding to the emotion score.
[0178] The above-described technical solution in this embodiment acquires multimodal emotional information of the user in the cockpit; performs emotional recognition based on the multimodal emotional information using a trained emotional recognition model to obtain the user's emotional score. The emotional recognition model is obtained through model training using a sample training set and a set of positive and negative samples. The positive and negative sample pairs contained in the set of positive and negative samples are generated by constructing sample pairs hierarchically from the single-modal representation sequence of the sample data in the sample training set; and controls the equipment in the cockpit according to the control method corresponding to the emotional score. This method trains the model by adding a set of positive and negative samples constructed from the single-modal representation sequences of the sample data, rather than introducing additional network parameters or structures into the emotion recognition model. This not only allows the trained emotion recognition model to fully learn the distributed representations, improving the accuracy of emotion recognition for users in the cockpit, but also improves the efficiency of emotion recognition, meeting the high real-time requirements of emotion recognition in smart cockpit scenarios. Furthermore, by constructing rich pairs of positive and negative samples, more comprehensive information mining and utilization of the limited training set is achieved, further improving the accuracy and reliability of emotion recognition. Based on this, the control methods corresponding to the emotion scores are used to control the equipment in the cockpit, enhancing the matching degree between vehicle cockpit control and user emotional states, improving the intelligence and emotionality of vehicle cockpit control, and ultimately improving the user experience.
[0179] Furthermore, the device also includes a model training module, which specifically may include:
[0180] The training set and model acquisition unit is used to acquire a sample training set and a pre-built initial recognition model. The sample training set includes multiple sample data and sentiment score labels corresponding to each sample data. Each sample data includes multiple sample sequences of different modalities. The initial recognition model includes a single-modal representation extraction sub-model and a multi-modal fusion sub-model.
[0181] The single-modal representation sequence acquisition unit is used to input the sample data into the initial recognition model to obtain each single-modal representation sequence output by the single-modal representation extraction sub-model, and each single-modal representation sequence corresponds one-to-one with each sample sequence in the sample data;
[0182] The sample construction and score prediction unit is used to input each of the single-modal representation sequences into the sample pair construction auxiliary model and the multimodal fusion sub-model, respectively, to obtain the positive and negative sample sets corresponding to each of the single-modal representation sequences output by the sample pair construction auxiliary model, and the predicted sentiment score corresponding to the sample data output by the multimodal fusion sub-model.
[0183] The loss function determination unit is used to determine the target loss function value based on the set of positive and negative samples, the predicted sentiment score, and the sentiment score label, combined with a set loss function.
[0184] The parameter adjustment unit is used to back-learn and adjust the network parameters in the initial recognition model according to the target loss function value, obtain the adjusted initial recognition model, and return to re-execute the input operation of sample data until the training end condition is met. The initial recognition model obtained after training is determined as the emotion recognition model.
[0185] Furthermore, the sample data includes text sample sequences, speech sample sequences, and visual sample sequences;
[0186] Accordingly, the single-modal characterization sequence acquisition unit can be specifically used for:
[0187] The text sample sequence is input into the first encoder included in the single-modal representation extraction sub-model, and the text sample sequence is encoded by the first encoder to obtain the first representation vector;
[0188] The speech sample sequence and the visual sample sequence are respectively input into the second encoder included in the single-modal representation extraction sub-model. The second encoder encodes the speech sample sequence and the visual sample sequence respectively to obtain the second representation vector and the third representation vector.
[0189] Each representation vector obtained through encoding is mapped according to a preset representation dimension, and the representation vectors of each preset representation dimension obtained by mapping are averaged and pooled to obtain each single-modal representation sequence corresponding to each sample sequence in the sample sequence data.
[0190] Furthermore, the sample pair construction auxiliary model includes at least one of the following: a primary-level modal intra-aggregate sub-model, a secondary-level inter-modal alignment sub-model, and a higher-level sample intra-connection sub-model;
[0191] Accordingly, the sample construction and score prediction unit can be specifically used for:
[0192] Each of the single-modal representation sequences is input into the initial-level modal intra-aggregator model to obtain the first positive sample set and the first negative sample set output by the initial-level modal intra-aggregator model;
[0193] Each of the single-modal representation sequences is input into the intermediate-level inter-modal alignment sub-model to obtain the second positive sample set and the second negative sample set output by the intermediate-level inter-modal alignment sub-model;
[0194] Each of the single-modal representation sequences is input into the high-level sample intralinkage sub-model to obtain the third positive sample set output by the high-level sample intralinkage sub-model.
[0195] Furthermore, the sample construction and score prediction unit can specifically be used for:
[0196] The input single-modal representation sequences are concatenated through the cascaded layers of the multimodal fusion sub-model to obtain the initial representation sequence of the sample data.
[0197] The initial representation sequence is mapped according to a preset representation dimension through the first linear layer of the multimodal fusion sub-model, and the initial representation sequence with the preset representation dimension obtained by the mapping is processed by a nonlinear activation function.
[0198] The multimodal representation sequence is obtained by performing deep representation extraction on the processed initial representation sequence through the second linear layer of the multimodal fusion sub-model.
[0199] The multimodal representation sequence is input into the regression layer of the multimodal fusion sub-model to obtain the predicted sentiment score corresponding to the sample data output by the regression layer.
[0200] Furthermore, the positive and negative sample sets include a first positive sample set, a first negative sample set, a second positive sample set, a second negative sample set, and a third positive sample set;
[0201] Accordingly, the loss function determination unit can specifically be used for:
[0202] For each single-modal representation sequence, based on the single-modal representation sequence and the positive and negative samples corresponding to the single-modal representation sequence in the first positive sample set and the first negative sample set, respectively, the first function value of the first contrastive learning loss function is calculated;
[0203] Based on the single-modal representation sequence, and the positive and negative samples corresponding to the single-modal representation sequence in the second positive sample set and the second negative sample set, respectively, the second function value of the second contrastive learning loss function is calculated. The second contrastive learning loss function is the sum of the first contrastive learning loss function and the function compensation value. The function compensation value is used to compensate the first contrastive learning loss function.
[0204] Based on the single-modal representation sequence and the positive samples in the third positive sample set corresponding to the single-modal representation sequence, calculate the third function value of the second-norm loss function;
[0205] Based on the predicted sentiment score and the sentiment score label, calculate the fourth function value of the mean absolute error loss function;
[0206] The target loss function is obtained by weighted summation of the first, second, third, and fourth function values.
[0207] Furthermore, the first positive sample set includes at least one first positive sample pair, which consists of two single-modal representation values that belong to the same modality, have the same sentiment tendency, and have a sentiment score label difference value less than the first positive sample selection threshold. The single-modal representation values are determined based on the single-modal representation sequence.
[0208] The first negative sample set includes at least one first negative sample pair, which consists of two single-modal representation values belonging to the same modality, having different sentiment tendencies, and having a sentiment score label difference value greater than the first negative sample selection threshold; the second positive sample set includes at least one second positive sample pair, which consists of two single-modal representation values belonging to different modalities, having the same sentiment tendency, and having a sentiment score label difference value less than the second positive sample selection threshold; the second positive sample selection threshold is the difference between the first positive sample selection threshold and the first threshold compensation value, and the first threshold compensation value is used to compensate for the first positive sample selection threshold;
[0209] The second negative sample set includes at least one second negative sample pair. The second negative sample pair consists of two single-modal representation values that belong to different modalities, have different sentiment tendencies, and have sentiment score label differences greater than the second negative sample selection threshold. The second negative sample selection threshold is the sum of the first negative sample selection threshold and the second threshold compensation value. The second threshold compensation value is used to compensate for the first negative sample selection threshold.
[0210] The third positive sample set includes at least one third positive sample pair, which consists of two single-modal representation values belonging to the same sample data but belonging to different modalities.
[0211] Furthermore, the first positive sample selection threshold increases with the number of iterations, while the first negative sample selection threshold decreases with the number of iterations.
[0212] The vehicle cockpit control device provided in the embodiments of the present invention can execute the vehicle cockpit control method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.
[0213] Figure 4 A schematic diagram of the structure of a vehicle 40 that can be used to implement an embodiment of the present invention is shown. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the invention described and / or claimed herein.
[0214] like Figure 4 As shown, vehicle 40 includes at least one processor 41 and a memory, such as read-only memory (ROM) 42 and random access memory (RAM) 43, communicatively connected to at least one processor 41. The memory stores computer programs executable by at least one processor. Processor 41 can perform various appropriate actions and processes based on the computer program stored in ROM 42 or loaded from storage unit 48 into RAM 43. RAM 43 can also store various programs and data required for the operation of vehicle 40. Processor 41, ROM 42, and RAM 43 are interconnected via bus 44. Input / output (I / O) interface 45 is also connected to bus 44.
[0215] Multiple components in vehicle 40 are connected to I / O interface 45, including: input unit 46, such as keyboard, mouse, etc.; output unit 47, such as various types of displays, speakers, etc.; storage unit 48, such as disk, optical disk, etc.; and communication unit 49, such as network card, modem, wireless transceiver, etc. Communication unit 49 allows vehicle 40 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0216] Processor 41 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 41 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 41 performs the various methods and processes described above, such as vehicle cockpit control methods.
[0217] In some embodiments, the vehicle cockpit control method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 48. In some embodiments, part or all of the computer program may be loaded and / or installed on vehicle 40 via ROM 42 and / or communication unit 49. When the computer program is loaded into RAM 43 and executed by processor 41, one or more steps of the vehicle cockpit control method described above may be performed. Alternatively, in other embodiments, processor 41 may be configured to perform the vehicle cockpit control method by any other suitable means (e.g., by means of firmware).
[0218] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0219] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0220] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0221] To provide interaction with the user, the systems and technologies described herein can be implemented in a vehicle having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the vehicle. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0222] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0223] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0224] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0225] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A vehicle cockpit control method, characterized in that, include: Acquire multimodal emotional information of users in the cockpit; The user's emotional score is obtained by performing emotion recognition based on the multimodal emotional information using a trained emotion recognition model. The emotion recognition model is obtained by training the model through a sample training set and a set of positive and negative samples. The positive and negative sample pairs contained in the set of positive and negative samples are generated by constructing sample pairs hierarchically from the single-modal representation sequence of the sample data in the sample training set. The equipment in the cockpit is controlled according to the control method corresponding to the emotional score.
2. The method according to claim 1, characterized in that, The emotion recognition model is obtained by training the model using a sample training set and a set of positive and negative samples, including: A sample training set and a pre-built initial recognition model are obtained. The sample training set includes multiple sample data and sentiment score labels corresponding to each sample data. Each sample data includes multiple sample sequences of different modalities. The initial recognition model includes a single-modal representation extraction sub-model and a multi-modal fusion sub-model. The sample data is input into the initial recognition model to obtain the single-modal representation sequences output by the single-modal representation extraction sub-model. Each single-modal representation sequence corresponds one-to-one with each sample sequence in the sample data. Each of the single-modal representation sequences is input into the sample pair construction auxiliary model and the multimodal fusion sub-model to obtain the positive and negative sample sets corresponding to each of the single-modal representation sequences output by the sample pair construction auxiliary model, and the predicted sentiment score corresponding to the sample data output by the multimodal fusion sub-model. Based on the set of positive and negative samples, the predicted sentiment score, and the sentiment score label, and in conjunction with the set loss function, the target loss function value is determined; The network parameters in the initial recognition model are back-learned and adjusted based on the target loss function value to obtain the adjusted initial recognition model. The input operation of sample data is then re-executed until the training termination condition is met. The initial recognition model obtained after training is determined as the emotion recognition model.
3. The method according to claim 2, characterized in that, The sample data includes text sample sequences, speech sample sequences, and visual sample sequences; Accordingly, the step of inputting the sample data into the initial recognition model to obtain the single-modal representation sequences output by the single-modal representation extraction sub-model includes: The text sample sequence is input into the first encoder included in the single-modal representation extraction sub-model, and the text sample sequence is encoded by the first encoder to obtain the first representation vector; The speech sample sequence and the visual sample sequence are respectively input into the second encoder included in the single-modal representation extraction sub-model. The second encoder encodes the speech sample sequence and the visual sample sequence respectively to obtain the second representation vector and the third representation vector. Each representation vector obtained through encoding is mapped according to a preset representation dimension, and the representation vectors of each preset representation dimension obtained by mapping are averaged and pooled to obtain each single-modal representation sequence corresponding to each sample sequence in the sample sequence data.
4. The method according to claim 2, characterized in that, The sample pair construction auxiliary model includes at least one of the following: a primary-level modal intra-modal aggregation sub-model, a mid-level inter-modal alignment sub-model, and a high-level sample intra-connection sub-model; Accordingly, each of the single-modal representation sequences is input into the sample pair construction auxiliary model to obtain the positive and negative sample sets corresponding to each of the single-modal representation sequences output by the sample pair construction auxiliary model, including: Each of the single-modal representation sequences is input into the initial-level modal intra-aggregator model to obtain the first positive sample set and the first negative sample set output by the initial-level modal intra-aggregator model; Each of the single-modal representation sequences is input into the intermediate-level inter-modal alignment sub-model to obtain the second positive sample set and the second negative sample set output by the intermediate-level inter-modal alignment sub-model; Each of the single-modal representation sequences is input into the high-level sample intralinkage sub-model to obtain the third positive sample set output by the high-level sample intralinkage sub-model.
5. The method according to claim 2, characterized in that, Each of the single-modal representation sequences is input into the multimodal fusion sub-model to obtain the predicted sentiment score corresponding to the sample data output by the multimodal fusion sub-model, including: The input single-modal representation sequences are concatenated through the cascaded layers of the multimodal fusion sub-model to obtain the initial representation sequence of the sample data. The initial representation sequence is mapped according to a preset representation dimension through the first linear layer of the multimodal fusion sub-model, and the initial representation sequence with the preset representation dimension obtained by the mapping is processed by a nonlinear activation function. The multimodal representation sequence is obtained by performing deep representation extraction on the processed initial representation sequence through the second linear layer of the multimodal fusion sub-model. The multimodal representation sequence is input into the regression layer of the multimodal fusion sub-model to obtain the predicted sentiment score corresponding to the sample data output by the regression layer.
6. The method according to claim 2, characterized in that, The positive and negative sample sets include a first positive sample set, a first negative sample set, a second positive sample set, a second negative sample set, and a third positive sample set; Accordingly, determining the target loss function value based on the positive and negative sample set, the predicted sentiment score, and the sentiment score label, combined with a set loss function, includes: For each single-modal representation sequence, based on the single-modal representation sequence and the positive and negative samples corresponding to the single-modal representation sequence in the first positive sample set and the first negative sample set, respectively, the first function value of the first contrastive learning loss function is calculated; Based on the single-modal representation sequence, and the positive and negative samples corresponding to the single-modal representation sequence in the second positive sample set and the second negative sample set, respectively, the second function value of the second contrastive learning loss function is calculated. The second contrastive learning loss function is the sum of the first contrastive learning loss function and the function compensation value. The function compensation value is used to compensate the first contrastive learning loss function. Based on the single-modal representation sequence and the positive samples in the third positive sample set corresponding to the single-modal representation sequence, calculate the third function value of the second-norm loss function; Based on the predicted sentiment score and the sentiment score label, calculate the fourth function value of the mean absolute error loss function; The target loss function is obtained by weighted summation of the first, second, third, and fourth function values.
7. The method according to claim 4 or 6, characterized in that, The first positive sample set includes at least one first positive sample pair. The first positive sample pair consists of two single-modal representation values that belong to the same modality, have the same sentiment tendency, and have a sentiment score label difference value that is less than the first positive sample selection threshold. The single-modal representation values are determined based on the single-modal representation sequence. The first negative sample set includes at least one first negative sample pair, which consists of two single-modal representation values belonging to the same modality, having different sentiment tendencies, and having a sentiment score label difference value greater than the first negative sample selection threshold; the second positive sample set includes at least one second positive sample pair, which consists of two single-modal representation values belonging to different modalities, having the same sentiment tendency, and having a sentiment score label difference value less than the second positive sample selection threshold; the second positive sample selection threshold is the difference between the first positive sample selection threshold and the first threshold compensation value, and the first threshold compensation value is used to compensate for the first positive sample selection threshold; The second negative sample set includes at least one second negative sample pair. The second negative sample pair consists of two single-modal representation values that belong to different modalities, have different sentiment tendencies, and have sentiment score label differences greater than the second negative sample selection threshold. The second negative sample selection threshold is the sum of the first negative sample selection threshold and the second threshold compensation value. The second threshold compensation value is used to compensate for the first negative sample selection threshold. The third positive sample set includes at least one third positive sample pair, which consists of two single-modal representation values belonging to the same sample data but belonging to different modalities.
8. The method according to claim 7, characterized in that, The first positive sample selection threshold increases with the number of iterations, and the first negative sample selection threshold decreases with the number of iterations.
9. A vehicle cockpit control device, characterized in that, include: The information acquisition module is used to acquire multimodal emotional information of users in the cockpit; The scoring module is used to perform emotion recognition based on the multimodal emotion information using a trained emotion recognition model to obtain the user's emotion score; The emotion recognition model is obtained by training the model through a sample training set and a set of positive and negative samples. The positive and negative sample pairs contained in the set of positive and negative samples are generated by constructing sample pairs hierarchically from the single-modal representation sequence of the sample data in the sample training set. The control module is used to control the equipment in the cockpit according to the control method corresponding to the emotion score.
10. A vehicle, characterized in that, The vehicles include: At least one processor; and a memory communicatively connected to the at least one processor; The memory stores a computer program that can be executed by the at least one processor, which is then executed by the at least one processor to enable the at least one processor to perform the vehicle cockpit control method according to any one of claims 1-8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the vehicle cockpit control method of any one of claims 1-8.
12. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the vehicle cockpit control method according to any one of claims 1-8.
Citation Information
Patent Citations
Multi-modal sentiment analysis method, system and equipment based on multivariate loss function and medium
CN116701996A
Multi-modal sentiment analysis method based on personality and generality comparison staged guidance
CN117809229A
Multi-granularity cross-modal comparative learning method and device for multi-modal sentiment analysis
CN119066543A