Semantic emotion analysis method, intelligent cabin, medium and vehicle
By segmenting and semantically encoding user speech, using gating networks to filter target word semantic codes from independent neural networks, and generating shared semantic features, the problem of disconnect between local and overall emotional bias in intelligent cockpits is solved, achieving more accurate sentiment analysis.
Patent Information
- Application Number
- CN202511294829.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-10
- Publication Date
- 2025-12-12
AI Technical Summary
Existing sentiment analysis methods sever the connection between local and overall user sentiment bias, thus affecting the sentiment analysis capabilities of smart cockpits.
By segmenting and semantically encoding user speech, a gating network is used to filter the target word semantic encoding of independent neural networks, and a new target neural network is added to generate shared semantic features. The sentiment category is then determined by combining global semantic information.
The improved emotion analysis capabilities of the intelligent cockpit make emotion analysis more comprehensive and accurate, taking into account both detailed expressions and overall emotional bias, and reducing misjudgments.
Smart Images

Figure CN121122331A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of intelligent cockpit technology, and in particular to a semantic sentiment analysis method, intelligent cockpit, medium, and vehicle. Background Technology
[0002] As smart cockpits are upgraded, traditional physical interaction methods (such as buttons and touch screens) are gradually failing to meet the smart cockpit's pursuit of natural interaction, and emotional computing is needed to achieve a user-centric interaction upgrade.
[0003] The emotional fluctuations of in-vehicle users (such as drivers) can significantly influence driving decisions. For example, when a driver experiences negative emotions such as anger or tension, recognizing the driver's emotional category will trigger safety interventions, such as proactively lowering the media volume or playing relaxing music to ensure driving safety. To conduct nuanced sentiment analysis of drivers, a corresponding label is typically created for each emotion category in the sentiment analysis model, and words are captured from the user's voice information to parse the semantic representation under each label.
[0004] Generally, a user's overall emotional bias influences the classification of emotions. For example, a driver's overall voice output leaning towards positive emotions will positively impact sentiment analysis. However, current sentiment analysis models rely solely on capturing relevant local words for the semantic representation of each label, ignoring the user's overall emotional bias. This approach severs the connection between the local and the overall context, resulting in semantic representations that deviate from the overall context and leading to poor sentiment analysis capabilities in smart cockpits. Summary of the Invention
[0005] This specification provides a semantic sentiment analysis method, a smart cockpit, a medium, and a vehicle to solve or partially solve the technical problem that existing sentiment analysis methods sever the connection between local and overall user sentiment bias, thus affecting the sentiment analysis capabilities of smart cockpits.
[0006] To address the aforementioned technical problems, this specification discloses a semantic sentiment analysis method applied to a smart cockpit. The method includes:
[0007] The user's speech is segmented into words to obtain a token sequence containing several word elements, and the global semantic information corresponding to the user's speech is placed into a set position in the token sequence;
[0008] The token sequence is semantically encoded to obtain the global semantic code corresponding to the global semantic information, and the semantic code of the plurality of words;
[0009] The semantic encoding of the several word units is processed by the gating network in the semantic capture model, and the semantic encoding of the target word units corresponding to several independent neural networks is selected; wherein, the several independent neural networks are arranged side by side in the semantic capture model, and each independent neural network corresponds to a sentiment category;
[0010] The semantic encoding of each target word is processed by the aforementioned independent neural networks to obtain the sentiment semantic features corresponding to each of the aforementioned independent neural networks;
[0011] A target neural network is added to the semantic capture model. The gating network is used to output the global semantic encoding and the semantic encoding of all words to the target neural network for processing, and output shared semantic features. The target neural network and the several independent neural networks are parallel. The shared semantic features are semantic features used by the several independent neural networks.
[0012] In each independent neural network, the shared semantic features, the global semantic encoding corresponding to the global semantic information, and the sentiment semantic features corresponding to the independent neural network are fused to obtain the fused features corresponding to the independent neural network.
[0013] Based on the fusion features corresponding to each independent neural network, the emotion probability corresponding to each emotion category is calculated to determine the emotion category to which the user's voice belongs.
[0014] This specification discloses an intelligent cockpit, including:
[0015] The lexical module is used to segment user speech into words, obtain a token sequence containing several lexical units, and place the global semantic information corresponding to the user speech into a set position in the token sequence;
[0016] The encoding module is used to perform semantic encoding on the token sequence to obtain the global semantic encoding corresponding to the global semantic information, as well as the semantic encoding of the plurality of word elements;
[0017] The first filtering module is used to process the semantic encoding of the plurality of lexical units using the gating network in the semantic capture model, and to filter out the target lexical unit semantic encoding corresponding to the plurality of independent neural networks; wherein, the plurality of independent neural networks are arranged side by side in the semantic capture model, and each independent neural network corresponds to a sentiment category;
[0018] The network processing module is used to process the semantic encoding of the target word units corresponding to the plurality of independent neural networks to obtain the sentiment semantic features corresponding to the plurality of independent neural networks.
[0019] The second filtering module is used to add a target neural network to the semantic capture model, and use the gating network to output the global semantic encoding and the semantic encoding of all words to the target neural network for processing, and output shared semantic features; wherein, the target neural network and the several independent neural networks are parallel; the shared semantic features are semantic features used by the several independent neural networks.
[0020] The fusion module is used to fuse the shared semantic features, the global semantic encoding corresponding to the global semantic information, and the sentiment semantic features corresponding to the independent neural network in each independent neural network to obtain the fused features corresponding to the independent neural network.
[0021] The calculation module is used to calculate the emotion probability corresponding to each emotion category based on the fusion features corresponding to each independent neural network, so as to determine the emotion category to which the user's voice belongs.
[0022] This specification discloses a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps of the above-described method.
[0023] This specification discloses a vehicle including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the method described above.
[0024] Through one or more embodiments of this specification, this specification has the following beneficial effects or advantages:
[0025] The technical solution described in this specification splits user speech and embeds global semantic information into a token sequence, using both local lexical units and the overall user semantics as the basis for analysis. After obtaining the global semantic encoding and the semantic encoding of each lexical unit, a gating network is used to filter the target lexical unit semantic encoding of each independent neural network, processing it to obtain the emotional semantic features adapted to each independent neural network. Then, a new target neural network is added, and combined with the gating network, shared semantic features are generated for use by all independent neural networks. Thus, each independent neural network integrates the shared semantic features, the global semantic encoding corresponding to the global semantic information, and its own emotional semantic features. This allows each independent neural network to analyze specific emotion categories by having both focused local emotional semantic features as detailed basis for accurate judgment, grasping the overall emotional bias through global semantic encoding, and obtaining semantic support common to all emotion categories through shared semantic features. This allows for in-depth mining of the local and overall correlation, ensuring that the judgment of each emotion category takes into account both detailed expression and global emotional bias, reducing misjudgments caused by partial information extraction, and making emotion analysis more comprehensive and accurate, thereby improving the emotion analysis capabilities of the intelligent cockpit.
[0026] The above description is merely an overview of the technical solution in this specification. In order to better understand the technical means in this specification and to implement it in accordance with the contents of this specification, and to make the above and other objects, features and advantages of this specification more apparent and understandable, specific embodiments of this specification are given below. Attached Figure Description
[0027] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit this specification. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:
[0028] Figure 1 A flowchart of a semantic sentiment analysis method according to an embodiment of this specification is shown;
[0029] Figure 2 The overall implementation logic of a semantic sentiment analysis method according to one embodiment of this specification is shown;
[0030] Figure 3 A schematic diagram of an intelligent cockpit according to one embodiment of this specification is shown. Detailed Implementation
[0031] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0032] Firstly, the embodiments of this specification provide a semantic sentiment analysis method applied to a smart cockpit.
[0033] See Figure 1 The semantic sentiment analysis method in this embodiment includes the following steps:
[0034] S101, the user's speech is segmented to obtain a token sequence containing several word elements, and the global semantic information corresponding to the user's speech is placed into the set position of the token sequence.
[0035] Specifically, user speech is parsed to obtain text content and its global semantic information. User speech refers to the speech produced by the driver in a driving scenario, but it can also include speech produced by passengers. Speech recognition technology is used to parse user speech to obtain text content; then, the overall intent, core needs, and emotional bias of the user's speech or text content are extracted to obtain global semantic information. Global semantic information is used to represent the overall emotional bias of the user's speech and is not limited to individual words or phrases, but is extracted through a comprehensive understanding of the entire speech segment.
[0036] The text content is segmented into tokens, resulting in a sequence of tokens containing several words, also known as a word sequence. A token is the smallest unit of word segmentation formed after dividing the text content according to semantic and syntactic rules. In practical applications, segmentation rules such as WordPiece, BPE (Byte Pair Encoding), and BBPE (Byte-level Byte Pair Encoding) can be referenced, but are not strictly limited. Furthermore, global semantic information is placed in designated positions, such as the beginning of the token sequence. A terminator is set at the end of the token sequence. The individual words obtained from the segmentation are placed in positions other than the beginning and end of the token sequence. For example, the token sequence might be: [CLS], w1, w2, w3, ..., w n [SEP]. Where [CLS] is the head position of the token sequence, containing global semantic information, and [SEP] is the tail position, containing the end marker, w1, w2, w3, ..., w n Each token is a single token. Other token sequences can be appended after the terminator [SEP].
[0037] S102, perform semantic encoding on the token sequence to obtain the global semantic encoding corresponding to the global semantic information, as well as the semantic encoding of several word elements.
[0038] The sequence code H is obtained by semantically encoding the global semantic information, each word, and the terminator. The sequence code H includes: the global semantic code corresponding to the global semantic information; the semantic codes for several words, where one word corresponds to one semantic code; and the terminator's terminator code. For example, the structure relative to the token sequence is: [CLS], w1, w2, w3, ..., w n After encoding [SEP], the corresponding sequence code H is obtained as: h cls h1, h2, h3, ..., h n h sep Among them, h is clsThe global semantic information corresponds to the global semantic encoding, h1, h2, h3, ..., h n It is the semantic encoding corresponding to each word, h sep It is the end code corresponding to the end character.
[0039] In the semantic encoding process, the global semantic information, each word, and the terminator are first converted into corresponding initial vectors, and then processed using a multi-layer self-attention mechanism to output the sequence code H.
[0040] When processing individual word vectors, a multi-layer self-attention mechanism is used to combine other word vectors in the token sequence and the global semantic information vector. Different attention weights are assigned based on the semantic relevance between the individual word vector and the token sequence, dynamically capturing the semantic association between the individual word vector and the token sequence to output the semantic code corresponding to the individual word. This design preserves the specific meaning of local words while integrating global contextual information, providing deep semantic representations for subsequent tasks such as sentiment analysis. Of course, the global semantic vector is also encoded according to the above process to obtain the global semantic code h corresponding to the global semantic information. cls The terminator will also undergo the above operation to obtain the corresponding terminator code h. sep .
[0041] S103 utilizes the gating network in the semantic capture model to process the semantic encoding of several lexical units, and filters out the target lexical semantic encoding corresponding to several independent neural networks.
[0042] In the semantic capture model, see Figure 2 The model comprises a gating network, several independent neural networks, and a target neural network. The independent neural networks are arranged side-by-side within the semantic capture model, each corresponding to a sentiment category. The number of independent neural networks is determined by the number of sentiment categories. The target neural network is arranged alongside the independent neural networks but does not correspond to any sentiment category. The gating network connects each of the independent neural networks and the target neural network.
[0043] To obtain semantic representations for each sentiment category, it's necessary to extract semantic representations related to the current sentiment category from the token sequence. Current technologies typically employ a combination of attention mechanisms and pooling operations to extract semantic representations for each sentiment category. However, the semantic representations of various lexical units in the token sequence differ; some lexical units are applicable to their own sentiment category but are considered interfering units in other sentiment categories. The aforementioned mechanism, when extracting semantic representations for each sentiment category, only reduces the weight of interfering lexical units, rather than discarding them directly. If the interfering lexical units themselves have strong semantic meaning, they will affect the final semantic representation of the sentiment category.
[0044] To address this issue, this specification employs a gating network to calculate the semantic probability of each word's semantic code mapping to the independent neural network for each independent neural network. Referring to the semantic probabilities of all words in the independent neural network, the top b word semantic codes, ranked by probability, are selected as the target word semantic codes in the independent neural network. This approach avoids considering all words, instead excluding those with low probabilities or that are irrelevant, thus obtaining semantic codes closely related to the independent neural network.
[0045] The gating network is shown below:
[0046] s j =topb(softmax(ffn) 1j (h i )))
[0047] Among them, s j h represents the semantic encoding of the target lexical unit selected by the j-th independent neural network. i h represents the semantic encoding of the i-th term. i ∈[h1, h2, ..., h n ], where n represents the number of semantic codes for lexical units, ffn 1j () represents the feedforward neural network in the gated network. The symbol j indicates that the computation is currently being performed on the j-th independent neural network. When computing other independent neural networks, the symbol will be changed accordingly. Softmax() represents the normalization function, softmax(ffn) 1j (h i )) represents calculating the semantic probability of the semantic code of the i-th word relative to the j-th independent neural network, and topb represents the probability from [h1, h2, ..., hj] to [hj]. n The first b semantic codes are selected according to their probability as the target semantic codes of the j-th independent neural network.
[0048] When selecting the target word semantic encoding for the j-th independent neural network, the feedforward neural network ffn1() in the gated network is used to calculate [h1, h2, ..., hj]. n The semantic code of each word in the array is relative to the semantic probability of the j-th independent neural network, resulting in n semantic probabilities. The first b word semantic codes are selected as the target word semantic codes for the j-th independent neural network according to their probability magnitude. All word semantic codes after the first b are excluded to prevent interference, ensuring that each independent neural network is closely related to its selected target word semantic codes. The processing method for other independent neural networks is similar, so it will not be elaborated further.
[0049] S104, using several independent neural networks to process the semantic encoding of their respective target lexical units, to obtain the sentiment semantic features corresponding to each of the several independent neural networks.
[0050] Within the independent neural network, the semantic encoding of the target lexical unit is processed to obtain the sentiment semantic features used to calculate the probability of sentiment category.
[0051] Each individual neural network is shown below:
[0052] e j =ffn downj (Gelu(ffn gate (s j ))×ffn upj (s j ))
[0053] Among them, e j Let s represent the sentiment semantic features of the j-th independent neural network. j ffn represents the semantic encoding of the target lexical unit selected by the j-th independent neural network. gate (s j ) indicates that s is calculated using a feedforward neural network with a gated network. j The gating weights involved in the j-th independent neural network required for each semantic encoding, ffn upj (s j ) indicates that the j-th independent neural network itself is used to process s. j The semantic output obtained after each semantic encoding, Gelu() represents the activation function, ffn downj () is used to map to a uniform dimension.
[0054] Inside the j-th independent neural network, a feedforward neural network ffn is passed through a gated network. gate () Calculate the target word semantic code s j The gating weights required for each semantic encoding in the j-th independent neural network are passed through the feedforward neural network ffn of the j-th independent neural network itself. upj () Processing target lexical semantic encodings j Each semantic code in the algorithm yields a corresponding semantic output. The gating weights for each semantic code are different, depending on the target word semantic code s. j Each semantic code in the algorithm is activated using the Gelu() activation function after performing a weighted sum of gating weights and semantic outputs, and then activated using the feedforward neural network ffn of the j-th independent neural network. downjBy mapping the () to a unified dimension, the sentiment semantic features can be extracted. After processing all the semantic codes in the target word semantic encoding in this way, the sentiment semantic features corresponding to the j-th independent neural network are obtained.
[0055] It is worth noting that the gated network's feedforward neural network ffn gate () and the feedforward neural network in the gated network ffn i () are two independent feedforward neural networks that are located at different positions in the gated network and process different parameters.
[0056] S105 adds a target neural network to the semantic capture model. It uses a gating network to output the global semantic code and the semantic codes of all words to the target neural network for processing, and outputs shared semantic features.
[0057] Shared semantic features are semantic features used by several independent neural networks.
[0058] To capture the shared semantic features corresponding to the global semantic encoding, a target neural network is added to the semantic capture model. In this case, the gating network does not perform filtering and outputs the global semantic encoding and the semantic encoding of all words together to the target neural network for processing.
[0059] The target neural network is shown below:
[0060] e s =ffn down ′(Gelu(ffn gate (H))×ffn up ′(H))
[0061] Among them, e s H represents shared semantic features, and H represents the encoded sequence, which contains at least the global semantic encoding h. cls and the set of semantic codes for all lexical units: h1, h2, ..., h n Of course, it can also include the semantic code h corresponding to the terminator. sep ;ffn gate (H) represents the weight assignment of each semantic code in H calculated using a feedforward neural network of a gated network, ffn up ′(H) represents the semantic output obtained after processing H using the target neural network's own feedforward neural network, Gelu() represents the activation function, and ffn down ′() is used to map to a uniform dimension.
[0062] Inside the target neural network, a feedforward neural network ffn with a gated network is used. gate () Calculate the weight assignment of each semantic code in the encoding sequence H, and then use the target neural network's own feedforward neural network ffn.up The encoding sequence H is processed by '() to obtain the semantic output corresponding to each semantic code in the encoding sequence H. The gating weights for each semantic code are different. For each semantic code, after performing a weighted sum of the gating weights and the semantic output, the activation function Gelu() is used for activation, and the target neural network's own feedforward neural network ffn is utilized. down By mapping ′() to a unified dimension, shared semantic features can be obtained.
[0063] It is worth noting that the gated network's feedforward neural network ffn gate () Calculates all semantic codes in the encoding sequence H. The purpose is to reset the weight of semantic codes belonging to common features to 1 and the weight of semantic codes not belonging to common features to 0, so as to obtain shared semantic features that can be adapted to all independent neural networks.
[0064] S106, in each independent neural network, the shared semantic features, the global semantic encoding corresponding to the global semantic information, and the sentiment semantic features corresponding to the independent neural network are fused to obtain the fused features corresponding to the independent neural network.
[0065] Among them, the dimensions of the shared semantic features, the global semantic information corresponding to the global semantic encoding, and the emotional semantic features corresponding to the independent neural network are the same.
[0066] In the fusion process, one possible fusion method is to add or multiply the shared semantic features, the global semantic encoding corresponding to the global semantic information, and the sentiment semantic features corresponding to the independent neural networks to obtain the fused features. Another possible fusion method is to concatenate the shared semantic features, the global semantic encoding corresponding to the global semantic information, and the sentiment semantic features corresponding to the independent neural networks, then integrate them through an attention mechanism or a feedforward neural network, autonomously learning the weight allocation of the three and performing a weighted sum to obtain the fused features. Yet another possible fusion method is to add the global semantic encoding corresponding to the global semantic information as a residual term to the fusion result of the shared features and sentiment features to obtain the fused features.
[0067] The fusion features of the j-th neural network are shown below:
[0068] z j =ffn3([e j :e s :h cls ])
[0069] Among them, z j Let e represent the fusion feature of the j-th neural network. j Let e represent the sentiment semantic features of the j-th independent neural network. s h represents shared semantic features clsffn3() represents global semantic encoding and feedforward neural network.
[0070] In the fusion scheme of this embodiment, considering that the user's overall emotional tendency will affect the classification of emotion categories, the global semantic encoding corresponding to the global semantic information is taken into account in order to integrate the user's overall emotional tendency. Specifically, the emotional semantic features corresponding to independent neural networks may focus more on the emotional bias in the text, such as emotional signals such as "anxious" and "irritable". However, such features may be biased due to the diversity of the same emotional expression. The integration of global semantic encoding can supplement the overall emotional bias and avoid misidentification of emotion categories due to the one-sidedness of emotional features.
[0071] Furthermore, taking shared semantic features into account can provide general and stable basic semantic support, not limited to specific emotional scenarios or tasks. After fusion, it can provide a universal semantic anchor for the model, reducing the generalization deficiency caused by the over-reliance of emotional features on local details, and enabling the intelligent cockpit to maintain stable semantic understanding capabilities in different voice scenarios.
[0072] In this embodiment, the integration of shared semantic features, global semantic encoding, and emotional semantic features corresponding to each independent neural network can form multi-level semantic complementarity, achieving deep association between the global and local. Through the basic semantic support of shared semantic features, the overall emotional bias guidance of global semantic encoding, and the detailed basis of emotional semantic features, the combination of the three can enhance the robustness of emotional category analysis. Even when faced with speech with unclear emotional expression, it can accurately identify its implicit emotional category, thereby improving the emotional analysis capability of the intelligent cockpit.
[0073] S107, based on the fusion features corresponding to each independent neural network, calculate the emotion probability corresponding to each emotion category to determine the emotion category to which the user's voice belongs.
[0074] Specifically, the probability of each emotion category is calculated based on a probability calculation formula. Each independent neural network is configured with this probability calculation formula.
[0075] The probability calculation formula is as follows:
[0076] y j =softmax(ffn4(z j ))
[0077] Among them, y j Let z represent the sentiment probability of the sentiment category corresponding to the j-th independent neural network. j represents the fusion feature of the j-th independent neural network, ffn4() represents the feedforward neural network, and softmax() represents the normalization function.
[0078] Furthermore, by referring to the emotional probability corresponding to each emotional category, the emotional category with the highest probability is selected as the emotional category to which the user's voice belongs.
[0079] To facilitate understanding of this technical solution, please refer to the following: Figure 2 This is the overall implementation logic of this technical solution.
[0080] First, the user's voice output during the vehicle driving scenario is parsed and segmented to obtain a token sequence containing several word units: [CLS], w1, w2, w3, ..., w n [SEP]. [CLS] is the head position of the token sequence, where global semantic information is placed, and [SEP] is the tail position, where the end character is placed, w1, w2, w3, ..., w n Each token is a single token. Other token sequences can be appended after the terminator [SEP].
[0081] The token sequence is encoded using the Roberta model to obtain the corresponding encoded sequence H:h cls h1, h2, h3, ..., h n h sep Among them, h is cls The global semantic information corresponds to the global semantic encoding, h1, h2, h3, ..., h n It is the semantic encoding corresponding to each word, h sep It is the end code corresponding to the end character.
[0082] The semantic capture model includes a gating network, several independent neural networks, and a target neural network. There are a total of k independent neural networks, the same number as the k emotion categories. Examples of independent neural networks are: independent neural network 1, ..., independent neural network k.
[0083] For each independent neural network, the gated network computes the semantic encoding sequence of the lexical units [h1, h2, ..., h...]. n For each lexical semantic code in the [ ], the semantic probability relative to the independent neural network is used to select the top b lexical semantic codes as the target lexical semantic codes in that independent neural network. In this way, the target lexical semantic codes corresponding to independent neural network 1, ..., independent neural network k are obtained.
[0084] In the target neural network, the encoded sequence H is processed to filter out shared semantic features e. s .
[0085] In the aggregation layer, semantic features e are shared for each individual neural network. s Global semantic encoding hsep The sentiment semantic feature e corresponding to this independent neural network j By fusing the data, the fused feature z corresponding to the independent neural network is obtained. j .
[0086] For example, for an independent neural network 1, the shared semantic feature e s Global semantic encoding h sep The sentiment semantic feature e1 corresponding to independent neural network 1 is fused to obtain the fused feature z1 corresponding to independent neural network 1. For example, for independent neural network k, the shared semantic feature e1 is fused to obtain the fused feature z1 corresponding to independent neural network 1. s Global semantic encoding h sep The sentiment semantic feature e corresponding to the independent neural network k k By fusing the data, we obtain the fused feature z corresponding to the independent neural networks k. k .
[0087] Each independent neural network is called with its corresponding feedforward neural network ffn and normalization function softmax to calculate the corresponding sentiment category probability. For example, independent neural network 1 outputs the sentiment probability for sentiment category 1, and independent neural network k outputs the sentiment probability for sentiment category k.
[0088] Select the emotion category with the highest probability from k emotion probabilities as the emotion category to which the user's voice belongs.
[0089] The above is the complete implementation logic of the semantic sentiment analysis method in this manual.
[0090] Secondly, based on the same inventive concept as the semantic sentiment analysis method in the first aspect, embodiments of this specification provide an intelligent cockpit, see [link to documentation]. Figure 3 ,include:
[0091] The word segmentation module 301 is used to segment the user's speech into words, obtain a token sequence containing several word segments, and place the global semantic information corresponding to the user's speech into a set position in the token sequence.
[0092] Encoding module 302 is used to perform semantic encoding on the token sequence to obtain the global semantic encoding corresponding to the global semantic information, and the semantic encoding of the plurality of word elements;
[0093] The first filtering module 303 is used to process the semantic codes of the plurality of lexical units using the gating network in the semantic capture model, and to filter out the target lexical unit semantic codes corresponding to the plurality of independent neural networks; wherein, the plurality of independent neural networks are arranged side by side in the semantic capture model, and each independent neural network corresponds to a sentiment category;
[0094] Network processing module 304 is used to process the semantic encoding of the target word units corresponding to the plurality of independent neural networks to obtain the sentiment semantic features corresponding to the plurality of independent neural networks.
[0095] The second filtering module 305 is used to add a target neural network to the semantic capture model, and use the gating network to output the global semantic encoding and the semantic encoding of all words to the target neural network for processing, and output shared semantic features; wherein, the target neural network and the several independent neural networks are parallel; the shared semantic features are semantic features used by the several independent neural networks.
[0096] The fusion module 306 is used to fuse the shared semantic features, the global semantic encoding corresponding to the global semantic information, and the sentiment semantic features corresponding to the independent neural network in each independent neural network to obtain the fused features corresponding to the independent neural network.
[0097] The calculation module 307 is used to calculate the emotion probability corresponding to each emotion category based on the fusion features corresponding to each independent neural network, so as to determine the emotion category to which the user's voice belongs.
[0098] Thirdly, based on the same inventive concept as the semantic sentiment analysis method in the first aspect, embodiments of this specification also provide a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps of any of the methods described above.
[0099] Fourthly, based on the same inventive concept as the semantic sentiment analysis method in the first aspect, embodiments of this specification also provide a vehicle including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of any of the methods described above.
[0100] The algorithms and displays provided herein are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems can also be used in conjunction with the teachings herein. The required structure for constructing such systems is apparent from the above description. Furthermore, this specification is not directed to any particular programming language. It should be understood that the contents of this specification can be implemented using various programming languages, and the above descriptions of specific languages are for the purpose of disclosing preferred embodiments of this specification.
[0101] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of this specification may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0102] Similarly, it should be understood that, in order to streamline this disclosure and aid in understanding one or more of the various inventive aspects, in the foregoing description of exemplary embodiments of this specification, various features of this specification are sometimes grouped together in a single embodiment, figure, or description thereof. However, this approach to disclosure should not be construed as reflecting an intention that the claimed specification requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, inventive aspects lie in fewer than all features of a single foregoing disclosed embodiment. Therefore, the claims following the detailed description are hereby expressly incorporated into that detailed description, wherein each claim itself is a separate embodiment of this specification.
[0103] Those skilled in the art will understand that modules in the device of the embodiments can be adaptively changed and placed in one or more devices different from that embodiment. Modules, units, or components in the embodiments can be combined into a single module, unit, or component, and further, they can be divided into multiple sub-modules, sub-units, or sub-components. Except where at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all processes or units of any method or device so disclosed. Unless expressly stated otherwise, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) may be replaced by an alternative feature that serves the same, equivalent, or similar purpose.
[0104] Furthermore, those skilled in the art will understand that although some embodiments herein include certain features included in other embodiments but not others, combinations of features from different embodiments are intended to be within the scope of this specification and form different embodiments. For example, in the following claims, any of the claimed embodiments can be used in any combination.
[0105] The various component embodiments of this specification can be implemented in hardware, or as software modules running on one or more processors, or a combination thereof. Those skilled in the art will understand that microprocessors or digital signal processors (DSPs) can be used in practice to implement some or all of the functions of some or all of the components of the gateway, proxy server, or system according to embodiments of this specification. This specification can also be implemented as a device or apparatus program (e.g., a computer program and computer program product) for performing some or all of the methods described herein. Such implementations of this specification can be stored on a computer-readable medium or can take the form of one or more signals. Such signals can be downloaded from an Internet website, provided on a carrier signal, or provided in any other form.
[0106] It should be noted that the above embodiments are illustrative of this specification and not limiting of it, and that those skilled in the art can devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. This specification can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names.
Claims
1. A semantic sentiment analysis method, the method being applied to a smart cockpit, the method comprising: The user's speech is segmented into words to obtain a token sequence containing several word elements, and the global semantic information corresponding to the user's speech is placed into a set position in the token sequence; The token sequence is semantically encoded to obtain the global semantic code corresponding to the global semantic information, and the semantic code of the plurality of words; The semantic encoding of the several word units is processed by the gating network in the semantic capture model, and the semantic encoding of the target word units corresponding to several independent neural networks is selected; wherein, the several independent neural networks are arranged side by side in the semantic capture model, and each independent neural network corresponds to a sentiment category; The semantic encoding of each target word is processed by the aforementioned independent neural networks to obtain the sentiment semantic features corresponding to each of the aforementioned independent neural networks; A target neural network is added to the semantic capture model. The gating network is used to output the global semantic encoding and the semantic encoding of all words to the target neural network for processing, and output shared semantic features. The target neural network and the several independent neural networks are parallel. The shared semantic features are semantic features used by the several independent neural networks. In each independent neural network, the shared semantic features, the global semantic encoding corresponding to the global semantic information, and the sentiment semantic features corresponding to the independent neural network are fused to obtain the fused features corresponding to the independent neural network. Based on the fusion features corresponding to each independent neural network, the emotion probability corresponding to each emotion category is calculated to determine the emotion category to which the user's voice belongs.
2. The method as described in claim 1, wherein segmenting the user's speech to obtain a token sequence containing several word units, and placing the global semantic information corresponding to the user's speech into a predetermined position in the token sequence, specifically includes: The user's speech is parsed to obtain the text content and its global semantic information; The text content is segmented to obtain the token sequence, and the global semantic information is placed in the set position.
3. The method as described in claim 1, wherein processing the semantic encoding of the plurality of lexical units using a gating network in the semantic capture model, and filtering out the target lexical unit semantic encoding corresponding to each of the plurality of independent neural networks, specifically includes: For each independent neural network, the semantic probability of mapping each word's semantic code to the independent neural network is calculated using the gating network; Referring to the semantic probabilities of all lexical units in the independent neural network, the b lexical semantic codes ranked first according to their probability are selected as the target lexical semantic codes in the independent neural network.
4. The method as described in claim 3, wherein the gating network is as follows: s j =topb(softax(ffn 1j (h i )))) in, s j h represents the semantic encoding of the target lexical unit selected by the j-th independent neural network. i h represents the semantic encoding of the i-th term. i ∈[h1, h2, ..., h n ], where n represents the number of semantic codes for lexical units, ffn 1j () represents a feedforward neural network, and softmax() represents the normalization function, softmax(ffn) 1j (h i )) represents calculating the semantic probability of the semantic code of the i-th word relative to the j-th independent neural network, and topb represents the probability from [h1, h2, ..., hj] to [hj]. n The first b semantic codes are selected according to their probability as the target semantic codes of the j-th independent neural network.
5. The method of claim 1, wherein each independent neural network is as follows: e j =ffn downj (Gelu (ffn gate (s j ))×ffn upj (s j )) in, e j S represents the sentiment semantic features of the j-th independent neural network. j ffn represents the semantic encoding of the target lexical unit selected by the j-th independent neural network. gate (s j ) indicates that s is calculated using the feedforward neural network of the gated network. j The gating weights involved in the j-th independent neural network required for each semantic encoding, ffn upj (s j ) indicates that the j-th independent neural network itself is used to process s. j The semantic output obtained after each semantic encoding, Gelu() represents the activation function, ffn downj () is used to map to a uniform dimension.
6. The method of claim 1, wherein the target neural network is as follows: e s =ffn down ′(Gelu (ffn gate (H))×ffn up (H)) in, e s H represents the shared semantic features, H represents the encoding sequence, which includes at least the global semantic encoding and the semantic encoding of all lexical units, and ffn gate (H) represents the weight allocation of each semantic code in H calculated using the feedforward neural network of the gated network, ffn up ′(H) represents the semantic output obtained after processing H using the feedforward neural network of the target neural network itself, Gelu() represents the activation function, and ffn down ′() is used to map to a uniform dimension.
7. The method as described in claim 1, wherein calculating the emotion probability corresponding to each emotion category based on the fusion features corresponding to each independent neural network to determine the emotion category to which the user's speech belongs specifically includes: Based on y j =softmax(ffn4(z j )) Calculate the sentiment probability corresponding to each sentiment category; where y j Let z represent the sentiment probability of the sentiment category corresponding to the j-th independent neural network. j represents the fusion feature of the j-th independent neural network, ffn4() represents the feedforward neural network, and softmax() represents the normalization function; Referring to the emotion probability corresponding to each emotion category, the emotion category with the highest probability is selected as the emotion category to which the user's voice belongs.
8. A smart cockpit, comprising: The lexical module is used to segment user speech into words, obtain a token sequence containing several lexical units, and place the global semantic information corresponding to the user speech into a set position in the token sequence; The encoding module is used to perform semantic encoding on the token sequence to obtain the global semantic encoding corresponding to the global semantic information, as well as the semantic encoding of the plurality of word elements; The first filtering module is used to process the semantic encoding of the plurality of lexical units using the gating network in the semantic capture model, and to filter out the target lexical unit semantic encoding corresponding to the plurality of independent neural networks; wherein, the plurality of independent neural networks are arranged side by side in the semantic capture model, and each independent neural network corresponds to a sentiment category; The network processing module is used to process the semantic encoding of the target word units corresponding to the plurality of independent neural networks to obtain the sentiment semantic features corresponding to the plurality of independent neural networks. The second filtering module is used to add a target neural network to the semantic capture model, and use the gating network to output the global semantic encoding and the semantic encoding of all words to the target neural network for processing, and output shared semantic features; wherein, the target neural network and the several independent neural networks are parallel; the shared semantic features are semantic features used by the several independent neural networks. The fusion module is used to fuse the shared semantic features, the global semantic encoding corresponding to the global semantic information, and the sentiment semantic features corresponding to the independent neural network in each independent neural network to obtain the fused features corresponding to the independent neural network. The calculation module is used to calculate the emotion probability corresponding to each emotion category based on the fusion features corresponding to each independent neural network, so as to determine the emotion category to which the user's voice belongs.
9. A computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps of the method according to any one of claims 1-7.
10. A vehicle comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the steps of the method according to any one of claims 1-7.