Multi-modal emotion calculation method and system based on hierarchical hypergraph knowledge extraction

Through the method based on hierarchical hypergraph structure, multimodal emotion calculation is performed using hypergraph convolution, the problem of extracting emotional interaction knowledge within and between modals is solved, and efficient emotional feature fusion and discrimination is achieved.

CN120372353APending Publication Date: 2025-07-25ECCOM NETWORK SYST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510448174.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The prior art is difficult to effectively extract the emotional interaction knowledge within and between modes in multimodal emotion calculation, resulting in redundancy or loss of information, and it is difficult for existing models to efficiently integrate multimodal information.

Method used

Using a method based on hierarchical hypergraph structure, the emotional interaction knowledge within and between modes is extracted through hypergraph multivariate connection relationships, and emotional knowledge is fusion by hypergraph convolution to construct single-modal and multimodal emotional representations.

Benefits of technology

Effectively overcome information loss and redundancy, extract highly expressed single-modal and multi-modal emotional characteristics, and improve emotional discrimination performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372353A_ABST
    Figure CN120372353A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal emotion calculation method and system based on hierarchical hypergraph knowledge extraction, and the method comprises the steps: obtaining a text, voice and visual sequences of a target object based on a video of the target object, and extracting text features, voice features and visual features based on the text, voice and visual sequences of the target object; the text features, the voice features and the visual features are subjected to feature encoding through a feature encoder to obtain text modes, voice modes and visual modes with sequential relations; performing single-mode hypergraph construction on the text, the voice and the visual modes with the sequential relationship, performing single-mode emotion knowledge extraction by using hypergraph convolution, and performing hypergraph output to obtain single-mode emotion representation; constructing a multi-modal full-connection hypergraph based on the single-modal emotion representation, carrying out multi-modal emotion knowledge extraction by utilizing hypergraph convolution, and carrying out hypergraph output to obtain final multi-modal emotion representation; and performing emotion classification and emotion polarity classification based on the final multi-modal emotion representation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and specifically, to a multi-modal emotion calculation method and system based on hierarchical hypergraph knowledge extraction. Background Art

[0002] Multi-modal emotion calculation refers to analyzing and understanding the emotions expressed by relying on multi-modal information such as text, speech, and vision, and exploring intelligent processing methods for human perception. In the prior art, there are the following defects in multi-modal emotion calculation: 1) How to effectively extract the emotion interaction knowledge within the modality, which mainly focuses on how to mine the interaction relationship between time points of each modality, and this requires the model to have the ability to encode long sequences; 2) How to effectively extract the emotion interaction knowledge between modalities. Due to the natural inconsistency of the feature distributions of multi-modal sequences, directly fusing them may lead to information redundancy or information loss, thus seriously damaging the model performance. Therefore, it is required to effectively extract the interaction information between the three modalities of text, speech, and vision on the basis of solving the first problem, and then form a highly unified multi-modal representation. Generally speaking, multi-modal emotion calculation is mainly divided into single-modal time series knowledge extraction and simultaneously modeling the interaction information between each single-modal sequence.

[0003] To address the above problems, previous methods often used recurrent neural networks, Transformers, and graph neural networks to model and solve the above problems. However, models based on recurrent neural networks are prone to gradient vanishing and long-term dependencies may have the problem of forgetting. And the model based on Transformer is still a sequence model and cannot efficiently fuse the information of all time steps. Therefore, it is difficult to extract highly expressive emotion knowledge. With the rapid development of graph convolutional neural networks, currently, models based on graph neural networks have been successively proposed to solve the problems existing in multi-modal emotion calculation, taking each time point as a node of the graph neural network and using graph convolution to model the relationship between them. However, the ordinary graph structure is a binary connection and cannot effectively mine the connection relationship of each time point.

[0004] Based on the above background, the present invention proposes a new multi-modal emotion calculation method based on hierarchical hypergraph knowledge extraction, and extracts the emotion interaction knowledge within the modality and between modalities through a hierarchical hypergraph structure. For the relationship within the modality, the hypergraph multi-element connection relationship is used to solve the long sequence dependence problem within the modality, avoid the problems existing in the previous sequence-based models and graph-based models, and effectively extract emotion knowledge. For the relationship between modalities, the high-order modeling ability of the hypergraph is used to solve the problem of insufficient information fusion between modalities existing in previous methods, and effectively extract the complementary emotion information between different modalities on the basis of removing redundant information, so as to obtain a unified multi-modal representation. Summary of the Invention

[0005] Aiming at the defects in the prior art, the purpose of the present invention is to provide a multi-modal sentiment calculation method and system based on hierarchical hypergraph knowledge extraction.

[0006] A multi-modal sentiment calculation method based on hierarchical hypergraph knowledge extraction provided by the present invention includes:

[0007] Step S1: Based on the video of the target object, obtain the text, speech, and visual sequences of the target object, and extract text features, speech features, and visual features based on the text, speech, and visual sequences of the target object;

[0008] Step S2: The text features, speech features, and visual features are respectively encoded through a feature encoder to obtain text, speech, and visual modalities with temporal relationships;

[0009] Step S3: Respectively construct single-modal hypergraphs for the text, speech, and visual modalities with temporal relationships, use hypergraph convolution for single-modal sentiment knowledge extraction, and then obtain single-modal sentiment representations through hypergraph output;

[0010] Step S4: Based on the single-modal sentiment representations, construct a multi-modal fully connected hypergraph, use hypergraph convolution for multi-modal sentiment knowledge extraction, and then obtain the final multi-modal sentiment representation through hypergraph output;

[0011] Step S5: Perform emotion classification and sentiment polarity classification based on the final multi-modal sentiment representation.

[0012] Preferably, the step S1 includes:

[0013] Step S1.1: Based on the video of the target object, obtain the text, speech, and visual sequences of the target object;

[0014] Step S1.2: Based on the obtained text sequence of the target object, extract text features through the BERT model;

[0015] Step S1.3: Based on the obtained speech sequence of the target object, extract speech features through COVAREP, and the speech features include: speech mel spectrogram features;

[0016] Step S1.4: Based on the obtained visual sequence of the target object, extract the video features of the speaker in the video frame through Facet, and the video features include facial localization information.

[0017] Preferably, the step S2 includes: The text features, speech features, and visual features are respectively encoded through a bidirectional long short-term memory network to obtain text, speech, and visual modalities with temporal relationships; then use a fully connected network to map different modal features to the same dimension;

[0018] Xm = FC(BiLSTM(U m ))

[0019] Wherein, U m represents the single-modal extracted features, and m ∈ {t, a, v} represent text, speech, and visual modalities respectively; BiLSTM(·) and FC(·) represent bidirectional long short-term memory network and fully connected network respectively; represents the encoded single-modal features, and T m represents the length of the time series, and d represents the unified mapping dimension.

[0020] Preferably, the step S3 includes:

[0021] Step S3.1: Use the KNN algorithm to construct single-modal hypergraphs for text, speech, and visual modality representations respectively;

[0022] Construct single-modal hypergraphs HG m = (N m , E m ) for text, speech, and vision respectively; wherein, the nodes N m of the hypergraph are each time point of the single-modal time series; E m represents the hyperedges of the single-modal hypergraph; then use the KNN algorithm for hypergraph node clustering, and each category is used as a hyperedge in the hypergraph, and its adjacency matrix is represented as follows:

[0023]

[0024] Step S3.2: After the construction of the single-modal hypergraph is completed, use hypergraph convolution to extract sentiment knowledge, which includes two aggregation processes: point-to-hyperedge and hyperedge-to-point:

[0025]

[0026] N m = H m E m

[0027] The entire aggregation process can be summarized as follows:

[0028]

[0029] Wherein, and represent the node degree matrix and hyperedge degree matrix respectively, θ m is the learning parameter in the convolution process, and W m is the unit weight matrix;

[0030] Step S3.3: Optimize the updated nodes using the self-attention mechanism and obtain the single-modal sentiment representation through average pooling;

[0031]

[0032] Among them, MeanPool(·) and SelfAtten(·) respectively represent the average pooling and self-attention functions, and the unimodal sentiment representation

[0033] Preferably, the step S5 includes: using a fully connected layer as a classifier, taking the final multimodal sentiment representation as input, and performing sentiment classification according to the preset task requirements to obtain a sentiment discrimination result.

[0034] Preferably, the method further includes: using the cross-entropy function as the loss of sentiment calculation, calculating the loss between the true category and the predicted category, and minimizing the loss through backpropagation, thereby updating the parameters in the multimodal sentiment calculation method based on hierarchical hypergraph knowledge extraction.

[0035] According to a multimodal sentiment calculation system based on hierarchical hypergraph knowledge extraction provided by the present invention, it includes:

[0036] Module M1: Based on the video of the target object, obtain the text, speech, and visual sequences of the target object, and extract text features, speech features, and visual features based on the text, speech, and visual sequences of the target object;

[0037] Module M2: The text features, speech features, and visual features are respectively encoded through a feature encoder to obtain text, speech, and visual modalities with temporal relationships;

[0038] Module M3: Respectively construct unimodal hypergraphs for the text, speech, and visual modalities with temporal relationships, use hypergraph convolution to extract unimodal sentiment knowledge, and then obtain unimodal sentiment representations through hypergraph output;

[0039] Module M4: Construct a multimodal fully connected hypergraph based on the unimodal sentiment representations, use hypergraph convolution to extract multimodal sentiment knowledge, and then obtain the final multimodal sentiment representation through hypergraph output;

[0040] Module M5: Perform emotion classification and sentiment polarity classification based on the final multimodal sentiment representation.

[0041] Preferably, the module M1 includes:

[0042] Module M1.1: Based on the video of the target object, obtain the text, speech, and visual sequences of the target object;

[0043] Module M1.2: Based on the obtained text sequence of the target object, extract text features through the BERT model;

[0044] Module M1.3: Based on the obtained speech sequence of the target object, extract speech features through COVAREP, and the speech features include: speech Mel spectrogram features;

[0045] Module M1.4: Based on the obtained visual sequence of the target object, extract the video features of the speaker in the video frame through Facet, and the video features include facial localization information;

[0046] The module M2 includes: The text features, speech features, and visual features are respectively encoded through a bidirectional long short-term memory network to obtain text, speech, and visual modalities with temporal relationships; then use a fully connected network to map the different modality features to the same dimension;

[0047] X m = FC(BiLSTM(U m ))

[0048] where U m represents the single-modal extraction features, m ∈ {t, a, v} respectively represent the text, speech, and visual modalities; BiLSTM(·) and FC(·) respectively represent the bidirectional long short-term memory network and the fully connected network; represents the encoded single-modal features, T m represents the length of the temporal sequence, and d represents the unified mapping dimension.

[0049] Preferably, the module M3 includes:

[0050] Module M3.1: Use the KNN algorithm to respectively construct single-modal hypergraphs for the text, speech, and visual modality representations;

[0051] Respectively construct single-modal hypergraphs HG m = (N m , E m ); where the nodes N m of the hypergraph are each time point of the single-modal temporal sequence; E m represents the hyperedges of the single-modal hypergraph; then use the KNN algorithm for hypergraph node clustering, and each category is used as a hyperedge in the hypergraph, and its adjacency matrix is represented as follows:

[0052]

[0053] Module M3.2: After the single-modal hypergraph construction is completed, use hypergraph convolution to extract emotion knowledge, which includes two aggregation processes: point-to-hyperedge and hyperedge-to-point:

[0054]

[0055] N m = H mE m

[0056] The entire aggregation process can be summarized as follows:

[0057]

[0058] Among them, and represent the node degree matrix and the hyperedge degree matrix, θ m is the learning parameter in the convolution process, and W m is the unit weight matrix;

[0059] Module M3.3: The updated nodes are optimized using the self-attention mechanism and the single-modal sentiment representation is obtained through average pooling;

[0060]

[0061] Among them, MeanPool(·) and SelfAtten(·) represent average pooling and the self-attention function respectively, and the single-modal sentiment representation

[0062] Preferably, the module M5 includes: using a fully connected layer as a classifier, taking the final multi-modal sentiment representation as the input, and performing sentiment classification according to the preset task requirements to obtain the sentiment discrimination result.

[0063] Compared with the prior art, the present invention has the following beneficial effects:

[0064] 1. The hierarchical hypergraph structure proposed by the present invention constructs a hypergraph using the clustering method for different time series points, fully excavates the direct connection relationship between time series points, and uses hypergraph convolution for sentiment knowledge fusion, effectively overcoming information loss and information redundancy; finally, the sentiment information is further refined through the hypergraph output, and then a highly expressive single-modal feature is obtained;

[0065] 2. On the basis of extracting hypergraph knowledge at one layer, the present invention constructs the second-layer hypergraph structure in a fully connected manner, avoiding the loss of modal information while using the high-order feature extraction ability of the hypergraph to maximally eliminate the fusion information loss and information misguidance caused by different modalities due to distribution;

[0066] 3. The present invention can better extract and fuse different modal information, and at the same time, the connection attribute of the hypergraph also provides a more intuitive model interpretation. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] By reading the following detailed description of the non-limiting embodiments with reference to the accompanying drawings, other features, objects and advantages of the present invention will become more apparent:

[0068] Figure 1Schematic diagram of a multimodal sentiment computing system based on hierarchical hypergraph knowledge extraction.

[0069] Figure 2 Flowchart of a multimodal sentiment computing method based on hierarchical hypergraph knowledge extraction. Specific implementation manner

[0070] The present invention will be described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several changes and improvements can still be made. These all fall within the protection scope of the present invention.

[0071] According to a multimodal sentiment computing method based on hierarchical hypergraph knowledge extraction provided by the present invention, as Figures 1 to 2 shown, it includes:

[0072] Step 1: Obtain the text (speech content), speech, and visual (video frame) sequences of the video speaker, and extract features through BERT, COVAREP, and Facet respectively.

[0073] Step 2: Perform feature encoding on the text, speech, and visual modalities respectively, optimize the initial features and add temporal information;

[0074] Step 3: Use the KNN algorithm to construct unimodal hypergraphs for the features encoded by the text, speech, and visual modalities respectively, and use hypergraph convolution to extract unimodal sentiment knowledge. Finally, obtain unimodal sentiment representations through hypergraph output.

[0075] Step 4: Construct a multimodal fully connected hypergraph from the unimodal sentiment representations, and use hypergraph convolution to extract multimodal sentiment knowledge. Finally, obtain the final multimodal sentiment representation through hypergraph output.

[0076] Step 5: Input the final sentiment representation into the sentiment discriminator, and perform emotion classification and sentiment polarity classification according to different tasks.

[0077] The present invention makes full use of the hierarchical hypergraph structure to extract sentiment knowledge within unimodal and between different modalities, obtains a unified multimodal sentiment representation, and thus performs more accurate sentiment discrimination.

[0078] Furthermore, in step 2, the feature encoding uses a bidirectional long short-term memory network for feature encoding, and uses a fully connected network to map features of different modalities to a unified dimension:

[0079] X m = FC(BiLSTM(U m ))

[0080] Among them, U m represents single-modal extraction features, and m ∈ {t, a, v} represent text, speech, and visual modalities respectively. BiLSTM(·) and FC(·) represent bidirectional long short-term memory network and fully connected network respectively. represents the encoded single-modal features, T m represents the length of the time series, and d represents the unified mapping dimension.

[0081] Furthermore, in step 3, single-modal hypergraphs HG m =(N m , E m ) of text, speech, and vision are constructed respectively, where the nodes N m of the hypergraph are each time point of the single-modal time series, that is, each hypergraph contains T m points. Subsequently, the KNN algorithm is used for hypergraph node clustering, and each category is used as a hyperedge in the hypergraph, and its incidence matrix is represented as follows:

[0082]

[0083] After the construction of the single-modal hypergraph is completed, hypergraph convolution is used to extract sentiment knowledge, which includes two aggregation processes: point-to-hyperedge and hyperedge-to-point:

[0084]

[0085] N m = H m E m

[0086] To improve the learning ability of hypergraph convolution, the entire aggregation process can be summarized as follows:

[0087]

[0088] Among them, and represent the node degree matrix and the hyperedge degree matrix, θ m is the learning parameter in the convolution process, and W m is the unit weight matrix.

[0089] The updated node features are input into the hypergraph output module, optimized by the self-attention mechanism, and the single-modal sentiment representation is obtained through average pooling, and the calculation is as follows:

[0090]

[0091] Among them, MeanPool(·) and SelfAtten(·) represent average pooling and self-attention function respectively, and the single-modal sentiment representation

[0092] Further, in step 4, for the unimodal sentiment representation, a fully-connected hypergraph is constructed to form a multimodal HG Mul , that is, all nodes are fully connected to each other, and the calculation is as follows.

[0093]

[0094] Subsequent hypergraph convolution and hypergraph output are the same as in step 3, and finally the multimodal sentiment representation Output is obtained.

[0095] Further, in step 5, a fully-connected layer is used as the classifier, and the final multimodal sentiment representation Output is used as the input to obtain the sentiment discrimination result.

[0096] Further, the cross-entropy function is used as the loss for sentiment calculation, the loss between the true class and the predicted class is calculated, and the loss is minimized through backpropagation, thereby updating the model parameters.

[0097] The present invention also provides a multimodal sentiment calculation system based on hierarchical hypergraph knowledge extraction. The multimodal sentiment calculation system based on hierarchical hypergraph knowledge extraction can be implemented by executing the process steps of the multimodal sentiment calculation method based on hierarchical hypergraph knowledge extraction. That is, those skilled in the art can understand the multimodal sentiment calculation method based on hierarchical hypergraph knowledge extraction as a preferred embodiment of the multimodal sentiment calculation system based on hierarchical hypergraph knowledge extraction.

[0098] The present invention adopts a hierarchical hypergraph structure, where the hypergraph can connect two or more vertices and has high-order modeling capabilities. The present invention uses a unimodal hypergraph network to model the mutual dependence relationship of long-time series, aiming to extract effective sentiment knowledge within a single modality. On this basis, a multimodal hypergraph network is used to establish emotional connections between different modalities, aiming to effectively extract emotional information from different modalities, thereby improving the final emotional discrimination performance.

[0099] Example 2

[0100] Example 2 is a preferred example of Example 1

[0101] According to the multimodal sentiment calculation system based on hierarchical hypergraph knowledge extraction provided by the present invention, it includes:

[0102] Unimodal encoder module: Since the present invention takes the time points of each modality feature as a point of the hypergraph, the first problem encountered when constructing a non-Euclidean structure hypergraph using the sequence features of the Euclidean structure is how to reduce the loss of time information. The present invention first uses a bidirectional long short-term memory network for feature encoding, and maps the features of different modalities to the same dimension through a fully-connected neural network to achieve feature alignment.

[0103] Single-modal sentiment knowledge extraction module: This module mainly includes three modules: single-modal hypergraph construction, hypergraph convolution, and hypergraph output. For the encoded single-modal time-series features, the present invention uses the KNN algorithm to construct a hypergraph structure, extracts the sentiment knowledge between time points through hypergraph convolution, and finally inputs it into the hypergraph output module to obtain a single-modal representation.

[0104] Multi-modal sentiment knowledge extraction module: This module mainly includes three modules: multi-modal hypergraph construction, hypergraph convolution, and hypergraph output. Among them, hypergraph convolution and hypergraph output are the same as those of the single-modal sentiment knowledge extraction module. For the representations output by different single-modal hypergraphs, the present invention constructs a hypergraph structure in a fully connected form, and then obtains the final multi-modal sentiment representation through hypergraph convolution and hypergraph output.

[0105] Sentiment discrimination module: The present invention uses a fully connected network as a sentiment classifier and inputs the final multi-modal sentiment representation for sentiment discrimination. According to different sentiment classification tasks, it can be divided into emotion classification and sentiment polarity classification.

[0106] Those skilled in the art know that in addition to implementing the systems, devices, and their respective modules provided by the present invention in the form of pure computer-readable program codes, the method steps can be logically programmed to enable the systems, devices, and their respective modules provided by the present invention to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, the systems, devices, and their respective modules provided by the present invention can be regarded as a kind of hardware component, and the modules included therein for implementing various programs can also be regarded as the structures within the hardware component; the modules for implementing various functions can also be regarded as both software programs for implementing the method and the structures within the hardware component.

[0107] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.

Claims

1. A multimodal sentiment computing method based on hierarchical hypergraph knowledge extraction, characterized in that Including: Step S1: Obtain the text, speech, and visual sequences of the target object based on the video of the target object, and extract text features, speech features, and visual features based on the text, speech, and visual sequences of the target object; Step S2: The text features, speech features, and visual features are respectively encoded through a feature encoder to obtain text, speech, and visual modalities with temporal relationships; Step S3: Perform single-modal hypergraph construction on the text, speech, and visual modalities with temporal relationships respectively, extract single-modal sentiment knowledge using hypergraph convolution, and then obtain single-modal sentiment representations through hypergraph output; Step S4: Construct a multi-modal fully-connected hypergraph based on the single-modal sentiment representations, extract multi-modal sentiment knowledge using hypergraph convolution, and then obtain the final multi-modal sentiment representations through hypergraph output; Step S5: Perform emotion classification and sentiment polarity classification based on the final multi-modal sentiment representations.

2. The multimodal sentiment calculation method based on hierarchical hypergraph knowledge extraction according to claim 1, wherein, The said Step S1 includes: Step S1.1: Obtain the text, speech, and visual sequences of the target object based on the video of the target object; Step S1.2: Extract text features through the BERT model based on the obtained text sequence of the target object; Step S1.3: Extract speech features through COVAREP based on the obtained speech sequence of the target object, and the speech features include: speech mel spectrogram features; Step S1.4: Extract the video features of the speaker in the video frame through Facet based on the obtained visual sequence of the target object, and the video features include facial localization information.

3. The multimodal sentiment calculation method based on hierarchical hypergraph knowledge extraction according to claim 1, characterized in that, The said Step S2 includes: The text features, speech features, and visual features are respectively encoded through a bidirectional long short-term memory network to obtain text, speech, and visual modalities with temporal relationships; then use a fully-connected network to map different modality features to the same dimension; X m = FC(BiLSTM(U m )) Among them, U m represents single-modal extraction features, where m ∈ {t, a, v} represent text, speech, and visual modalities respectively; BiLSTM(·) and FC(·) represent bidirectional long short-term memory network and fully connected network respectively; represents the encoded single-modal features, T m represents the length of the time series, and d represents the unified mapping dimension.

4. The multimodal sentiment calculation method based on hierarchical hypergraph knowledge extraction according to claim 1, wherein, The said Step S3 includes: Step S3.1: Perform single-modal hypergraph construction on the text, speech, and visual modality representations respectively using the KNN algorithm; Construct the single-modal hypergraph HG of text, speech, and vision respectively m =(N m , E m ); where the nodes N of the hypergraph m are each time point of a single-modal time series; E m represents the hyperedges of the single-modal hypergraph; Subsequently, the KNN algorithm is used for hypergraph node clustering, and each category is used as a hyperedge in the hypergraph, and its adjacency matrix is represented as follows: Step S3.2: After the single-modal hypergraph construction is completed, extract sentiment knowledge using hypergraph convolution, which includes two aggregation processes: point-to-hyperedge and hyperedge-to-point; N m = H m E m The entire aggregation process can be summarized as follows: Among them, and represent the node degree matrix and the hyperedge degree matrix, and θ m is the learning parameter in the convolution process, and W m is the unit weight matrix; Step S3.3: The updated nodes are optimized using the self-attention mechanism and the single-modal sentiment representations are obtained through average pooling; Among them, MeanPool(·) and SelfAtten(·) represent average pooling and self-attention functions respectively, and the unimodal sentiment representation 5. The multimodal sentiment calculation method based on hierarchical hypergraph knowledge extraction according to claim 1, wherein, The said Step S5 includes: Use a fully-connected layer as a classifier, take the final multi-modal sentiment representations as input, and perform sentiment classification according to the preset task requirements to obtain sentiment discrimination results.

6. The multimodal sentiment calculation method based on hierarchical hypergraph knowledge extraction according to claim 1, wherein The said method further includes: Use the cross-entropy function as the loss for sentiment calculation, calculate the loss between the true category and the predicted category, and minimize the loss through backpropagation, thereby updating the parameters in the multi-modal sentiment calculation method based on hierarchical hypergraph knowledge extraction.

7. A multimodal sentiment computing system based on hierarchical hypergraph knowledge extraction, characterized in that Including: Module M1: Obtain the text, speech, and visual sequences of the target object based on the video of the target object, and extract text features, speech features, and visual features based on the text, speech, and visual sequences of the target object; Module M2: The text features, speech features, and visual features are respectively encoded through a feature encoder to obtain text, speech, and visual modalities with temporal relationships; Module M3: Construct single-modal hypergraphs for text, speech, and visual modalities with temporal relationships respectively, extract single-modal sentiment knowledge using hypergraph convolution, and obtain single-modal sentiment representations through hypergraph output; Module M4: Construct a multi-modal fully connected hypergraph based on single-modal sentiment representations, extract multi-modal sentiment knowledge using hypergraph convolution, and obtain the final multi-modal sentiment representations through hypergraph output; Module M5: Perform emotion classification and sentiment polarity classification based on the final multi-modal sentiment representations.

8. The multimodal sentiment computing system based on hierarchical hypergraph knowledge extraction according to claim 7, wherein The module M1 includes: Module M1.1: Obtain the text, speech, and visual sequences of the target object based on the video of the target object; Module M1.2: Extract text features through the BERT model based on the obtained text sequence of the target object; Module M1.3: Extract speech features through COVAREP based on the obtained speech sequence of the target object, where the speech features include: speech mel spectrogram features; Module M1.4: Extract the video features of the speaker in the video frame through Facet based on the obtained visual sequence of the target object, where the video features include face localization information; The module M2 includes: The text features, speech features, and visual features are respectively feature-encoded through a bidirectional long short-term memory network to obtain text, speech, and visual modalities with temporal relationships; then the different modality features are mapped to the same dimension using a fully connected network; X m = FC(BiLSTM(U m )) Among them, U m represents single-modal extraction features, where m ∈ {t, a, v} represent text, speech, and visual modalities respectively; BiLSTM(·) and FC(·) represent bidirectional long short-term memory network and fully connected network respectively; represents the encoded single-modal features, and T m represents the length of the time series, and d represents the unified mapping dimension.

9. The multimodal sentiment computing system based on hierarchical hypergraph knowledge extraction according to claim 7, wherein, The module M3 includes: Module M3.1: Construct single-modal hypergraphs for text, speech, and visual modality representations respectively using the KNN algorithm; Construct single - modality hypergraphs HG for text, speech, and vision respectively m =(N m , E m ); where the nodes N of the hypergraph m are each time point of a single - modality time series; E m represents the hyper - edges of a single - modality hypergraph; Subsequently, use the KNN algorithm for hypergraph node clustering, and each category is used as a hyper - edge in the hypergraph, and its adjacency matrix is represented as follows: Module M3.2: After the single-modal hypergraph construction is completed, extract sentiment knowledge using hypergraph convolution, which includes two aggregation processes: point-to-hyperedge and hyperedge-to-point; N m = H m E m The entire aggregation process can be summarized as follows: Among them, and represent the node degree matrix and the hyperedge degree matrix, and θ m is the learning parameter in the convolution process, and W m is the unit weight matrix; Module M3.3: The updated nodes are optimized using the self-attention mechanism and the single-modal sentiment representations are obtained through average pooling; where MeanPool(·) and SelfAtten(·) represent the average pooling and self-attention functions respectively, and the unimodal sentiment representation 10. The multimodal sentiment computing system based on hierarchical hypergraph knowledge extraction according to claim 7, wherein The module M5 includes: Use a fully connected layer as a classifier, take the final multi-modal sentiment representations as input, and perform sentiment classification according to the preset task requirements to obtain sentiment discrimination results.