A Multimodal Data Sentiment Analysis Method for Open-Source Intelligence
By encapsulating the multimodal sentiment analysis model within the Spark Streaming framework, designing the multimodal resource classification matrix, graphic and text data pair enhancement method and multi-label content label matrix, the problem of multimodal data processing in new media information is solved, efficient and accurate multimodal sentiment analysis is achieved, and the intelligent mining needs of massive data is met.
Patent Information
- Application Number
- CN202310596095.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-24
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2043-05-24
AI Technical Summary
The prior art is difficult to effectively process multimodal data in new media information, especially in fusion processing, video key information extraction, graphic and text data enhancement and computing model construction, and there are problems such as high computational complexity and difficulty in dealing with massive information.
A multimodal data sentiment analysis method for open source intelligence is proposed. By encapsulating the multimodal sentiment analysis model within the Spark Streaming framework, a multimodal resource classification matrix, a graphic and text data pair enhancement method and a multi-label content tag matrix are designed to realize multimodal sentiment analysis of new media information.
It realizes efficient, accurate and real-time multi-modal sentiment analysis of new media information, which can meet various application needs such as intelligence mining, public opinion monitoring, topic monitoring and tracking, brand reputation mining, etc., reduces the computational complexity and supports intelligent mining and analysis of massive data.
Smart Images

Figure CN116561639B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of artificial intelligence, big data, and sentiment analysis, and particularly relates to a multi-modal data sentiment analysis method for open-source intelligence. Background Art
[0002] With the development of digital technology, network technology, and mobile communication technology, new media has become an important dissemination form for providing information and services to users. The sentiment analysis of new media information has also become an important research direction for Internet content security supervision and control. The content composition of new media information is more diverse, including both single-modal forms such as pure text, pure pictures, and pure videos, and multi-modal forms such as text + pictures and text + videos. Traditional text feature-based sentiment analysis methods are no longer suitable for processing the sentiment analysis of new media information due to the lack of modeling of multi-modal data. Existing sentiment analysis methods combining text and pictures are only applicable to specific platforms, such as the sentiment analysis of Internet users on self-media platforms like Weibo and WeChat. For the sentiment analysis of video information, it mainly utilizes the text, images, and sounds in the video, etc., and calculates the emotional tendency of the characters in the video by extracting key features, with insufficient utilization of video description information and lack of overall sentiment analysis. In addition, the complexity of existing methods is generally high, making it difficult to handle the massive new media information on the Internet.
[0003] In view of the above deficiencies, the present invention proposes a multi-modal sentiment analysis method and system for open-source intelligence, which establishes a multi-modal sentiment analysis model for the text, video, and images contained in new media information, and constructs an extensible multi-modal sentiment analysis system in combination with big data technology, realizing the efficient, accurate, and real-time multi-modal sentiment analysis of new media information, and being able to meet the real-time / quasi-real-time mining and analysis of various applications such as intelligence mining, public opinion monitoring, topic monitoring and tracking, and brand reputation mining. Summary of the Invention
[0004] (1) Technical Problems to be Solved
[0005] The technical problem to be solved by the present invention is how to provide a multi-modal data sentiment analysis method for open-source intelligence to solve the problems of fusion processing of new media multi-modal information, extraction of key video information, enhancement of graphic and text data, and construction of a calculation model, and to solve the problems of cross-domain technology integration and low resources.
[0006] (2) Technical Solutions
[0007] To solve the above technical problems, the present invention proposes a multi-modal data sentiment analysis method for open source intelligence. The method includes: encapsulating a multi-modal sentiment analysis model within the Spark Streaming framework to implement a resource classification matrix operator, a graphic-text data pair enhancement operator, a multi-modal algorithm operator, and a multi-label content operator. The processing process of the method is as follows: First, preprocess the input data received from HDFS. Second, call the resource classification matrix operator to classify text, video, and images. Third, call the graphic-text data pair enhancement operator to enhance graphic-text data, call the multi-modal algorithm operator and the multi-label content operator to achieve sentiment prediction. Finally, write the predicted results to Kafka to complete the entire process of sentiment prediction.
[0008] (III) Beneficial Effects
[0009] The present invention proposes a multi-modal data sentiment analysis method for open source intelligence, and the beneficial effects are reflected in the following aspects:
[0010] (1) A multi-modal resource classification matrix is designed. For the text, video, and image information of new media, the resource classification matrix is divided into a video frame extraction and decomposition into picture category, a pure text word vector extraction category, a video frame extraction of text category, and a pure image feature extraction category. Through the classification and aggregation extraction of different information, it is uniformly output as network structure feature information conforming to feature fusion, solving the problem of unified representation of new media multi-modal information.
[0011] (2) A data enhancement method for graphic-text pairs is proposed. For the pictures in the picture-text pairs, the enhanced pHash algorithm is used to calculate the Hamming distance through discrete cosine transform to increase the positive picture samples; for the descriptive text in the picture-text pairs, the TF-IDF similarity between relevant sentences in the corpus and the descriptive text is calculated and increased as the positive text samples, making up for the shortage of training samples, reducing the imbalance of training samples, and improving the effect of model training.
[0012] (3) A multi-label sentiment classification method based on a content label matrix is proposed. This method can further fuse the multi-level label features of the content on the basis of the output results of the multi-modal resource classification matrix, realize the multi-label sentiment classification of new media information, make up for the deficiency of simply relying on the content itself for analysis, and improve the rationality of the sentiment polarity classification of new media information.
[0013] (4) A new media information multi-modal sentiment analysis system is proposed. Based on the streaming processing framework Spark Streaming, by functionalizing video frame extraction technology, Faster-RCNN object detection network, GRU model, and graphic-text feature fusion, the technical fusion of big data + deep learning is realized, meeting the requirements of scalability and low-resource applications, and supporting the intelligent mining and analysis of massive data. Brief Description of the Drawings
[0014] Figure 1 It is a flowchart of the resource classification matrix of the present invention;
[0015] Figure 2 It is a flowchart of enhancing graphic and text data pairs of the present invention;
[0016] Figure 3 It is a schematic diagram of the multi-modal sentiment analysis model of the present invention;
[0017] Figure 4 It is a flowchart of enhancing processing of graphic and text data pairs of the present invention;
[0018] Figure 5 It is a framework diagram of the multi-modal sentiment analysis system for massive new media information of the present invention. Detailed Embodiments
[0019] To make the objectives, content and advantages of the present invention clearer, the following further describes in detail the specific embodiments of the present invention in conjunction with the drawings and embodiments.
[0020] The present invention discloses a multi-modal sentiment analysis method and system for open-source intelligence, mainly solving the multi-modal sentiment analysis of massive new media information, specifically including: (1) It is necessary to solve the problems of fusion processing of new media multi-modal information, extraction of key video information, enhancement of graphic and text data, and construction of a calculation model, and provide a high-accuracy multi-modal sentiment polarity calculation model; (2) It is necessary to solve the problems of cross-domain technology integration and low resources, support expansion according to the scale of calculation tasks, provide an efficient mining system for massive data, and meet the real-time prediction of the sentiment polarity of massive multi-modal data.
[0021] 1. Data Preprocessing
[0022] In the data preprocessing stage, data cleaning and word segmentation are carried out. Among them, regular matching and the like are used for data cleaning, mainly filtering out interference information that affects the semantic continuity of words, including link parts, special characters in other encodings, meaningless information such as #@¥%……&*, and partial information of numbers and English.
[0023] 2. Resource Classification Matrix Operator
[0024] The resource classification matrix includes the classification and processing process of multi-modal data, and the input data is divided into three cases: video, image, and text for processing.
[0025] Among them
[0026] For video information, key frame extraction needs to be carried out through the FFmpge frame extraction technology to obtain image information;
[0027] For the image information, it is necessary to determine whether there is text information in the image. For pictures containing text information, text extraction technology is used to achieve text extraction;
[0028] For the text information, operations such as text content filtering need to be performed.
[0029] The specific resource classification matrix process is as Figure 1 shown.
[0030] (1) Video frame extraction
[0031] Video data is very similar to image data, both are data composed of pixel points. Video data in the non-audio part can basically be regarded as the splicing of multiple frames (images) of image data, that is, the combination of three-dimensional images. The video stream data analysis of the present invention adopts the key frame extraction technology (IPB frames), and uses FFmpeg to extract I frames. The number of I frames is small within a certain period of time, but the information contained is the most. After extracting multiple frames of image data, image processing is performed. After text extraction, it can be stored in the graph database and text database, and UUID is used as the unique identifier of the graph-text pair.
[0032] (2) Image text extraction
[0033] There will be text in the image data extracted from the video stream or in the pure image data. Extracting this text has an important impact on the emotional expression of the overall multi-modal data news information. Therefore, only by maximizing the acquisition of the semantic information of the video and the image can the overall emotional expression be better analyzed. For image text extraction, this method adopts the backbone network MobileNetV3 of the Differentiable Binarization+CRNN algorithm of the latest technology PaddleOCR for text detection and recognition, improving the overall detection rate and recognition rate.
[0034] 3. Graph-text data pair enhancement operator
[0035] The enhancement of image-text data pairs is to build an image-text database on the existing data and achieve the balance of image and text data in the image-text pairs through data augmentation. The image-text data pair enhancement operator first judges the proportion of images and texts in the processed image-text pairs. For the case where the proportion of images is small, it selects to augment the image data. For the case where the proportion of texts is small, it selects to augment the text data. For the image augmentation in the image-text pair, the enhanced pHash algorithm is used to perform similarity comparison with the existing images in the image database. The Hamming distance is calculated through discrete cosine transform. If the similarity threshold is met, the image data is augmented. If the threshold condition is not met after traversing the database, image sample augmentation is performed using techniques such as edge expansion, random cropping, size scaling, horizontal and vertical flipping, etc. For the text augmentation in the image-text pair, the TF-IDF algorithm is used to calculate the similarity with the sentences in the corpus. If the threshold condition is met after traversing the database, text data augmentation is performed using techniques such as synonym replacement, random addition, random exchange, etc. The specific process flow of the image-text data pair enhancement module is as Figure 2 shown.
[0036] When enhancing the image-text data pair, the image uses the enhanced pHash algorithm to calculate the Hamming distance through discrete cosine transform for data augmentation; the text uses the TF-IDF similarity between the relevant sentences in the corpus and the description text to expand the data. The specific relevant processing flow is as Figure 4 shown.
[0037] As Figure 4 shown, the image processing flow is as follows:
[0038] S401. Perform size transformation on the image, for example, shrink the picture to 32*32 size;
[0039] S402. Perform grayscale processing on the image, for example, process it through the average method to improve the image processing speed;
[0040] S403. Perform discrete cosine transform and region selection, calculate the DCT and its mean value, and select the representative region, for example, select the upper left 8*8 matrix;
[0041] S404. Calculate the Hash value, convert each DCT value into 0 or 1, and generate a binary array;
[0042] S405. Calculate the image similarity by calculating the Hamming distance;
[0043] S406. Compare with the predefined threshold and output the result.
[0044] The text processing flow is as follows:
[0045] S411. Calculate the term frequency TF of the word in the document;
[0046] S412. Standardize TF and TF-IDF to avoid being affected by text length;
[0047] S413. Calculate the inverse document frequency IDF of words;
[0048] S414. Calculate the TF-IDF values of words to obtain the multi-dimensional numerical vector of each text;
[0049] S415. Calculate the similarity value between two texts through cosine similarity;
[0050] S416. Compare with the predefined threshold and output the result.
[0051] 4. Multi-modal sentiment analysis model
[0052] The multi-modal sentiment analysis model is a pre-trained offline model. The training data comes from a large amount of new media information on the Internet and is annotated with multi-label content. The multi-modal sentiment analysis model is encapsulated as a multi-modal algorithm operator.
[0053] Such as Figure 3 shown, the algorithm model uses the Attention mechanism to fuse image and text information.
[0054] For the image part, the pre-trained Faster-RCNN model is used to extract the pooled-ROI features and localization features of each region. After passing through the FC layer, the two types of features are projected into the same embedding space.
[0055] For the text part, all the words of a sentence are first input into the GRU layer to obtain word vectors. Then, these word vectors are calculated through the self-attention mechanism to obtain the corresponding weights. Finally, the weighted sum is used to obtain the vector representation of the sentence. All the sentence vectors of a document are input into the GRU layer to obtain the sentence vectors with enhanced semantics.
[0056] The last layer is that each sentence vector and each image vector use attention to calculate the corresponding weights, and then the weighted sum is used to obtain a document vector. If there are M images, M document vectors will be obtained, representing different vector descriptions corresponding to different images. Multiple document vectors calculate the corresponding weights through self-attention, and then the weighted sum is used to obtain the final document vector description D. Finally, the task layer is connected to perform Softmax to obtain the multi-classification result.
[0057] Such as Figure 3 shown, the specific processing steps of the model are as follows:
[0058] S31. Multi-modal feature extraction and fusion
[0059] Feature extraction is the core of the multi-modal framework. First, the semantic features of text information are constructed. At the word vector stage, the pre-trained word vectors of Bert are first input and represented by W it where i represents the word number, t represents the current sentence, and after passing through a bidirectional GRU, hidden state representations in two directions are obtained Then, Attention is used to calculate the importance weight α it of each h it . After normalizing the weights by Softmax, a weighted sum of h it is taken to obtain the embedded vector representation s i of the sentence
[0060] In the text-image feature fusion stage, first, the sentence embedded vector s i is input. After passing through a bidirectional GRU, hidden states in two directions are obtained and concatenated to get the hidden state h i of each sentence. The feature vector m j of each image is extracted using Faster-RCNN. Then, m j is used to perform Attention on h i . By calculating the inner product of the two, a non-linear transformation between the image vector and the sentence vector is achieved. After passing through Softmax, the importance weight β i corresponding to each transformed h j is obtained. Finally, a weighted sum of the transformed h i is taken to obtain the text representation D of the document for each image i .
[0061] The text representations D i generated for different images are input. The corresponding weights r i are calculated using Attention, and then a weighted sum is taken to obtain the final document vector D, and the overall feature representation is I n , where n represents the total number of fused text and images
[0062] S32. Achieve multi-label content sentiment output through the multi-label content operator
[0063] The multi-label content operator is used for multi-label content sentiment output. Using the sentiment set S = {S 1(l,p,q) , S 2(l,p,q) ,..., S n(l,p,q) , n ∈ the amount of input data, (l, p, q) ∈ three-level multi-label combination}, according to the overall feature vector I nThe model parameters are updated using cross - entropy as the target loss function, and the labels of the S emotion set are output through the Softmax function. The multi - label content is classified into multiple levels according to the new media news information type, that is, the overall emotion label is no longer the three types of positive, negative, and neutral, but fine - grained multi - level classification labels. The label classification is stored in three levels. The first level is the information flow type, and the information flow type set = {video stream, image stream, text stream, mixed stream, l ∈ information flow type}. The second level is the news information type, and the news information type set = {law, finance, entertainment, technology, sports, military}, p ∈ news information type. The third level is the emotion expression type, and the emotion expression type set = {praise, neutral, resistance, criticism}. Conducting multi - level emotion classification on the news information type can enable readers to more accurately grasp the emotions in multiple fields of the news.
[0064] 5. Multi - modal Emotion Analysis System for Massive New Media Information
[0065] The processing framework of the multi - modal emotion analysis system for massive new media information is as Figure 5 shown:
[0066] As Figure 5 shown, the multi - modal emotion analysis system for massive new media information integrates big data + deep learning technologies. By encapsulating the multi - modal emotion analysis model within the Spark Streaming framework, resource classification matrix operators, graphic - text data pair enhancement operators, multi - modal algorithm operators, and multi - label content operators are implemented. The system processing process is as follows: First, perform pre - processing operations such as cleaning and word segmentation on the input data received from HDFS. Second, call the resource classification matrix operator to classify text, video, and images. Third, call the graphic - text data pair enhancement operator to enhance graphic - text data, call the multi - modal algorithm operator and the multi - label content operator to achieve emotion prediction. Finally, write the predicted results to Kafka to complete the entire process of emotion prediction.
[0067] The present invention discloses a multi - modal emotion analysis method and system for open - source intelligence, and the main advantages are reflected in the following aspects:
[0068] (1) A multi - modal resource classification matrix is designed. For the text, video, and image information of new media, the resource classification matrix is divided into categories such as video frame extraction and decomposition into pictures, pure text word vector extraction, video frame extraction of text, and pure image feature extraction. Through the classification and aggregation extraction of different information, it is uniformly output as network structure feature information that conforms to feature fusion, solving the problem of unified representation of new media multi - modal information.
[0069] (2) A data augmentation method for image-text pairs is proposed. For the images in the image-text pairs, the enhanced pHash algorithm is used to calculate the Hamming distance through discrete cosine transform to increase the positive image samples. For the descriptive texts in the image-text pairs, the TF-IDF similarity between the relevant sentences in the corpus and the descriptive text is calculated to increase the positive text samples, making up for the shortage of training samples, reducing the imbalance of training samples, and improving the effect of model training.
[0070] (3) A multi-label sentiment classification method based on a content label matrix is proposed. This method can further fuse the multi-level label features of the content on the basis of the output results of the multi-modal resource classification matrix, realizing the multi-label sentiment classification of new media information, making up for the deficiency of simply relying on the content itself for analysis, and improving the rationality of the sentiment polarity classification of new media information.
[0071] (4) A multi-modal sentiment analysis system for new media information is proposed. Based on the streaming processing framework Spark Streaming, through the functionalization of video frame extraction technology, Faster-RCNN object detection network, GRU model, and image-text feature fusion, the technical integration of big data + deep learning is realized, meeting the requirements of scalability and low-resource applications, and supporting the intelligent mining and analysis of massive data.
[0072] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and deformations can be made, and these improvements and deformations should also be regarded as the protection scope of the present invention.
Claims
1. A multi-modal data sentiment analysis method for open source intelligence, characterized in that, the method includes: encapsulating a multi-modal sentiment analysis model within the Spark Streaming framework to implement a resource classification matrix operator, a graphic-text data pair enhancement operator, a multi-modal algorithm operator, and a multi-label content operator; the processing process of the method is as follows: First, perform preprocessing operations on the input data received from HDFS. Second, call the resource classification matrix operator to classify and process text, video, and images. Third, call the graphic-text data pair enhancement operator to enhance the graphic-text data, call the multi-modal algorithm operator and the multi-label content operator to achieve sentiment prediction. Finally, write the predicted results to Kafka to complete the entire process of sentiment prediction; wherein, calling the multi-modal algorithm operator and the multi-label content operator to achieve sentiment prediction includes: S31. Multi-modal feature extraction and fusion First, construct the semantic features of the text information. At the word vector stage, first input the word vectors pre-trained by Bert, denoted by W it where i represents the word number, t represents the current sentence, and the hidden state representations in two directions are obtained through bidirectional GRU Then, use Attention to calculate the importance weight α it of each h it . After normalizing the weights by Softmax, perform weighted summation on h it to obtain the embedded vector representation s of the sentence i ; In the image-text feature fusion stage, the sentence embedding vector s is first input i , after bidirectional GRU, we get the hidden states in two directions, and concatenate them to get the hidden state h of each sentence i , use Faster-RCNN to extract the feature vector m of each image j , then use m j For h i As Attention, by calculating the inner product of the two, the nonlinear transformation of image vector and sentence vector is realized, and then the Softmax is used to obtain each transformed h i The corresponding importance weight β j , and finally the converted h i The weighted summation is used to obtain the text representation D of the document for each image. i ; Input the text representation D generated for different images i , and calculate the corresponding weight r using Attention i , then perform weighted summation to obtain the final document vector D, and the overall feature representation is I n , where n represents the total number of fused text and images; S32. Implement multi-label content sentiment output through the multi-label content operator The multi-label content operator is used for multi-label content sentiment output. Using the sentiment set S = {S 1(l,p,q) , S 2(l,p,q) ,..., S n(l,p,q) , n ∈ the amount of input data, (l, p, q) ∈ three-level multi-label combination}, according to the overall feature vector I n The model parameters are updated using cross-entropy as the target loss function, and the labels of the S sentiment set are output through the Softmax function; the multi-label content is classified into multiple levels according to the new media news information type, that is, the overall sentiment label is no longer the three types of positive, negative, and neutral, but fine-grained multi-level classification labels; the label classification is stored in three levels. The first level is the information flow type, and the information flow type set = {video stream, image stream, text stream, mixed stream, l ∈ information flow type; the second level is the news information type, and the news information type set = {law, finance, entertainment, technology, sports, military}, p ∈ news information type; the third level is the sentiment expression type, and the sentiment expression type set = {praise, neutral, resistance, criticism}, q ∈ sentiment expression type; multi-level sentiment classification of news information types enables readers to more accurately grasp the multi-field sentiment of news.
2. The multi-modal data sentiment analysis method for open source intelligence according to claim 1, characterized in that, the resource classification matrix operator includes the classification and processing process of multi-modal data, and the input data is divided into three cases: video, image, and text for processing. Among them, for video information, key frame extraction is performed through the FFmpge frame extraction technology to obtain image information; for image information, it is judged whether there is text information in the image. For pictures containing text information, text extraction is realized by using text extraction technology; for text information, text content filtering processing is performed.
3. The multi-modal data sentiment analysis method for open source intelligence according to claim 2, characterized in that, for video information, the video stream data analysis adopts the key frame extraction technology, and the I frame is extracted by using FFmpeg. The number of I frames is small within a period of time, but the information contained is the most. After extracting multiple frame image data, image processing is performed. After text extraction, it is stored in the graph database and the text database, and UUID is used as the unique identification code of the graphic-text pair.
4. The multi-modal data sentiment analysis method for open source intelligence according to claim 2, characterized in that, for image information, the text extraction technology adopts the backbone network MobileNetV3 of the Differentiable Binarization+CRNN algorithm of PaddleOCR to detect and recognize text.
5. The multi-modal data sentiment analysis method for open source intelligence according to any one of claims 2-4, characterized in that, The graphic-text data pair enhancement operator first determines the proportion of images and text in the processed graphic-text pair. For the case where the proportion of images is small, it selects to perform image data augmentation. For the case where the proportion of text is small, it selects to perform text data augmentation. For the image augmentation in the graphic-text pair, it uses the enhanced pHash algorithm to perform similarity comparison with the existing images in the image database, calculates the Hamming distance through discrete cosine transform. If the similarity threshold is met, it performs image data augmentation. If the threshold condition is not met after traversing the database, it uses techniques such as edge expansion, random cropping, size scaling, and horizontal and vertical flipping to perform image sample augmentation. For the text augmentation in the graphic-text pair, it uses the TF-IDF algorithm to calculate the similarity with the sentences in the corpus. If the threshold condition is met after traversing the database, it uses techniques such as synonym replacement, random addition, and random exchange to perform text data augmentation.
6. The multi-modal data sentiment analysis method for open-source intelligence as claimed in claim 5, characterized in that calculating the Hamming distance through discrete cosine transform, and if the similarity threshold is met, performing image data augmentation includes: S401. Performing size transformation on the image; S402. Performing grayscale processing on the image; S403. Performing discrete cosine transform and region selection, calculating the DCT and its mean value, and selecting the representative region; S404. Calculating the Hash value, converting each DCT value into 0 or 1 to generate a binary array; S405. Calculating the image similarity by calculating the Hamming distance; S406. Comparing with the predefined threshold and outputting the result.
7. The multi-modal data sentiment analysis method for open-source intelligence as claimed in claim 5, characterized in that using the TF-IDF algorithm to calculate the similarity with the sentences in the corpus, and if the threshold condition is met after traversing the database includes: S411. Calculating the term frequency TF of the word in the document; S412. Standardizing the TF to avoid being affected by the text length; S413. Calculating the inverse document frequency IDF of the word; S414. Calculating the TF-IDF value of the word to obtain the multi-dimensional numerical vector of each text; S415. Calculating the similarity value between two texts through cosine similarity; S416. Comparing with the predefined threshold and outputting the result.
8. The multi-modal data sentiment analysis method for open-source intelligence as claimed in claim 5, characterized in that invoking the multi-modal algorithm operator and the multi-label content operator to implement sentiment prediction includes: The algorithm model uses the Attention mechanism to fuse image and text information; For the image part, the pre-trained Faster-RCNN model is used to extract the pooled-ROI features and localization features of each region. After the two types of features pass through the FC, they are projected into the same embedding space; For the text part, first, all the words of a sentence are input into the GRU layer to obtain word vectors. Then, these word vectors are calculated through the self-attention mechanism to obtain the corresponding weights. Finally, the weighted sum is used to obtain the vector representation of the sentence. All the sentence vectors of a document are input into the GRU layer to obtain the sentence vectors with enhanced semantics; The last layer is that each sentence vector and each image vector are used to calculate the corresponding weights through attention, and then the weighted sum is used to obtain a document vector. If there are M images, M document vectors will be obtained, representing different vector descriptions corresponding to different images; multiple document vectors calculate the corresponding weights through self-attention, and then the weighted sum is used to obtain the final document vector description D. Finally, the task layer is connected to perform Softmax to obtain the multi-classification result.
Citation Information
Patent Citations
Multi-mode detection method for malicious images and texts
CN112417194A
Line bolt defect identification method based on combination of Phash algorithm and deep learning
CN113627378A
Text enhancement method and device, electronic equipment and storage medium
CN113822047A