File travel material feature extraction and classification method and system

By receiving and processing multimodal cultural and tourism materials and combining feature fusion models of text, images, and video data, the accuracy and confidence issues of sentiment classification of multimodal cultural and tourism materials have been solved, achieving efficient and accurate sentiment analysis.

CN121365265APending Publication Date: 2026-01-20SICHUAN TOURISM UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411724745.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-28
Publication Date
2026-01-20

AI Technical Summary

Technical Problem

Existing technologies struggle to fully reflect emotional expression when processing multimodal cultural and tourism materials. Single-modal analysis can easily lead to biased emotional classification results, and the confidence level is low when semantics are ambiguous or information is insufficient. Furthermore, there is a lack of integrated utilization of multimodal information.

Method used

By receiving text, image, and video data, word segmentation, sentiment word recognition, and syntactic analysis are performed to generate text features. When the confidence level is insufficient, associated multimedia data is extracted, and multimodal information is combined with adaptive weighted integration to construct a multimodal feature fusion model for sentiment classification.

Benefits of technology

It significantly improves the accuracy and reliability of sentiment classification, especially performing well in complex sentiment expressions and ambiguous semantic scenarios, ensuring the accuracy and robustness of classification results and meeting the actual needs of the cultural and tourism industry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121365265A_ABST
    Figure CN121365265A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of big data processing, and provides a text travel material feature extraction and classification method and system, and the method comprises the steps: receiving text travel materials, including text, picture and / or video data; performing word segmentation, sentiment word recognition and syntactic analysis on the text to generate a first text feature, and determining a first sentiment classification of the first text feature; if the first sentiment classification confidence coefficient is smaller than a first threshold value, extracting picture and / or video data associated with the text; performing target detection on the picture and / or video data associated with the text to generate a first multimedia feature; determining a second sentiment classification of the first text feature according to the first text feature, the first multimedia feature and the first sentiment classification; and associating and storing the first text feature and the second sentiment classification.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of big data processing, and specifically relates to a travel material feature extraction and classification method and system. BACKGROUND

[0002] In recent years, the cultural tourism industry has developed rapidly, and tourists have generated a large amount of travel materials through social media, short video platforms and online travel websites. These materials exist in the form of text, pictures and videos, covering multi-angle descriptions of tourists on scenic spots, services and activities, reflecting not only the actual experience of tourists, but also carrying rich emotional expressions. How to efficiently extract the features of these materials and classify them emotionally has become an important research direction to improve the digital and intelligent level of the travel industry.

[0003] Traditional travel material analysis methods are usually limited to the processing of single-modal data, such as keyword extraction of text, object detection of images, etc. These methods may perform well in certain specific tasks, but they have the following limitations when faced with multi-modal travel materials:

[0004] The multi-modal characteristics of travel materials (such as text descriptions and related pictures or videos) make the analysis results of a single modality often unable to fully reflect the emotional expression of the material. For example, a user may express dissatisfaction with the crowd in the text, but the associated picture may show beautiful natural scenery, and analyzing a single modality may lead to biased emotional classification results.

[0005] The emotional expression of travel materials often contains a mixture of positive and negative emotions, or may even be expressed through ambiguous language. For example, "There are too many people, but the scenery is indeed beautiful" contains positive evaluation of the scenery and negative evaluation of the crowd. Keyword or syntax analysis alone cannot accurately capture this complex emotion.

[0006] The emotional classification results of single-modal data sometimes have low confidence, especially when the text content is semantically ambiguous or lacks information. Lack of integration and use of multi-modal information may result in insufficient reliability of emotional classification. SUMMARY

[0007] To solve the problems in the prior art, the present application provides a travel material feature extraction and classification method, which comprises the following steps:

[0008] Receiving travel materials, including text, picture and / or video data;

[0009] Performing word segmentation, sentiment word recognition and syntax analysis on the text to generate first text features and determine the first emotional classification of the first text features;

[0010] If the first sentiment classification confidence is less than a first threshold, extracting picture and / or video data associated with the text; performing target detection on the picture and / or video data associated with the text to generate a first multimedia feature;

[0011] Determining a second sentiment classification of the first text feature according to the first text feature, the first multimedia feature, and the first sentiment classification;

[0012] Associating and saving the first text feature and the second sentiment classification.

[0013] Another aspect of the present application also provides a travel material feature extraction and classification system, the system comprising the following modules:

[0014] A first collection module for receiving travel materials, including text, picture and / or video data;

[0015] A first classification module for performing word segmentation, sentiment word recognition and syntax analysis on the text to generate a first text feature and determine a first sentiment classification of the first text feature;

[0016] A second collection module for extracting picture and / or video data associated with the text if the first sentiment classification confidence is less than a first threshold;

[0017] An extraction module for performing target detection on the picture and / or video data associated with the text to generate a first multimedia feature;

[0018] A second classification module for determining a second sentiment classification of the first text feature according to the first text feature, the first multimedia feature, and the first sentiment classification;

[0019] A storage module for associating and saving the first text feature and the second sentiment classification.

[0020] The present application proposes a travel material feature extraction and classification method and system, which realizes efficient and accurate sentiment classification and data management by combining multi-modal information of text, picture and / or video, and has the following beneficial effects:

[0021] By constructing a multi-modal feature fusion model, the text feature and the picture and / or video feature are adaptively weighted and integrated, the complementary information between different modalities is fully mined, and the accuracy of sentiment classification is significantly improved, especially in complex emotional expression and fuzzy semantic scenarios.

[0022] When the confidence of the first sentiment classification result is insufficient, the associated multimedia data is automatically extracted, and the sentiment tendency is re-evaluated through dynamic fusion of multi-modal information, so as to avoid classification deviation caused by insufficient or incorrect single modal information, and ensure that the classification result is more reliable.

[0023] In combination with specific semantics and emotional characteristics in the field of tourism and travel (such as "crowded", "beautiful", etc.), a sentiment classification model is constructed, so that it is more suitable for the characteristics of tourism and travel materials and meets the actual needs of the industry.

[0024] The text features, classification results and related multimedia materials are structurally associated and efficiently stored, which not only facilitates subsequent data query and analysis, but also provides reliable basic data support for intelligent applications such as sentiment trend analysis, personalized recommendation and tourist experience optimization.

[0025] In summary, the present application breaks through the limitations of traditional sentiment classification methods in single modal analysis, semantic fuzzy processing and multi-modal association management, has significant technical progress and industry application value, and provides comprehensive technical support for the development of the tourism and travel industry. BRIEF DESCRIPTION OF DRAWINGS

[0026] In order to more clearly illustrate the technical solutions in the embodiments or the prior art, the drawings needed in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0027] Figure 1 is a flowchart of the method of the present application. DETAILED DESCRIPTION

[0028] In the following, the preferred description of the application is made in combination with the drawings and the specific embodiments.

[0029] The present embodiment solves the above problems by the following steps:

[0030] In one embodiment, with reference to Figure 1 The present application provides a method for extracting and classifying features of tourism and travel materials, aiming to solve the problem of insufficient precision in feature extraction and sentiment classification of tourism and travel materials in the prior art, especially the demand for correlation analysis and accuracy improvement in sentiment classification of multi-modal tourism and travel materials.

[0031] The tourism and travel materials are received, including text, picture and / or video data. The tourism and travel material receiving step of the present application is mainly optimized for various social media and user-generated content (UGC), and specifically includes the following content:

[0032] The main sources of travel and tourism materials are social media and other user-generated content (UGC) platforms, including but not limited to:

[0033] Social media platforms: such as Weibo, WeChat, Xiaohongshu, Douyin, Kuaishou, etc. Users publish text, pictures, short videos, comments and other travel-related content on these platforms.

[0034] Travel review websites: such as Ma Bensuo, Dazhongdianping, etc. Users share their comments on scenic spots, travel notes and recommendations.

[0035] Video sharing platform: such as B Station, users upload travel records, scenic videos and cultural activity displays.

[0036] Community forums and Q&A platforms: such as Baidu Tieba, Zhihu, users discuss and share their experiences on travel topics.

[0037] Scenic spot official or activity-related content: dynamic comments, forwards, likes and other content generated by user interaction.

[0038] The travel materials include but are not limited to the following types:

[0039] Text content: user-generated dynamic text, comments, tags, Q&A content, and keywords related to travel.

[0040] Picture content: photos of scenic spots, food pictures, architectural scenery and activity photos taken and uploaded by users.

[0041] Video content: travel short videos, live broadcast replays, scenery records, activity clips uploaded by users.

[0042] Metadata: publishing time, geographic location (such as GPS positioning), user-labeled topic tags (such as #scenic beauty#) related to travel materials.

[0043] For social media and UGC materials, the following receiving strategies are adopted:

[0044] Real-time acquisition:

[0045] Use social media open API (such as Weibo API, Douyin open platform API) to receive user-generated travel materials;

[0046] Through the message push mechanism to realize real-time data collection and update, ensure the timeliness of the materials.

[0047] Crawling mechanism:

[0048] For platforms without open API, use web crawler technology to perform targeted crawling based on keywords, geographic location or topic tags to obtain user-generated materials.

[0049] Batch Download:

[0050] Batch download historical content, such as collecting high-heat travel-related materials within a certain time range in the past.

[0051] Authorized Cooperation:

[0052] Cooperate with tourism platforms, scenic spots, or content creators to directly receive user-uploaded original content.

[0053] For the received social media and UGC materials, the following processing is performed:

[0054] Content Screening:

[0055] Filter content unrelated to travel themes by keywords, hashtags, or geographic locations;

[0056] Remove low-quality materials, such as low-resolution images or invalid videos.

[0057] Format Standardization:

[0058] Convert various data formats (such as JPEG, MP4, TXT, etc.) to a unified standard format supported by the system.

[0059] Privacy Processing:

[0060] Anonymize user names and hide sensitive geographic information for data involving user privacy.

[0061] Text segmentation, sentiment word recognition, and syntax analysis are performed on the text to generate first text features and determine the first sentiment classification of the first text features. Specifically, the text processing in this step includes the following content:

[0062] Text Segmentation

[0063] Text segmentation is performed on the received text data to split continuous natural language sentences into several independent word units. Specifically, it includes:

[0064] Based on a predefined segmentation dictionary, the text is segmented to identify basic words, phrases, or professional terms;

[0065] Word classification is performed on the segmentation results through part-of-speech tagging (such as nouns, verbs, adjectives, etc.);

[0066] Specific annotations are performed on domain-specific words in the travel field (such as scenic spot names, landmark buildings, and activity types) to ensure domain adaptability of the segmentation results.

[0067] Sentiment Word Recognition

[0068] On the basis of the word segmentation result, the sentiment dictionary or sentiment classification model is used to identify the sentiment expression contained in the text, specifically including:

[0069] The words or phrases related to the travel and tourism sentiment are extracted, such as "beautiful", "comfortable", "many people", "crowded", etc.

[0070] The sentiment words are classified as positive or negative, and the sentiment intensity value is attached.

[0071] The sentiment tendency of the sentiment words is corrected using negative words and transitional words (such as "not" and "but").

[0072] Syntactic analysis

[0073] The syntax structure of the text is analyzed to extract deeper semantic features, specifically including:

[0074] Using dependency syntax analysis method, the subject-predicate-object, modifier and other syntax relationship trees in the sentence are constructed;

[0075] The logical relationship of multiple sentiment expressions in complex sentences is identified, such as transition, cause and effect, progression, etc.

[0076] According to the syntax structure of the sentence where the sentiment word is located, the semantic weight of the sentiment word in the sentence is determined.

[0077] Generating the first text feature

[0078] The results of word segmentation, sentiment word recognition and syntactic analysis are integrated to generate the first text feature of the text; specifically including:

[0079] Constructing a text sentiment vector, which contains sentiment word categories, intensity values and their syntax positions in the sentence;

[0080] Extracting the domain relevance of sentiment features, such as whether the sentiment words are related to the travel and tourism theme;

[0081] Generating the context features of the text, including the sentiment tendency distribution and its trend in the whole text.

[0082] First sentiment classification

[0083] Using the existing sentiment classification model to classify the first text feature to determine the first sentiment classification of the text; specifically including:

[0084] Based on the supervised learning model (such as BERT, RoBERTa), the first text feature is classified to get the overall sentiment classification result of the text, and the classification categories include but are not limited to "positive", "negative", "neutral";

[0085] The reliability of the classification result is determined by the output confidence of the sentiment classification model.

[0086] For the text containing multiple emotions, the weighted result of the overall sentiment classification of the text is calculated, and a first sentiment classification is generated.

[0087] Result output

[0088] The first text feature of the text is associated with the corresponding first sentiment classification result, and is temporarily stored in the system for subsequent multi-modal fusion analysis.

[0089] Through the above steps, the present application can efficiently extract emotional features from text data and complete preliminary classification, providing accurate and context-related text features for travel material sentiment analysis.

[0090] If the first sentiment classification confidence is less than the first threshold, the picture and / or video data associated with the text are extracted, which includes the following contents:

[0091] The confidence value of the first sentiment classification result is obtained:

[0092] The confidence value is the output probability of the sentiment classification model (such as BERT or RoBERTa) based on the sentiment classification of the text data, indicating the reliability of the classification result.

[0093] For example, the probability of the model output text being "positive" emotion is 0.65, and the confidence value is 0.65.

[0094] Determine whether the confidence is lower than the first preset threshold:

[0095] The first threshold is a system predefined reliability threshold for distinguishing high confidence and low confidence classification results.

[0096] When the confidence of the first sentiment classification is lower than the first threshold (such as 0.7), it is determined that the classification result is unreliable, and further extraction of the picture and / or video data associated with the text is required.

[0097] Determine the picture and / or video data associated with the text:

[0098] Determine the association by checking the metadata (such as timestamp, upload user, geographic location information, etc.) of the text, picture and / or video;

[0099] For example, the text and picture with the same upload time and user can be considered as associated materials.

[0100] Use a pre-trained language model (such as CLIP) to calculate the semantic similarity of the text and picture / video description;

[0101] For example, the text description is "beautiful scenery, but too many people", and the associated picture contains a large number of people, which can be determined to be associated with the text.

[0102] When both picture and video data exist, preferentially extract data more relevant to the text semantics;

[0103] If there is too much multimedia data, sort by upload time and extract the picture and / or video closest to the text generation time.

[0104] According to the matching result of the text and the picture metadata, locate the picture associated with the text from the data repository.

[0105] Format normalization processing is performed on the extracted picture data, such as resolution adjustment or denoising.

[0106] Confirm the video material associated with the text:

[0107] According to the matching result of the text and the video metadata, locate the video file associated with the text from the data repository;

[0108] Transcode the extracted video data to standardize the video format.

[0109] Video key frame extraction:

[0110] Use a time series analysis algorithm to extract key frames from the video, ensuring that the extracted frames have the highest semantic relevance;

[0111] Apply the same object detection algorithm to the key frames as the picture to extract visual features that can be used for sentiment classification.

[0112] Audio feature extraction:

[0113] If the video contains an audio track, extract the audio signal and identify features such as noisy, cheerful, or natural sound background through a sentiment audio classification algorithm.

[0114] If the text is associated with multiple pictures and multiple video files, sort them by relevance and preferentially extract the multimedia data that best fits the text content;

[0115] Temporarily store the extracted multimedia data and associate it with the text data for subsequent multi-modal sentiment classification analysis.

[0116] Through the above steps, when the first sentiment classification result of the text is unreliable, the invention can efficiently extract the picture and / or video data associated with the text, providing rich visual and auditory information support for subsequent multi-modal fusion, thereby significantly improving the accuracy and robustness of sentiment classification.

[0117] Target detection is performed on the text-associated picture and / or video data to generate first multimedia features. The target detection on the text-associated picture and / or video data and the generation of the first multimedia features in this step specifically include the following:

[0118] Target detection of picture data:

[0119] Apply a target detection model (such as YOLO, Faster R-CNN, DETR, etc.) to the associated picture data to identify the main objects in the picture and extract object categories, positions, and size information.

[0120] The detection results include bounding boxes, category labels, and confidence scores, for example:

[0121] The main objects contained in the picture may include crowds, natural landscapes (such as mountains, water, and plants), buildings (such as historical sites and landmark buildings), etc.

[0122] Extract the feature vector of the picture from the target detection results as part of the first multimedia features, including:

[0123] Object category distribution: for example, crowds account for 40% of the image, natural landscapes account for 50%, and buildings account for 10%;

[0124] Scene features: supplement the recognition of overall environmental features (such as "crowded", "open", "natural") through scene classification technology (such as ResNet or Vision Transformer);

[0125] Visual emotional features: preliminary assessment of the emotional tendency of the picture by combining image color, brightness, and texture information, for example, warm tones may convey positive emotions.

[0126] Target detection of video data

[0127] Apply a key frame extraction algorithm (such as a time segmentation algorithm based on SHOT detection) to the video data to extract representative frames;

[0128] Each key frame is treated as a static picture and undergoes the same target detection and feature extraction as a picture.

[0129] Combine the temporal characteristics of video frames to analyze the frequency and distribution of object categories appearing in the video, for example:

[0130] The detected "crowd" exists in 50% of the frames, and "natural landscape" appears in 80% of the frames.

[0131] Extract dynamic features of the video, such as object motion trajectories, change trends, etc.

[0132] If the video contains audio signals, extract audio features such as volume, spectral features to judge the nature of ambient sound (e.g. noisy, natural sound).

[0133] Integrate the target detection results of pictures and videos to generate a unified first multimedia feature, including the following:

[0134] Object class statistics: Describe the types and proportions of all detected objects in the picture and video;

[0135] Spatial information: Describe the location distribution of target objects in the picture or video frame;

[0136] Temporal sequence characteristics (video-specific): Describe the frequency and dynamic changes of target objects.

[0137] Example

[0138] Input

[0139] Text content: "The scenery is beautiful, but there are too many people."

[0140] Associated picture: A photo containing crowds and natural scenery.

[0141] Associated video: A video clip containing mountain scenery, running water sound, and background human voices.

[0142] Processing steps

[0143] Picture target detection:

[0144] The detection model identifies the following objects:

[0145] Crowd: Occupies 40% of the picture, confidence 0.95;

[0146] Mountain: Occupies 30% of the picture, confidence 0.85;

[0147] Trees: Occupies 20% of the picture, confidence 0.90.

[0148] Generate multimedia features for the picture:

[0149] F image = {Crowd: 40%, Mountain: 30%, Trees: 20%}

[0150] Video target detection:

[0151] Extract key frames and detect the following objects in the frames:

[0152] Frame 1: Mountain (confidence 0.9), running water (confidence 0.8);

[0153] Frame 2: Mountain (confidence 0.85), crowd (confidence 0.88).

[0154] Dynamic feature analysis:

[0155] People appear in 50% of frames, mountains and running water appear in 100% of frames.

[0156] Audio analysis results: background noise (confidence 0.7) and running water sound (confidence 0.8) are detected in the audio.

[0157] Multimedia features of the generated video:

[0158] F video ={mountains: 100%, people: 50%, running water: 100%}

[0159] Integrated multimedia features:

[0160] Picture and video feature fusion:

[0161] F multimedia ={people: 45%, mountains: 65%, running water: 50%, trees: 10%}

[0162] Through the above steps, the present application can efficiently analyze text-related picture and / or video data, extract meaningful target detection features, and provide a solid data foundation for subsequent multi-modal sentiment classification.

[0163] Determine the second sentiment classification of the first text feature according to the first text feature, the first multimedia feature, and the first sentiment classification.

[0164] In this step, a multi-modal adaptive sentiment classification network is constructed to comprehensively process text features, picture and video features, and preliminary classification results, and generate more accurate second sentiment classification results. Specifically:

[0165] Text features are semantic embedding vectors extracted from word segmentation, sentiment recognition, and syntax analysis, denoted as F text ; visual features (such as target detection results, scene classification results) and audio features extracted from pictures or videos are integrated into a multi-modal feature vector F media .

[0166] Preliminary sentiment classification results are generated by a text classification model, including categories C first and their confidence P first .

[0167] Use a multi-head attention mechanism to assign dynamic weights to each modality feature, and the weights are calculated based on the correlation between modalities.

[0168]

[0169] where αtext denotes the weight of text features, a media denotes the weight of multimedia features, Q denotes the query vector of text and multimedia features, K denotes the key vector of text and multimedia features; T denotes the transpose operation; d k is the scaling factor of the feature vector.

[0170] The role of the query vector and the key vector is to measure the relevance between different modal features, so as to realize the dynamic weighted fusion of cross-modal features.

[0171] Query vector

[0172] Text modality: The query vector represents what information is expected to be "obtained" from the multimedia modality in the text features. For example, the sentiment expressed in the text (such as "crowded" and "beautiful scenery") can request visual or audio information related to it through the query vector from the multimedia modality.

[0173] Multimedia modality: The query vector represents the semantic content expected to be "sought" from the text modality in the multimedia features. For example, the "crowded scene" in the multimedia features can request a text sentiment description related to it through the query vector.

[0174] Key vector

[0175] Text modality: The key vector represents the semantic features of the text, which is used to provide the overall semantic content of the current text (such as "beautiful scenery but crowded").

[0176] Multimedia modality: The key vector represents the characteristic information of the multimedia modality (such as object class "crowd", "natural scenery" and its distribution ratio), which is used to respond to the query request of the text modality.

[0177] Specific query vectors and key vectors can be obtained by multiplying text features, multimedia features and their corresponding learnable weight matrix of query and key.

[0178] Fusion feature calculation:

[0179] F fused = a text · F text + a media · F media

[0180] Project the text and multimedia features into a shared semantic space through a cross-modal alignment network (such as CLIP or a custom Transformer model):

[0181] F aligned = AlignmentNetwork(F fused )

[0182] Where F aligned represents the aligned vector, AlignmentNetwork() represents the cross-modal alignment processing, which can be processed using existing CLIP or custom Transformer model.

[0183] The aligned features are classified using deep neural networks such as Transformer or stacked LSTM networks.

[0184] Model structure:

[0185] Input: Aligned feature vector F aligned ;

[0186] Intermediate layer: Use self-attention mechanism to capture the interaction between text and multimedia features;

[0187] Output layer: Classification probability vector P second , corresponding to sentiment categories (such as positive, negative, neutral).

[0188] P second = Softmax(W·F aligned +b)

[0189] Where W is the weight matrix and b is the bias term.

[0190] Combine the first sentiment classification confidence P first to modify P second :

[0191] P final = λ·P second +(1-λ)·P first

[0192] Where λ is the weight adjustment factor, dynamically adjusted according to the credibility of P first .

[0193] Exemplarily

[0194] Input data

[0195] Text content: "There are too many people, but the scenery is still very beautiful."

[0196] F text =[0.5,-0.6,0.3], representing the positive, negative, and neutral weights respectively;

[0197] First sentiment classification result: C text ={negative}, P first =0.6.

[0198] Picture features:

[0199] Detected objects: 50% human, 50% natural scenery.

[0200] Feature vector: F media = [0.5, 0.5].

[0201] Video features:

[0202] Keyframe detection: Humans appear in 40% of frames, natural scenery in 80% of frames.

[0203] Audio features: Background sound is flowing water, confidence 0.8.

[0204] Dynamic weights:

[0205] α text = 0.4, α media = 0.6

[0206] Fused features:

[0207] F fused = 0.4·[0.5, -0.6, 0.3] + 0.6·[0.5, 0.5]

[0208] F fused = [0.5, 0.06, 0.38]

[0209] Aligned features:

[0210] F aligned = AlignmentNetwork([0.5, 0.06, 0.38])

[0211] Input F aligned to classification network:

[0212] P second = [positive: 0.7, negative: 0.3, neutral: 0.2]

[0213] Classification result: C second = positive, confidence 0.7.

[0214] Final sentiment classification: positive.

[0215] Explanation: Despite the presence of negative expressions ("too many people") in the text, the final sentiment leans towards positive due to the beautiful natural scenery features in the pictures and videos.

[0216] In this step, by dynamically adjusting the weights of text and multimedia, the sentiment classification effect in complex scenarios is enhanced. Using deep learning models to achieve semantic mapping between modalities improves the accuracy of feature fusion. Combined with the preliminary classification results, the impact of incorrect classification on the final result is reduced.

[0217] This method can provide more accurate sentiment classification results under complex emotional expression and multi-modal association, laying a solid foundation for sentiment analysis of travel materials.

[0218] Associating and saving the first text features and the second emotion classification. This step aims to associate the processed and classified text features with the final generated emotion classification results, and structure these data for subsequent analysis and retrieval. Specifically, it includes the following:

[0219] The first text features and the second emotion classification results are stored as a unified record, mainly including the following data items:

[0220] Text ID (Text ID): Used to mark the unique identifier of each text (such as an incremental ID or hash value).

[0221] First text features (Feature Vector): Text feature vector, recording the text embedding information extracted from sentiment analysis.

[0222] Second emotion classification (Emotion Classification): The final emotion classification result of the text, including:

[0223] Classification label (Label): Such as positive, negative or neutral;

[0224] Classification confidence (Confidence): The probability value of the classification result, used to evaluate the reliability of the result.

[0225] Related information (Related Media): If any, record the identifiers of the pictures or video files associated with the text.

[0226] For easy retrieval, use a distributed system (such as MongoDB, Elasticsearch) to store records in JSON format, exemplarily:

[0227] {

[0228] "text_id":"001",

[0229] "text_feature_vector":[0.5,-0.8,0.3],

[0230] "emotion_label":"Positive",

[0231] "confidence":0.85,

[0232] "related_media":["IMG_001"]

[0233] }

[0234] The text feature vector of each segment (first text feature) is associated with the corresponding second sentiment classification result, using the text unique identifier Text ID as the association key, ensuring one-to-one correspondence.

[0235] If the text processing involves picture and / or video auxiliary classification results, the multimedia data identifier associated with the text also needs to be recorded. For example:

[0236] Picture file name (such as "IMG_001");

[0237] Video clip number or key frame index.

[0238] Each record adds a processing timestamp field (such as "2024-11-15T12:30:00Z") for subsequent audit, traceability or classification trend analysis.

[0239] In another aspect, the present application also provides a travel material feature extraction and classification system, comprising:

[0240] A first collection module for receiving travel materials, including text, picture and / or video data;

[0241] A first classification module for performing word segmentation, sentiment word recognition and syntax analysis on the text to generate first text features and determine the first sentiment classification of the first text features;

[0242] A second collection module for extracting picture and / or video data associated with the text if the first sentiment classification confidence is less than a first threshold;

[0243] An extraction module for performing target detection on the picture and / or video data associated with the text to generate first multimedia features;

[0244] A second classification module for determining the second sentiment classification of the first text features based on the first text features, the first multimedia features and the first sentiment classification;

[0245] A storage module for associating and saving the first text features and the second sentiment classification.

[0246] The part of the module structure of the present application is not particularly clear, and the content recorded in the prior art is used as the criterion. The prior art mentioned in the foregoing background section and the specific embodiment section of the present application can be used as a part of the present application to understand the meaning of some technical features or parameters.

Claims

1. A method for extracting and classifying travel materials, characterized in that, The method comprises the following steps: Receiving travel materials, including text, pictures and / or video data; Performing word segmentation, sentiment word recognition and syntax analysis on the text to generate first text features and determine a first sentiment classification of the first text features; If the first sentiment classification confidence is less than a first threshold, extracting picture and / or video data associated with the text; performing target detection on the picture and / or video data associated with the text to generate first multimedia features; Determining a second sentiment classification of the first text features according to the first text features, the first multimedia features and the first sentiment classification; Associating and saving the first text features and the second sentiment classification.

2. The method of claim 1, wherein, The travel materials come from social media and / or user-generated content.

3. The method of claim 1, wherein, The extraction of the picture and / or video data associated with the text includes determining the association relationship by checking the timestamps or upload user or geographical location information of the text, pictures and / or videos.

4. The method of claim 1, wherein, The target detection on the picture and / or video data associated with the text to generate first multimedia features includes: Performing target detection on the picture to obtain a feature vector of the picture; Performing target detection on each frame of the video to obtain a feature vector of the video; Fusing the feature vector of the picture and the feature vector of the video to obtain the first multimedia features.

5. The method of claim 1, wherein, The determination of the second sentiment classification of the first text features according to the first text features, the first multimedia features and the first sentiment classification includes: The first text feature is denoted as F text The first multimedia feature is denoted as F media ; The first sentiment classification comprises a class C first and a confidence P first ; Calculating the weight of the text features and the weight of the multimedia features wherein α text represents the weight of the text feature, α media represents the weight of the multimedia feature, Q represents a query vector representing the text and multimedia features, K represents a key vector of the text and multimedia features; T represents a transposition operation; d k is a scaling factor of the feature vector, and Softmax() is a softmax function. Performing fusion feature calculation: F fused = a text · F text + a media · F media Projecting the text and multimedia features into a shared semantic space through a cross-modal alignment network: F aligned = AlignmentNetwork(F fused ) where F aligned represents the aligned vector, AlignmentNetwork() represents the cross-modal alignment processing; Classifying the aligned features using a deep neural network Input: aligned feature vectors F aligned ; Intermediate layer: using self-attention mechanism to capture the interaction relationship between text and multimedia features; Output layer: classification probability vector P second corresponding sentiment class; P second = Softmax(W · F aligned + b) Where W is the weight matrix, b is the bias term, and Softmax() is the softmax function. combining the first sentiment classification confidence P first to P second to obtain a final classification confidence P final : P final = λ · P second + (1 - λ) · P first where λ is a weight adjustment factor, according to P first the credibility is dynamically adjusted.

6. A system for feature extraction and classification of travel materials, characterized in that, The method comprises the following modules: A first collection module for receiving travel materials, including text, pictures and / or video data; A first classification module for performing word segmentation, sentiment word recognition and syntax analysis on the text to generate first text features and determine a first sentiment classification of the first text features; A second collection module for extracting picture and / or video data associated with the text if the first sentiment classification confidence is less than a first threshold; An extraction module for performing target detection on the picture and / or video data associated with the text to generate first multimedia features; A second classification module for determining a second sentiment classification of the first text features according to the first text features, the first multimedia features and the first sentiment classification; A storage module for associating and saving the first text features and the second sentiment classification.

7. The system according to claim 6, characterized in that, The travel materials come from social media and / or user-generated content.

8. The system according to claim 6, characterized in that, The extraction of the picture and / or video data associated with the text includes determining the association relationship by checking the timestamps or upload user or geographical location information of the text, pictures and / or videos.

9. The system according to claim 6, wherein, The target detection is performed on the picture and / or video data associated with the text to generate a first multimedia feature, including: performing target detection on the picture to obtain a feature vector of the picture; performing target detection on each frame of the video to obtain a feature vector of the video; fusing the feature vector of the picture and the feature vector of the video to obtain the first multimedia feature.

10. The system according to claim 6, wherein, The second emotion classification of the first text feature is determined according to the first text feature, the first multimedia feature and the first emotion classification, including: The first text feature is denoted as F text The first multimedia feature is denoted as F media ; The first sentiment classification comprises a class C first and a confidence P first ; calculating the weight of the text feature and the weight of the multimedia feature wherein α text represents the weight of the text feature, α media represents the weight of the multimedia feature, Q represents a query vector representing the text and multimedia features, K represents a key vector of the text and multimedia features; T represents a transposition operation; d k is a scaling factor of the feature vector, and Softmax() is a softmax function. performing fusion feature calculation: F fused = a text · F text + a media · F media projecting the text and multimedia features into a shared semantic space through a cross-modal alignment network: F aligned = AlignmentNetwork(F fused ) where F aligned denotes the aligned vector, AlignmentNetwork() denotes the cross-modal alignment process; classifying the aligned features using a deep neural network Input: aligned feature vectors F aligned ; Intermediate layer: use self-attention mechanism to capture the interaction between text and multimedia features; Output layer: classification probability vector P second corresponding sentiment class; P second = Softmax(W · F aligned + b) where W is a weight matrix, b is a bias term, and Softmax() is a softmax function. combining the first sentiment classification confidence P first to P second to obtain a final classification confidence P final : P final = λ · P second + (1 - λ) · P first where λ is a weight adjustment factor, according to P first the credibility is dynamically adjusted.