A Quantification Method of User Travel Intention Based on the Presentation Style of Short Video Visual Content

Through machine learning image classification model, the visual scenes of portraits and scenery in short videos are analyzed, and the impact of visual presentation style on users' travel intentions is quantified, which solves the problem of neglecting the impact of visual content in the existing technology, and achieves more precise quantification of travel intentions and optimization of marketing strategies.

CN119784414BActive Publication Date: 2025-05-27NANJING UNIV OF INFORMATION SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510279005.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-05-27
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

When analyzing the impact of visual content on travel intention in short videos, most studies ignore the impact of visual content and rely on subjective audience feedback, which has deviations.

Method used

The machine learning image classification model is used to analyze two key visual scenes: portraits and scenery. Through steps such as data collection, image preprocessing, keyframe extraction and labeling, travel intention coding, model training and evaluation, image classification and style judgment, and impact quantitative analysis, the impact of short video visual presentation style on users' travel intentions is quantified.

Benefits of technology

The objective evaluation of short video visual content is achieved, the objectivity and accuracy of data processing is improved, and more accurate methods for quantifying users' travel intentions is provided, helping tourism marketers design more attractive marketing content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119784414B_ABST
    Figure CN119784414B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for quantifying the tourism willingness of users based on the visual content presentation style of short videos. By collecting tourism-related short videos, key frames are extracted through the inter-frame difference method and data annotation is performed. With the powerful feature extraction ability of the pre-trained model Inception-v3, the model is optimized by removing the top layer, adding a global average pooling layer, a custom fully connected layer and an output layer, and in-depth analysis of the image content is carried out to identify the visual presentation style of the short video, including portraits, landscapes, etc. The method of the present invention can detect and quantify the content of users' tourism willingness by analyzing the potential impact of the visual style on users' tourism willingness, accurately identify whether users have the tourism willingness, and further evaluate the intensity of intention, helping content creators and platforms to optimize video content and improve the marketing effect and user engagement of tourism-related content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of image classification and short video analysis, and particularly to a method for quantifying users' travel willingness based on the visual content presentation style of short videos. Background Art

[0002] With the advent of the digital age, the fields of cultural entertainment and tourism consumption have undergone major changes. In particular, short video platforms such as TikTok have become important tools for promoting tourist destinations and building tourism brands. Short videos, with their low cost and rapid dissemination characteristics, have shown remarkable effects in attracting audiences and enhancing travel willingness.

[0003] In addition, combined with new media forms such as virtual reality (VR), short videos provide users with immersive travel experiences, thus changing the traditional tourism marketing model to some extent.

[0004] In the prior art, although existing research has confirmed the effectiveness of short videos in tourism marketing, most research focuses on analyzing text content and audience interaction, such as comments and online word-of-mouth, and often ignores the impact of visual content on travel willingness. Although some research focuses on shooting perspectives and visual factors that trigger travel inspiration, these studies mainly rely on subjective audience feedback and may therefore have certain biases.

[0005] Therefore, it is necessary to further conduct in-depth analysis of the influence of short video visual elements, provide a more accurate and objective analysis framework, quantify users' travel willingness, and objectively evaluate how different visual presentation styles actually affect the travel willingness of audiences. Summary of the Invention

[0006] The problem to be solved by the present invention is: to provide a method for quantifying users' travel willingness based on the visual content presentation style of short videos, using a machine learning image classification model to analyze two key visual scenes of portraits and landscapes, and objectively evaluate how the visual content in short videos affects the travel willingness of audiences.

[0007] The present invention adopts the following technical solutions: A method for quantifying users' travel willingness based on the visual content presentation style of short videos, comprising the following steps:

[0008] Step S1, data collection and image preprocessing: Collect short video samples of tourism and perform image preprocessing to construct a sample data set;

[0009] Step S2, key frame extraction and annotation: For the short video data in the sample data set, extract local maxima through the inter-frame difference method as the key frames of each short video, and randomly select part of the data from the key frames for annotation;

[0010] Step S3, tourism intention coding: based on the comment information of short videos in the sample data set, randomly select some data, classify tourism intention through the three-level coding method, and count the number of tourism intention comments of all samples in the sample data set through the keyword matching method;

[0011] Step S4, Inception-v3 model training and evaluation: Divide the preprocessed image data into a training set, a validation set, and a test set, adjust the pretrained Inception-v3 model, remove the top layer, add a global average pooling layer, a custom fully connected layer, and an output layer, use the validation set to evaluate the performance during training, adjust the Inception-v3 model parameters according to the performance chart, and verify the classification accuracy and generalization ability of the Inception-v3 model through the test set;

[0012] Step S5, image classification and style judgment: input the image to be classified into the Inception-v3 model trained in step S4, automatically extract image features and classify them, and output a binary result for each classification, which respectively indicates that the image does not contain or contains a specific target element. Based on the classification result of the Inception-v3 model, the visual presentation style of the image is judged;

[0013] Step S6, quantitative analysis of impact: Based on the image classification results obtained in step S5 and the number of travel intention comments in the video obtained in step S3, a multivariate linear regression model is constructed to quantitatively analyze the visual presentation style of the short video and its impact on the user's travel intention.

[0014] Preferably, in step S1, the travel short video samples include: comments, number of likes, number of favorites, number of shares, number of fans, video length and release time information; the image preprocessing includes: deleting duplicate data, correcting missing values, and removing text and image videos.

[0015] Preferably, in step S2, key frame extraction and labeling includes the following sub-steps:

[0016] S2.1, split the short video data in the sample data set into frame sequences frame by frame, calculate the pixel difference between each frame and the previous frame, and identify the significant changes between frames;

[0017] S2.2, using the moving average method to smooth the difference between frames, using the local maximum method to detect key change points in the smoothed difference data, and extracting the key frames in the video;

[0018] S2.3. Randomly select a part of the data from the key frames according to a preset ratio, perform manual annotation to form a standardized data set, load the key frame data in the standardized data set and perform image data processing, including: color correction, size adjustment, contrast enhancement, and set the image resolution to a unified size, maintain the color format, and scale the image pixel values to between 0 and 1.

[0019] Preferably, the three-level coding method described in step S3 includes: open coding, axial coding, and selective coding; the classification of travel willingness includes: expressing interest and plans, asking for travel destination information, and interactive behaviors.

[0020] Preferably, in step S3, the travel willingness coding includes the following sub-steps:

[0021] S3.1. Based on the comment information of the short videos in the labeled sample data set, randomly select a part of the data, filter out the comments that reflect travel willingness, and perform preliminary coding.

[0022] S3.2. Classify the content of travel willingness comments through the three-level coding method:

[0023] S3.2.1. In the open coding stage, use software to read the original comment information word by word, conceptualize to obtain initial concepts, use the double coding method to perform annotation and coding respectively, compare and unify the coding results, and then classify the initial concepts;

[0024] S3.2.2. In the axial coding stage, logically classify and merge the categories with similar and related meanings to form super categories;

[0025] S3.2.3. In the selective coding stage, converge the super categories formed in the axial coding stage into core categories, including: strong willingness, exploration interest, specific plans, action preparation, play activities, location query, contact information, cost information, travel preparation, transportation route, looking for companions, and participation in social activities;

[0026] S3.2.4. Summarize the core of the user's travel willingness: integrate the comments in the short video comment area and classify the short video comments containing travel willingness.

[0027] S3.3. Statistically calculate the number of travel willingness comments of all samples in the sample data set through the keyword matching method:

[0028] S3.3.1. Selection of keywords based on coding categories: Based on the classification results of the short video comments containing travel willingness in step S3.2.4, select several keywords representing the core content of each category;

[0029] S3.3.2, Keyword Matching: Conduct keyword matching in all comments of the sample dataset, and mark each comment containing one or more keywords as a comment with specific travel intentions;

[0030] S3.3.3, Statistical Analysis: Conduct statistical analysis on the number of comments with travel intentions in a single short video to quantify the interest distribution of users in different travel destinations, as well as the specific concerns and behavior patterns of users when expressing travel intentions.

[0031] Preferably, in step S3.2.4, classify the short video comments with travel intentions and summarize them into: expressing interest and plans, asking for travel destination information, and interactive behavior types.

[0032] Classify the comments involving users' expression of interest in travel destinations or future travel plans as expressing interest and plans;

[0033] Classify the behavior of seeking information such as the location, cost, and strategy of a specific travel destination as asking for travel destination information;

[0034] Define the behavior of interacting with others such as looking for companions and participating in social activities as interactive behavior types.

[0035] Preferably, in step S4, adjust the pre-trained Inception-v3 model, including the following sub-steps:

[0036] S4.1, Based on the pre-trained Inception-v3 model, load the model without the top layer and set the base layer to be non-trainable;

[0037] S4.2, Add a custom output layer, including: a global average pooling layer and a fully connected layer, as well as a sigmoid activation output layer for binary classification;

[0038] S4.3, Compile using the Adam optimizer and the binary cross-entropy loss function, and use early stopping to prevent overfitting;

[0039] S4.4, Enhance the generalization ability of the Inception-v3 model through data augmentation and class weight adjustment.

[0040] Preferably, for the adjusted Inception-v3 model, the data processing method is as follows:

[0041] Input the feature map data stream into the adjusted Inception-v3 model, starting from an input layer with a size of 299×299×3, applicable to RGB image data;

[0042] The input data stream enters a series of basic layers consisting of convolutional layers, pooling layers, and concatenation layers, which change the data dimension and gradually extract and combine feature maps of different scales.

[0043] The first basic layer contains three convolutional pooling modules. Each module contains convolutional kernels of sizes 1×1, 3×3, and 5×5, as well as a max pooling layer. First, each module uses 1×1 convolution to reduce the number of channels, then applies 3×3 and 5×5 convolutions respectively to expand the receptive field, and at the same time combines the max pooling layer to capture spatial information at different levels. The results are merged in the concatenation layer.

[0044] The second basic layer is used to deepen the network and includes five convolutional pooling modules, with each module adding a 7×7 convolutional kernel to capture more extensive context information.

[0045] The third basic layer refines feature extraction through several convolutional pooling modules.

[0046] After the input data stream passes through the basic layers, it is passed to the global average pooling layer, which compresses the entire spatial dimension of each feature map into a single value. For each feature map, the average value of all pixel values is calculated, and a vector of fixed size is output.

[0047] The vector of fixed size is fed into one or more fully connected layers to learn higher-level feature representations; finally, a binary classification task is performed through an output layer with a sigmoid activation function to produce the final prediction result.

[0048] Preferably, in step S5, the Inception-v3 model classifies based on image features, and each classification outputs a binary result of 0 or 1, indicating that the image does not contain or contains a specific target element respectively.

[0049] The visual presentation styles include: an experience style and a landscape style. If the target element of the image is mainly a portrait close-up, it is classified as the experience style; if the target element of the image is mainly a landscape display, it is classified as the landscape style.

[0050] Preferably, in step S6, the multiple linear regression model is the generalized linear model Poisson regression. The variables of the multiple linear regression model include: the number of user travel intention comments , the number of landscape-based pictures in the video key frames , the number of portrait-based pictures in the video key frames , the video duration is the time interval from video publication to video collection , the number of fans of the publisher , the number of video likes , the number of video comments , the number of video shares , the number of video collections .

[0051] Analyze the relationship between variables through a multiple regression model, and quantify the influence degree of variables on the number of comments on travel willingness through the regression parameters β 0 ~β 8 respectively; determine the setting principle of the regression parameters, statistical significance, and test to judge which variables have a significant impact on the number of comments on travel willingness.

[0052] Specifically, first determine the setting principle of the regression parameters, that is, judge which variables have a significant impact on the number of comments on travel willingness through statistical significance tests. For example, if the regression coefficient of a certain variable is significantly non-zero, it is considered that the variable has a significant impact on the number of comments on travel willingness.

[0053] In this way, the following conclusions are obtained:

[0054] S6.1. The number of landscape pictures and portrait pictures in the video key frames and the number of portrait pictures have a significant impact on the number of comments on travel willingness;

[0055] S6.2. The video duration and the time interval from video release to collection also have a significant impact on the number of comments on travel willingness;

[0056] S6.3. The number of fans of the publisher , the number of video likes , the number of video comments , the number of video shares and the number of video collections also have a significant impact on the number of comments on travel willingness.

[0057] Preferably, before applying the multiple linear regression model, perform logarithmic transformation and standardization on the key variables , , , . Take the natural logarithm after adding 1 to each variable, and then perform centering and scaling to unit variance to reduce the skewness and scaling differences of the key variable data, and enhance the statistical power and interpretability.

[0058] Compared with the prior art, the present invention adopts the above technical solutions and has the following technical effects:

[0059] 1. The user travel willingness quantification method of the present invention combines machine learning image classification technology with econometric regression analysis, enhancing the scientificity and effectiveness of tourism marketing strategies. By using an automated image classification method, it reduces manual intervention, improves the objectivity and accuracy of data processing, and ensures the consistency and repeatability of results. Through the precise analysis of key visual elements - portraits and landscapes in short videos, it objectively evaluates how each element affects the travel willingness of the audience, helping tourism marketers design more attractive marketing content for target audiences more effectively.

[0060] 2. Through the application of the user travel willingness quantification method of the present invention, the attractiveness of tourist destinations can be enhanced. By providing scientific decision-making support, it promotes the sustainable development and innovation of the tourism industry. It can provide a new and more objective means of evaluation and optimization for tourism marketing, and has important value for improving the market performance of tourist destinations and brands. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 is a flowchart of the user travel willingness quantification method based on the visual content presentation style of short videos of the present invention;

[0062] Figure 2 is a schematic diagram of key frames of the user travel willingness quantification method based on the visual content presentation style of short videos of the present invention;

[0063] Figure 3 is a diagram of the three-level coding result provided by an embodiment of the present invention;

[0064] Figure 4 is a structural diagram of the Inception-v3 model provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0065] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the application will be further elaborated in detail below with reference to the accompanying drawings. The described embodiments are only a part of the embodiments involved in the present invention. All non-innovative embodiments made by other researchers in this field based on this embodiment fall within the protection scope of the present invention. At the same time, for the step numbers in the embodiments of the present invention, they are only set for the convenience of explanation and do not limit the order between steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0066] In an embodiment of the present invention, a user travel willingness quantification method based on the visual content presentation style of short videos is used to complete the image classification of short videos and further quantify the travel willingness of users, as Figure 1 shown, and includes the following steps:

[0067] Step S1. Data collection and image preprocessing: Collect short video samples related to tourism, including information such as comments, likes, collections, shares, the number of followers, video duration, and release time, to construct a dataset.

[0068] Specifically, in this embodiment, the purposive sampling technique is adopted, and keywords such as "beach", "island", "Yalong Bay", "Nan'ao Island", and "Dongshan Island" are selected to search for short videos related to seaside scenery categories and coastal city attractions. A total of 1,941 short video data released between January 1 and July 17, 2024 are crawled using a Python program. The collected data includes user names, the number of followers, release time, video duration, video description, number of likes, number of shares, number of comments, number of collections, and video links, etc.

[0069] In view of the research focus on video content analysis, all picture-and-text videos that only contain static pictures are excluded in this embodiment. In addition, strict preprocessing is performed on the collected data, including deleting duplicates and correcting missing values.

[0070] Furthermore, in addition to video data, the comment data (excluding sub-comments) of each video is crawled in batches in this embodiment. A total of 676,500 comments are collected, which ensures data integrity and high accuracy of analysis, and provides a solid data foundation for the subsequent quantification of users' tourism willingness.

[0071] Step S2. Key frame extraction and image preprocessing: Apply the inter-frame difference method to the downloaded short videos to extract local maxima, which represent the key frames of each video. Then, randomly select 15% of the images from these key frames for manual annotation for subsequent analysis.

[0072] First, the downloaded short videos are segmented into frame sequences frame by frame; then, the pixel difference between each frame and its previous frame is calculated to identify significant inter-frame changes.

[0073] To reduce noise interference and improve the accuracy of detection, the moving average method is used to smooth the difference values between frames, as Figure 2 shown. The local maxima method is used to detect key change points in the smoothed difference data, so as to accurately extract the key frames in the video.

[0074] Specifically, in this embodiment, after processing 1,941 videos, a total of 23,693 key frames are extracted, with an average of about 12 key frames extracted per video. The specific quantity is adjusted according to the complexity of the video content and the frequency of scene changes.

[0075] Subsequently, 15% of these key frames are randomly selected for manual annotation to prepare a standardized data set for subsequent video content analysis and machine learning training. The manually annotated key frames will be loaded and subjected to necessary preprocessing operations, including but not limited to: color correction, resizing, and contrast enhancement, to ensure that the data quality meets the requirements of subsequent processing.

[0076] Step S3, travel intention coding: For the video comments collected in step S1, 5% are randomly selected for three-level coding, and the travel intention comments are divided into three categories: expressing interest and plans, asking for travel destination information, and interactive behavior. The number of travel intention comments in all video samples is counted through keyword matching method.

[0077] In the data coding and analysis process of this embodiment, a systematic coding method is adopted to ensure the consistency and reliability of data processing.

[0078] First, 20,000 reviews were randomly selected from the 676,500 crawled reviews for manual annotation, and the reviews reflecting travel intentions were screened and preliminarily coded.

[0079] Then, if Figure 3 As shown in the figure, the content of tourism intention comments is divided into types through a three-level coding method:

[0080] In the open coding stage, NVivo 11.0 software was used to read the comments word by word and repeatedly to conceptualize and categorize the original data to ensure a coding process without subjective bias. To improve the reliability of coding, a double coding procedure was adopted, with two coders annotating the data separately and comparing the coding results. It was found that 95% of the coding results were consistent. For the inconsistent parts, after discussion and consensus, the initial concepts were finally classified into 83 categories.

[0081] Subsequently, the axial coding phase was entered to logically classify and merge categories with similar and related meanings, forming 20 supercategories.

[0082] In the selective coding stage, the supercategories formed during the axial coding process were further focused and aggregated into 12 core categories: strong intention, exploratory interest, specific plan, action preparation, play activities, location inquiry, contact information, cost information, travel preparation, transportation route, finding companions, and participation in social activities.

[0083] On this basis, the comments in the short video comment area were deeply mined and logically integrated. Comments related to users expressing interest in a certain tourist destination or future travel plans were classified as the "expressing interest and plans" category; behaviors of seeking information such as the location, cost, and travel guides of a specific tourist destination were classified as the "inquiring about tourist destination information" category; and behaviors of interacting with others such as looking for companions and participating in social activities were defined as the "interaction behaviors" category.

[0084] Through this process, the short video comments containing travel intentions were classified into three main categories: the "expressing interest and plans" category, the "inquiring about tourist destination information" category, and the "interaction behaviors" category. These three categories not only effectively summarize the core aspects of users' travel intentions but also reflect the internal connections and importance among the core categories.

[0085] Finally, the number of travel intention comments in the sample dataset was counted through the keyword matching method. The specific method is as follows:

[0086] First, keyword selection based on coding categories was carried out.

[0087] In this embodiment, after completing the selective coding, three main categories, namely "expressing interest and plans", "inquiring about tourist destination information", and "interaction behaviors", were determined. For each category, keywords representing its core content were selected respectively.

[0088] For example, keywords such as "want to go", "go go go", and "intend to go" were used to mark the "expressing interest and plans" category; keywords such as "where", "location", and "how to go" were used to mark the "inquiring about tourist destination information" category; and keywords such as "partner" and "check in" were used to mark the "interaction behaviors" category.

[0089] Then, keyword matching was carried out.

[0090] In this embodiment, keyword matching was performed on all comments in the sample dataset, and each comment containing any of the above keywords or a combination of multiple keywords was marked as a comment containing a specific travel intention. This matching method ensures that all relevant comments can be comprehensively and accurately identified.

[0091] Finally, statistical analysis was carried out.

[0092] In this embodiment, the number of travel intention comments in a single short video was statistically analyzed. By quantitatively analyzing the travel intention comments in each short video, not only can the interest distribution of users in different tourist destinations be understood, but also the specific concerns and behavior patterns of users when expressing travel intentions can be further explored.

[0093] Step S4, Training and Evaluation of Inception-v3 Model: Divide the preprocessed image data into a training set (accounting for 64%), a validation set (accounting for 16%), and a test set (accounting for 20%); Adjust the pre-trained Inception-v3 model, remove the top layer, add a global average pooling layer, a custom fully connected layer, and an output layer; Use the validation set to evaluate the performance during training, and adjust the model parameters according to the performance chart to optimize the effect; Finally, verify the classification accuracy and generalization ability of the model through the test set.

[0094] Specifically, in this embodiment, step S4 includes the following sub-steps:

[0095] Step S4.1: For 3169 manually annotated image samples, divide 2028 of them into the training set, 507 into the validation set, and 634 into the test set.

[0096] Step S4.2: As Figure 4 shown, adjust the pre-trained InceptionV3 model to adapt to the image classification task.

[0097] First, load the InceptionV3 model without the top layer and set its base layer to be non-trainable to retain its efficient feature extraction ability. Through this strategy, it is possible to focus on customizing the top of the model without retraining the entire network.

[0098] Specifically, add a global average pooling layer (GlobalAveragePooling2D) on the basis of the original InceptionV3 model architecture to compress the feature map into a vector of a fixed length.

[0099] Subsequently, introduce a fully connected layer (Fully Connected Layer) to further process these feature vectors, and finally implement a binary classification task through an output layer using the sigmoid activation function.

[0100] Next, in the model compilation stage, select the Adam optimizer and the binary cross-entropy loss function to optimize the classification performance. To avoid overfitting problems, this implementation also adopts the early stopping method (Early Stopping).

[0101] In addition, to enhance the generalization ability of the model and solve the problem of data imbalance, this embodiment also introduces data augmentation technology and class weight adjustment strategies. Through these measures, the improved InceptionV3 model can not only effectively extract high-level features from the input image, but also show good performance and adaptability in practical applications.

[0102] In this embodiment, the improved InceptionV3 model starts from an input layer with a size of 299×299×3, which is applicable to RGB images. The subsequent data stream enters a series of complex basic layers (Blocks) composed of convolutional layers, pooling layers, and concatenation layers. The design of the basic layer aims to gradually extract and combine feature maps of different scales.

[0103] For example, Block 1 contains multiple modules. Each module contains convolutional kernels of sizes 1×1, 3×3, and 5×5, as well as a max pooling layer. The results of these operations are merged in the concatenation layer. Such a design allows the model to capture both fine-grained and coarse-grained feature information simultaneously.

[0104] The specific module details are as follows:

[0105] Block 1: This block contains three modules. Each module first uses a 1×1 convolution to reduce the number of channels, then applies 3×3 and 5×5 convolutions respectively to expand the receptive field, and combines a max pooling layer to capture spatial information at different levels.

[0106] Block 2: Further deepen the network, including five modules, adopting a similar but more complex pattern, adding 7×7 convolutional kernels to capture more extensive context information.

[0107] Block 3: The last few modules continue this process, but are simplified in architecture, focusing on refined feature extraction.

[0108] The operations of each layer change the data dimensions. For example, after multiple convolutions and poolings, the original input data gradually shrinks and deepens its feature representation ability.

[0109] When the data flows through the basic layer, it is passed to the global average pooling layer. This step compresses the entire spatial dimension of each feature map into a single value, that is, for each feature map, calculate the average value of all its pixel values.

[0110] In this way, regardless of the size of the input image, the output will be a vector of a fixed size, which is 8×8×2048 in this embodiment. This process not only reduces the number of parameters but also effectively prevents overfitting.

[0111] Next, this 8×8×2048 vector is fed into one or more fully connected layers for further processing to learn higher-level feature representations. Finally, a binary classification task is performed through an output layer with a sigmoid activation function to produce the final prediction result.

[0112] Through the above process, the improved InceptionV3 model can efficiently process input images, gradually extract useful feature information, and finally make accurate classification decisions based on these features. This design not only improves the performance on specific tasks but also enhances the adaptability of the model to various application scenarios.

[0113] Step S4.3: To select an image classification model suitable for Step S4, comprehensive tests and performance comparisons were conducted on mainstream image classification models such as ResNet50, InceptionV3, VGG16, VGG19, and DenseNet121.

[0114] By generating performance charts such as confusion matrices, loss rate curves, precision curves, accuracy curves, and ROC curves, the classification performance of each model was comprehensively evaluated. To prevent overfitting, early stopping, regularization, and data augmentation techniques were uniformly adopted during the training process to improve the generalization ability of the model.

[0115] After comprehensive performance evaluation, this embodiment finally determined the fine-tuned InceptionV3 as the main image classification model. To adapt to the processing requirements of different models, the main adjustment includes the size of the input pictures. The learning rate of all models was uniformly set to 0.0001 to ensure the stability and efficiency of the training process. In addition, all models used a batch size of 32, which is a balanced choice considering memory efficiency and model performance. The optimizer used was Adam, and due to its ability to adapt the learning rate, the training was more stable. The loss function selected was binary_crossentropy, which is suitable for binary classification problems and can effectively handle class outputs.

[0116] Furthermore, in this embodiment, the loss of the validation set was monitored through early stopping (EarlyStopping), and the patience parameter was set to 10, aiming to prevent overfitting and reduce unnecessary training time. Data augmentation includes operations such as random rotation, scaling, and flipping, which enhanced the generalization ability of the model when facing new and unseen samples. Through these meticulous configurations and parameter optimizations, the scientificity and applicability of model selection were ensured, enabling the model to achieve better performance in practical applications.

[0117] Table 1 further shows the detailed measurement results of the performance of different machine learning models. For a given test set, the calculation formulas for precision, recall, and F1-score are:

[0118] ;

[0119] ;

[0120] ;

[0121] Among them, TP, FP and FN represent true positive, false positive and false negative, respectively.

[0122] Table 1: Machine Learning Model Performance

[0123]

[0124] Step S5, image classification and style judgment: Input the image to be classified into the trained Inception-v3 model, and the model automatically extracts image features and classifies them. Each classification outputs a binary result, "0" or "1", indicating that the image does not contain or contains a specific target element (such as scenery or portraits). Based on the model classification results, the visual presentation style of the image is judged.

[0125] In this example, the image classification and style judgment process is implemented by using the trained InceptionV3 model. First, the image to be classified is input into the model, where the model uses its deep learning architecture to automatically extract the key features in the image. Then, the model classifies according to these features, and the output of each classification is a binary result, "0" or "1", which represents whether the image contains a specific target element. For example, "0" may mean that there is no portrait but only scenery, while "1" means that the image contains a portrait.

[0126] This classification mechanism allows the system to determine the visual presentation style of an image in an efficient and automated manner. For example, based on the presence and proportion of portrait and landscape elements in the image, the system can determine whether the image is more inclined to a person-centric style or a landscape-centric style. The method of this embodiment is extremely valuable in a variety of application scenarios, such as automatic content classification, augmented reality applications, or personalized media recommendations.

[0127] Step S6, quantitative analysis of impact: Based on the image classification results and the number of travel intention comments in the video, a multivariate linear regression model is constructed to quantitatively analyze the impact of the visual presentation style and other video features of the short video on the user's travel intention.

[0128] Specifically, this embodiment studies the influence of the visual presentation style of short videos on the user's travel intention by establishing a generalized linear model Poisson regression, which is as follows:

[0129] ;

[0130] ;

[0131] ;

[0132] ;

[0133] In the formula, is the number of user travel intention comments, is the number of landscape-dominated pictures in the key frames of the video, is the number of portrait-dominated pictures in the key frames of the video, is the video duration, is the time interval from video release to video collection, is the number of fans of the publisher, is the number of video likes, is the number of video comments, is the number of video shares, is the number of video collections; β 0 ~β 8 are regression parameters, and ε is a random disturbance term.

[0134] Preferably, before applying the generalized linear model, perform logarithmic transformation and subsequent standardization on , , , and several other key variables, in order to reduce the skewness and scaling differences of the data, and increase its statistical power and explanatory power in the subsequent model. The specific operation is to take the natural logarithm after adding 1 to each variable, and then perform centering and scaling to unit variance, so that the data is closer to a normal distribution, which is particularly important for the application of the regression model.

[0135] Then, apply this model to explore the impact of the visual content presentation style of short videos on users' travel intentions, and the results are shown in Table 2.

[0136] Table 2. Results of the impact of the visual content presentation style of short videos on users' travel intentions

[0137]

[0138] *p < 0.05; **p < 0.01, ***p < 0.001

[0139] Specifically, first, analyze the regression results of Model 1 and Model 2.

[0140] Regression analysis shows that the number of key frames dominated by portraits in the video has a significant negative impact on users' travel intentions. Specifically, as the number of portrait pictures in the video increases, the number of travel intention comments from the audience significantly decreases (regression parameter , , Model 1).

[0141] The regression parameter β represents the degree of influence of the independent variable on the dependent variable. For example, in Model 1, β 1It is shown that for each additional key frame mainly featuring a human figure, the average number of comments on the audience's travel intention decreases by 0.024.

[0142] The p-value is used to measure statistical significance. In this embodiment, p < 0.01 and p < 0.05 respectively indicate that at the 1% and 5% significance levels, the observed results are statistically significant. For the case of p < 0.01, the null hypothesis can be rejected at the 99% confidence level, while for the case of p < 0.05, it is at the 95% confidence level. This means there is sufficient evidence to reject the null hypothesis, indicating that the number of key frames has a significant impact on travel intention.

[0143] This result shows that when the video contains more human-related elements, it may distract the audience's attention from the travel destination or scenic spots, thus reducing their travel interest.

[0144] Contrary to the number of key frames mainly featuring human figures, the number of key frames mainly featuring scenery in the video shows a significant positive impact. The regression results show that as more natural landscapes or beautiful scenic spots are shown in the video, the audience's travel intention increases (regression parameter , , Model 2). This may be because natural landscapes themselves have strong attraction and can stimulate the audience's exploration desire. Especially when it comes to travel destinations, rich scenery pictures can effectively enhance the audience's interest.

[0145] When considering the squared term of the number of key frames mainly featuring human figures , the regression coefficient is negative and significant (regression parameter , , Model 1). This indicates that as the number of human pictures increases, its negative impact on travel intention is increasing. Too many human pictures may cause visual fatigue in the audience, thus reducing their interest in travel content. Therefore, although an appropriate amount of human elements may play a certain role in emotional resonance, when their number is excessive, it will instead weaken the audience's travel intention.

[0146] Similar to the number of key frames mainly featuring human figures, the squared term of the number of key frames mainly featuring scenery ) also shows a significant negative impact (regression parameter , , Model 2). This means that although increasing the number of scenery pictures can improve the audience's travel interest, when the number of scenery pictures is excessive, the audience's interest may start to decline. Over-displaying natural landscapes may cause visual fatigue or information overload, thus weakening its positive impact on travel intention, reminding video producers to control the degree of scenery display.

[0147] Regarding the video duration , the regression coefficients are positive and negative in different models. When the number of key frames mainly featuring people and its squared term are used as independent variables (Model 1), the regression coefficient of the video duration is , and it is significant at . This result indicates that an increase in the video duration will promote an increase in the number of comments on travel willingness. Longer videos can provide more information and details. Especially in videos mainly featuring people, viewers are more likely to have an emotional resonance. If the people in the video show emotions such as joy and happiness, it may infect the viewers and stimulate their travel interest. A longer duration can further strengthen the transmission of this emotion, and viewers are also more likely to generate more travel willingness and leave comments during the viewing process. Therefore, in videos mainly featuring people, a longer duration helps to improve the viewers' engagement and willingness to comment.

[0148] When the number of key frames mainly featuring scenery and its squared term are used as independent variables (Model 2), the regression coefficient of the video duration is , and it is significant at . This result indicates that an increase in the video duration will instead reduce the number of comments on travel willingness. Videos mainly featuring scenery usually focus on showing the beauty and practical value of the scenic spots (for example, whether the scenic spots are worth visiting), but if the shooting quality of the scenery is not high or the scenic spots themselves lack sufficient attraction, longer videos may expose these deficiencies. For scenery videos, viewers usually expect to see more visually impactful and creative content, and a longer duration may make the disadvantages of the scenic spots more obvious, leading to a decline in viewers' interest and thus reducing comments and interactions. Therefore, if the duration of a scenery-based video is too long, it may cause viewers' fatigue and reduce their engagement.

[0149] The time interval from video release to video collection ) also has a significant positive and negative difference in regression coefficients in different models. When the number of key frames mainly featuring people and its squared term are used as independent variables (Model 1), the regression coefficient is , and it is significant at , indicating that a longer time interval has a negative impact on the number of comments on travel willingness. This may be because a longer time interval means that the attention and engagement of viewers may decline after the video is released, resulting in weakened interactivity. Especially when it comes to people's pictures, a long time interval may gradually reduce viewers' interest, thus reducing the number of comments. On the other hand, when the number of key frames mainly featuring scenery and its squared term are used as independent variables (Model 2), the regression coefficient is , and it is significant at It is significant at that time, indicating that a longer time interval is actually positively correlated with the number of comments on travel willingness. This may suggest that a longer time interval can increase the sedimentation effect of video content, enabling the video to regain the attention of viewers after a period of time. Especially when it comes to content related to scenery, the continuous exposure of the video may attract more viewers to participate in the comments.

[0150] In addition to these independent variables and the above two control variables, in the four models, the regression coefficients of the remaining control variables show a consistent significant trend. Among them, the number of video comments , the number of video shares and the number of video collections all have positive and significant regression coefficients.

[0151] The positive relationship between the number of video shares and the number of comments on travel willingness indicates that when viewers actively share a video, they not only to a certain extent approve of the video content but may also influence others in their social circle and network. When the video is widely shared, it will bring higher exposure, attracting more potential viewers to watch and comment on the video. Moreover, the act of sharing itself indicates a relatively high level of interest and engagement of the viewers with the video content, which further motivates them to leave comments. Especially when it comes to travel destinations, the sharing behavior can deepen the viewers' interest in the destination and prompt them to express their personal opinions.

[0152] An increase in the number of video comments is usually a reflection of the interactive activity level and may also drive more viewers to participate in the comments. People often generate a certain sense of social identity or curiosity after seeing other viewers' comments, which in turn stimulates their further interest in the video content. Especially in travel-related videos, viewers share their views, travel experiences or ask relevant questions in the comments. This interaction enhances the social nature and sense of participation of the video content, thereby increasing the willingness of other viewers to comment and forming a positive feedback loop.

[0153] An increase in the number of video collections indicates that viewers have a relatively high level of interest and approval of the video content. The act of collection usually means that viewers hope to review or refer to the video in the future, indicating that they consider the content of the video valuable or relevant to their personal needs. In travel videos, viewers may collect content related to destinations they are interested in or travel suggestions for future reference or travel planning. An increase in the number of video collections not only reflects the long-term interest of viewers in the content but also encourages them to share, discuss with other viewers or actively participate in more interactions in the future, thereby increasing the number of comments on travel willingness.

[0154] And the number of fans of the publisher and the number of video likes The regression coefficients are all negative and significant, indicating that an increase in these factors is negatively correlated with the number of comments on travel willingness. A possible reason is that when a scenic spot is too popular, viewers may hesitate to go there, worried about overcrowding and its impact on the travel experience, which may lead to a decrease in their interest in travel-related content and thus a reduction in the number of comments on travel willingness. In addition, a larger number of followers of the blogger may mean that the video has reached a large audience during the dissemination process, which may also make the atmosphere of short-video travel marketing more intense and further raise the expectations of the audience. For some viewers, overly "commercial" marketing may cause disgust or a "deterrent" effect, resulting in a decrease in their participation and willingness to interact. Therefore, although the number of followers and likes can increase the exposure of the video, an increase in these factors may, under certain conditions, have a negative relationship with the travel willingness of the audience.

[0155] For Models 3 and 4, the main focus is on studying the independent effects of two variables and their non-linear relationships.

[0156] In Model 3, the regression results show that the number of key frames featuring scenery and the number of key frames featuring people both have a significant positive impact on the number of comments on travel willingness of the audience.

[0157] Specifically, the regression coefficient is 0.290, which means that for every additional key frame featuring scenery, the number of comments on travel willingness will increase by approximately 0.290. This indicates that the scenery element has a strong attraction to the travel willingness of the audience and can effectively stimulate the interest and comment desire of the audience. In contrast, the regression coefficient is 0.105, that is, for every additional key frame featuring people, the number of comments on travel willingness increases by 0.105. This shows that although the portrait image also has a certain positive impact on travel willingness, its influence is smaller than that of the scenery.

[0158] In addition, the time interval from video release to video collection the regression coefficient is 0.031, significant at the 1% significance level, indicating that a longer time interval is related to a larger number of comments on travel willingness. This may be because the video continues to be exposed after release, attracting the attention of the audience and stimulating them to leave comments.

[0159] And the regression coefficient of the video duration is It is -0.215, indicating a negative relationship between the longer video duration and the number of comments on travel willingness. This may mean that an overly long video can lead to a decline in the audience's attention and affect their willingness to interact.

[0160] In Model 4, in addition to adding and we also considered and , that is, the square term of the number of key frames was introduced. The regression results show that the introduction of the square term further reveals the non-linear effect of the scenery and portrait elements in the video on the audience's travel willingness.

[0161] Specifically, the regression coefficient of is 0.277, while the regression coefficient of is -0.014, which indicates that the impact of the number of key frames dominated by scenery on travel willingness shows a decreasing trend, that is, when the number of scenery pictures exceeds a certain amount, its positive effect will weaken and may even bring negative effects. The regression coefficient of is 0.117, while the regression coefficient of is -0.015, which also indicates that the impact of the number of key frames dominated by portraits on travel willingness is also non-linear. As the number of portrait pictures increases, the positive effect gradually weakens and may even bring negative effects.

[0162] In summary, according to the regression results of Model 1 and Model 2, the composition of video content, especially the display of portraits and scenery, significantly affects the audience's travel willingness. The impact of video duration is closely related to the nature of the content: in videos mainly featuring people, a longer duration helps enhance emotional resonance and improve travel willingness; while in scenery videos, an overly long duration may expose the drawbacks of the scenic spots and instead reduce the audience's interest. The impact of the time interval between video release and collection on the number of comments varies depending on the content type: the number of comments on videos mainly featuring portraits decreases after a long time interval, while the number of comments on scenery videos increases after a long time interval, indicating that the cumulative effects of video type and release time affect interactivity. Video sharing, comments, and collections reflect the high degree of audience participation in the content and promote the increase in the number of comments on travel willingness. Although the number of fans and likes of the publisher helps with video exposure, they do not determine the success of travel marketing. Therefore, creators should focus on improving the quality of video content. Even starting from scratch, this will be more effective in attracting targeted audiences, increasing interactions and comments, and thus enhancing travel willingness.

[0163] In contrast, the regression results of Model 3 and Model 4 reveal the complexity of video content, especially the diminishing effect of portrait and scenery elements. In Model 3, when the number of key frames dominated by portraits is considered together with the number of key frames dominated by scenery, the positive impact of portrait elements on travel willingness is relatively significant, while the impact of scenery elements remains positive, indicating their relative importance. In Model 4, the squared terms of portraits and scenery are further introduced, and it is found that the impact of the increase in the number of portrait and scenery elements on travel willingness is not linear, but there is a diminishing effect. This shows that during the creation of short videos, moderately increasing portrait and scenery elements helps to increase the audience's interest, but after exceeding a certain amount, it may cause visual fatigue and instead reduce travel willingness. Therefore, appropriately balancing the number of video content elements and controlling the video duration and release time interval can significantly enhance the audience's interactivity and travel willingness.

[0164] In summary, the present invention proposes a method for quantifying users' travel willingness based on the visual content presentation style of short videos, and verifies its effectiveness in short video tourism marketing. In particular, by optimizing the ratio of portraits to scenery in the video, this method can effectively promote users' travel intention, which has important theoretical value and practical application significance.

[0165] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A method for quantifying user travel intention based on the presentation style of short video visual content, characterized in that: The steps include: Step S1, data collection and image preprocessing: collect tourism short video samples and perform image preprocessing to build a sample data set; Step S2, key frame extraction and labeling: for the short video data in the sample data set, the local maximum is extracted by the frame difference method as the key frame of each short video, and some data is randomly selected from the key frame for labeling; Step S3, tourism intention coding: based on the comment information of short videos in the sample data set, randomly select some data, classify tourism intention through the three-level coding method, and count the number of tourism intention comments of all samples in the sample data set through the keyword matching method; Step S4, Inception-v3 model training and evaluation: Divide the preprocessed image data into a training set, a validation set, and a test set, adjust the pretrained Inception-v3 model, remove the top layer, add a global average pooling layer, a custom fully connected layer, and an output layer, use the validation set to evaluate the performance during training, adjust the Inception-v3 model parameters according to the performance chart, and verify the classification accuracy and generalization ability of the Inception-v3 model through the test set; Step S5, image classification and style judgment: input the image to be classified into the Inception-v3 model trained in step S4, automatically extract image features and classify them, and output a binary result for each classification, indicating that the image does not contain or contains the target element. Based on the classification result of the Inception-v3 model, judge the visual presentation style of the image; Step S6, quantitative analysis of impact: Based on the image classification results obtained in step S5 and the number of travel intention comments in the video obtained in step S3, a multivariate linear regression model is constructed to quantitatively analyze the visual presentation style of the short video and its impact on the user's travel intention.

2. The method for quantifying user travel intention based on short video visual content presentation style according to claim 1 is characterized in that: In step S1, the travel short video sample includes: comments, number of likes, number of favorites, number of shares, number of fans, video length and release time information; the image preprocessing includes: deleting duplicate data, correcting missing values, and removing text and image videos.

3. The method for quantifying user travel intention based on short video visual content presentation style according to claim 1 is characterized in that: In step S2, the key frame extraction and labeling includes the following sub-steps: S2.1, split the short video data in the sample data set into frame sequences frame by frame, calculate the pixel difference between each frame and the previous frame, and identify the significant changes between frames; S2.2, using the moving average method to smooth the difference between frames, using the local maximum method to detect key change points in the smoothed difference data, and extracting the key frames in the video; S2.

3. Randomly select part of the data from the key frames according to a preset ratio, perform manual annotation, form a standardized data set, load the key frame data in the standardized data set and perform image data processing, including: color correction, size adjustment, contrast enhancement, and set the image resolution to a uniform size, maintain the color format, and scale the image pixel value to between 0 and 1.

4. The method for quantifying user travel intention based on short video visual content presentation style according to claim 3 is characterized in that: The three-level coding method in step S3 includes: open coding, main axis coding, and selective coding; the tourism intention classification includes: expressing interest and plan, asking for information about tourist destinations, and interactive behavior.

5. The method for quantifying user travel intention based on short video visual content presentation style according to claim 4 is characterized in that: The tourism intention coding in step S3 includes the following sub-steps: S3.

1. Based on the comment information of short videos in the labeled sample data set, randomly select some data, screen out the comments that reflect travel intentions, and perform preliminary coding; S3.

2. Classify the content of travel intention comments by three-level coding method: S3.2.

1. In the open coding stage, the original review information was read word by word using software to conceptualize the initial concepts. The double coding method was used to mark and code them respectively. After comparing and unifying the coding results, the initial concepts were classified. S3.2.

2. In the axial coding stage, categories with similar and related meanings were logically classified and merged to form supercategories; S3.2.

3. In the selective coding stage, the supercategories formed in the main axis coding stage were aggregated into core categories, including: strong intention, exploration interest, specific plan, action preparation, play activities, location inquiry, contact information, cost information, travel preparation, transportation route, finding companions and social activity participation; S3.2.

4. Summarize the core of users’ travel intentions: Integrate the comments in the short video comment area and classify the short video comments containing travel intentions; S3.

3. Count the number of travel intention comments of all samples in the sample data set by keyword matching method: S3.3.

1. Keyword selection based on coding categories: Based on the classification results of short video comments containing travel intentions in step S3.2.4, select several keywords representing the core content of each category; S3.3.2, Keyword matching: Keyword matching is performed in all comments in the sample data set, and each comment containing one or more keywords is marked as a comment containing travel intention; S3.3.

3. Statistical analysis: Conduct statistical analysis on the number of comments containing travel intentions in a single short video, quantify the distribution of users’ interests in different tourist destinations, and the specific concerns and behavior patterns of users when expressing their travel intentions.

6. The method for quantifying user travel intention based on short video visual content presentation style according to claim 5 is characterized in that: In step S3.2.4, the short video comments containing travel intention are classified into: expressing interest and plan, asking for travel destination information, and interactive behavior; Comments involving users expressing interest in a travel destination or future travel intentions were classified as expressed interest and plans; Comments seeking information about the location, cost, and guide of a tourist destination were classified as asking for information about the tourist destination; Comments about seeking companions, participating in social activities, and interacting with others were classified as interactive behaviors.

7. The method for quantifying user travel intention based on short video visual content presentation style according to claim 5 is characterized in that: In step S4, the pre-trained Inception-v3 model is adjusted, including the following sub-steps: S4.

1. Based on the pre-trained Inception-v3 model, load the model without the top layer and set the base layer to be non-trainable. S4.

2. Add a custom output layer, including: a global average pooling layer and a fully connected layer, wherein the global average pooling layer is used to compress the feature map into a feature vector of fixed length, and the fully connected layer is used to process the feature vector and perform binary classification through an output layer using a sigmoid activation function; S4.

3. Compile using the Adam optimizer and binary cross entropy loss function, using early stopping to prevent overfitting. S4.

4. Enhance the generalization ability of the Inception-v3 model through data augmentation and category weight adjustment.

8. The method for quantifying user travel intention based on short video visual content presentation style according to claim 7 is characterized in that: In step S5, the Inception-v3 model performs classification based on image features, and outputs a binary result of 0 or 1 for each classification, indicating that the image does not contain or contains the target element, respectively; The visual presentation style includes: experience style and landscape style. If the target element of the image is mainly a close-up of a portrait, it is classified as the experience style; If the target element of the image is mainly landscape display, it is classified as landscape style.

9. The method for quantifying user travel intention based on short video visual content presentation style according to claim 8 is characterized in that: In step S6, the multivariate linear regression model is a generalized linear model Poisson regression, and the variables of the multivariate linear regression model include: the number of comments on user travel intention , the number of pictures with mainly landscape in the video keyframes , the number of images with portraits as the main feature in the video keyframes , video duration The time interval from video publishing to video collection , the number of followers of the publisher , video likes , the number of video comments , video sharing quantity , video collection number ; The generalized linear model Poisson regression is expressed as follows: ; ; ; ; In the formula, Indicates a and The vector of β 0 ~ β 8 is the regression parameter, ε is the random interference term; The relationship between variables is analyzed by multivariate regression model, and the regression parameters are β 0 ~ β 8 Quantify the degree of influence of variables on the number of comments on tourism intention respectively; by determining the setting principles of regression parameters and statistical significance, test and determine which variables have a significant impact on the number of comments on tourism intention.

10. The method for quantifying user travel intention based on short video visual content presentation style according to claim 9, characterized in that: Before applying the multiple linear regression model, the key variables , , , Logarithmic transformation and standardization were performed, and each variable was added with 1 and then the natural logarithm was taken. The variables were then centered and scaled to unit variance to reduce the skewness and scaling differences of the key variable data.