Dual-Modal Scoring Method and System for Coupling Park Social Media Review Texts and Images
Through deep learning technology, combining text and images in social media comments, and using the fine-tuned Park-CN-CLIP model for park perception evaluation, it solves the problem that it is difficult to effectively combine text and images in the existing technology, and achieves efficient and accurate park comment scores, which improves the participation and satisfaction of urban green space planning and management.
Patent Information
- Application Number
- CN202410016760.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-05
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-01-05
AI Technical Summary
The prior art is difficult to effectively combine text and images in social media comments to conduct park perception assessments, and traditional methods rely on manpower, are costly and have limited data sources.
The dual-modal scoring method based on deep learning is adopted, and the Park-CN-CLIP model is fine-tuned by the CN-CLIP model, combined with the Transformer encoder to fuse text and image features, and use a full-connection layer to output prediction to achieve the perceived scoring of park reviews.
It achieves efficient and accurate scoring of the perception in the park reviews, improves public participation and satisfaction with urban green space planning and management, and overcomes the problems of high cost of traditional methods and limited data sources.
Smart Images

Figure CN117932404B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a bimodal scoring method and system for coupling park social media review texts and images, and belongs to the field of digital landscape technology in the discipline of landscape architecture. Background Art
[0002] Urban green space is an ideal leisure destination for urban residents, referring to the natural or semi-natural land use state in the city, mainly in the form of vegetation, which provides residents with a wide range of ecosystem services and opportunities for getting close to nature, outdoor activities and recreational interactions, and can have a positive impact on people's physical health, mental health and social environment. In the past few decades, due to the historical pattern of urban development, high residential density and explosive urbanization speed, there has been a fierce contradiction between supply and demand between the limited urban ecological space and the surging population. The increase in urban density has also led to a decline in the quality of the ecosystem. In order to achieve the sustainable, rapid and healthy development of the city, the United Nations Sustainable Development Goals point out that cities and communities should be inclusive, safe, resilient and sustainable, and ensure a healthy lifestyle and well-being for people of all ages. Among them, in order to ensure that the urban environment can best serve the public, it is necessary to understand various natural, cultural and other elements that affect public perception in the environment from the perspective of users, which is crucial for the sustainable development of the urban environment. Therefore, people-centered green space perception assessment is becoming an important issue in many disciplines.
[0003] Traditional research mostly relies on direct on-site observation of landscape scenes or social surveys. These methods mainly rely on the cost-intensive manual assessment of experts or interviewees on field scenes, images or texts, which is very laborious, expensive, time-consuming and the data sources are relatively limited, and it is more challenging when facing larger spatial and temporal scales. At the same time, in the survey, tourists are required to evaluate specific functions, but it is inevitable that some tourists will evaluate them when they do not see the function during their visit, resulting in bias. In the era of social perception, every citizen plays the role of a "sensor", and the amount of available data has increased greatly. Among them, social media reviews are spontaneous descriptions of users, which can directly reflect the feelings and experiences of the receiving-end users, and have the characteristics of low interference and high richness, and are more likely to reflect people's preferences.
[0004] However, due to the unstructured and open nature of social media data, it is difficult to process it using traditional methods. Currently, in the research on park perception based on social media data, the sentiment classification model in machine learning is mostly used to convert the intangible perception information in the review text into a satisfaction score, and to explore the association between environmental satisfaction and different environmental characteristics. But there are still two potential research gaps:
[0005] First, current social media comment analysis is mostly based on open natural language processing APIs (such as Baidu, Tencent, and Google natural language processing platforms). Such models are trained on sentiment classification datasets with strong generality and cannot evaluate different dimensions of park comments. Currently, few studies have established a dataset with specific evaluation dimensions for park comments based on landscape research theory and trained the latest deep learning models based on this to analyze park comments more effectively and adaptively.
[0006] Second, the current understanding of social media comments often only targets the text or pictures published by users alone. However, when people share comments, they tend to post pictures and text together to express their emotions or views. There is currently no method to consider both simultaneously for comment scoring.
[0007] Therefore, in order to maximize the benefits of parks in the urban environment, it is very important to design a scoring method that couples park social media comment text and image text based on landscape perception theory. Summary of the Invention
[0008] In view of this, the present invention provides a dual-modal scoring method, system, computer device, and storage medium for coupling park social media comment text and image, which can efficiently and accurately score the perception reflected in park comments, help improve the public's participation and satisfaction in urban green space planning and management, and provide useful references for urban green space planning and management.
[0009] The first object of the present invention is to provide a dual-modal scoring method for coupling park social media comment text and image.
[0010] The second object of the present invention is to provide a dual-modal scoring system for coupling park social media comment text and image.
[0011] The third object of the present invention is to provide a computer device.
[0012] The fourth object of the present invention is to provide a storage medium.
[0013] The first object of the present invention can be achieved by adopting the following technical solutions:
[0014] A dual-modal scoring method for coupling park social media comment text and image, the method comprising:
[0015] Obtain park social media comment data, classify and label it by experts, and establish a park comment perception classification dataset, where the park social media comment data includes comment date, text, and image;
[0016] Fine-tune the CN-CLIP model based on the park review perception classification dataset to obtain the Park-CN-CLIP model;
[0017] Use the text and image encoders in the Park-CN-CLIP model to extract the features of text and images in the park review perception classification dataset, use the Transformer encoder to fuse the two types of features, and use the fully connected layer to output predictions to complete the training of the park dual-modal scoring model;
[0018] For the social media review data collected for the target park, use the trained park dual-modal scoring model to predict the perception score of the review to obtain the perception score of the park review.
[0019] Further, the obtaining of the classified and labeled park social media review data and the establishment of the park review perception classification dataset specifically include:
[0020] Obtain the initial park social media review data;
[0021] Clean the initial park social media review data and delete blank, meaningless, and duplicate review data;
[0022] Randomly select a preset number of review data from the crawled review data as the park social media review data;
[0023] Combine the text and images in the park social media review data to classify and label the reviews and establish the park review perception classification dataset.
[0024] Further, the fine-tuning of the CN-CLIP model based on the park review perception classification dataset to obtain the Park-CN-CLIP model specifically includes:
[0025] Use the image augmentation strategy for the images in the park review perception classification dataset, use the AdamW optimizer to update the parameters of the CN-CLIP model, and use the contrast loss function to minimize the distance between positive pairs and maximize the distance between negative pairs to achieve the fine-tuning of the CN-CLIP model and obtain the Park-CN-CLIP model. The image augmentation strategy includes random cropping, random flipping, and color distortion.
[0026] Further, the use of the contrast loss function to minimize the distance between positive pairs and maximize the distance between negative pairs is as follows:
[0027]
[0028] Among them, N is the total number of samples, f i is the feature vector of the i-th image, t jis the feature vector of the j-th text, and α is the scaling factor.
[0029] Furthermore, use the text and image encoders in the Park-CN-CLIP model to extract the features of the text and images in the park review perception classification dataset, use the Transformer encoder to fuse the two types of features, and use the fully connected layer to output predictions to complete the training of the park bimodal scoring model, specifically including:
[0030] Use the text encoder in the Park-CN-CLIP model to extract the feature vector of the text in the park review perception classification dataset;
[0031] Use the image encoder in the Park-CN-CLIP model to extract the feature vector of the image in the park review perception classification dataset;
[0032] Use the Transformer encoder layer based on the Multi-head self-attention module to fuse the text and image features;
[0033] Use the fully connected layer to transform the fused features so that the features are suitable for the classification task, and output the feature vector corresponding to the category of the park review perception classification dataset;
[0034] Use the cross-entropy loss function to evaluate the prediction performance of the park bimodal scoring model and perform backpropagation;
[0035] Use the accuracy rate to judge the effectiveness of the park bimodal scoring model;
[0036] According to the effectiveness of the park bimodal scoring model, save the parameters of the park bimodal scoring model with the best effect to complete the training of the park bimodal scoring model.
[0037] Furthermore, the use of the cross-entropy loss function to evaluate the prediction performance of the park bimodal scoring model and perform backpropagation is as follows:
[0038]
[0039] where y ij is the label indicating whether the i-th sample belongs to the category j, is the probability that the model predicts that the i-th sample belongs to the category j, and N is the number of samples.
[0040] Furthermore, the use of the accuracy rate to judge the effectiveness of the park bimodal scoring model is as follows:
[0041]
[0042] where n is the total number of all reviews, yi is the actual label of the i-th comment, is the predicted label of the model, is the number of comments with correct classification results.
[0043] The second object of the present invention can be achieved by adopting the following technical solutions:
[0044] A bimodal scoring system that couples park social media comment text and images, the system includes:
[0045] A building module, used to obtain park social media comment data, classify and label it, and establish a park comment perception classification data set, where the park social media comment data includes comment date, text, and images;
[0046] A fine-tuning module, used to fine-tune the CN-CLIP model based on the park comment perception classification data set to obtain the Park-CN-CLIP model;
[0047] A training module, used to use the text and image encoders in the Park-CN-CLIP model to extract the features of text and images in the park comment perception classification data set, use the Transformer encoder to fuse the two types of features, and use a fully connected layer to output predictions to complete the training of the park bimodal scoring model;
[0048] A scoring module, used to predict the perception score of comments for the social media comment data collected for the target park using the trained park bimodal scoring model to obtain the perception score of the park comment.
[0049] The third object of the present invention can be achieved by adopting the following technical solutions:
[0050] A computer device, including a processor and a memory for storing programs executable by the processor. When the processor executes the program stored in the memory, the above-mentioned bimodal scoring method is implemented.
[0051] The fourth object of the present invention can be achieved by adopting the following technical solutions:
[0052] A storage medium stores a program, and when the program is executed by a processor, the above-mentioned bimodal scoring method is implemented.
[0053] The present invention has the following beneficial effects compared with the prior art:
[0054] 1. Based on the landscape perception theory in the landscape architecture discipline, the present invention establishes a data set with specific evaluation dimensions for park comments and trains the latest deep learning model based on this. Compared with the park comment research that calls natural language processing APIs, it can score park comments more effectively and adaptively.
[0055] 2. The understanding of social media comments in the present invention simultaneously considers the text or pictures published by users, performs bimodal fusion of the visual model and the language model, has higher prediction accuracy compared with the method that only uses text or images, can enhance the recognition ability of the model, and capture the emotions or opinions expressed by users in an automatic and effective manner.
[0056] 3. The present invention can be applied to empirical studies at different scales in different cities, has strong transferability and universality, overcomes the problem that traditional methods highly rely on manpower, and the application results of the invention have important reference value for the planning, design and management of green spaces. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following-described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on the structures shown in these drawings without creative efforts.
[0058] Figure 1 It is a flowchart of the bimodal scoring method for coupling park social media comment text and images in Embodiment 1 of the present invention.
[0059] Figure 2 It is a relational graph of three dimensions of park comment scoring involved in Embodiment 1 of the present invention.
[0060] Figure 3 It is a model architecture diagram of the bimodal scoring method for coupling park social media comment text and images in Embodiment 1 of the present invention.
[0061] Figures 4 to 6 It is a visualization result of the scores of 130 parks in Guangzhou City in three dimensions of accessibility, usability and attractiveness in Embodiment 1 of the present invention.
[0062] Figure 7 It is a structural block diagram of the bimodal scoring system for coupling park social media comment text and images in Embodiment 2 of the present invention.
[0063] Figure 8 It is a structural block diagram of the computer device in Embodiment 3 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0064] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0065] Embodiment 1:
[0066] like Figure 1 As shown, this embodiment provides a dual-modal scoring method for coupling park social media review text and image, and the method includes the following steps:
[0067] S101. Obtain park social media comment data, classify and annotate them, and establish a park comment perception classification dataset.
[0068] Furthermore, step S101 specifically includes:
[0069] S1011. Obtain initial review data on the park's social media.
[0070] Specifically, the application programming interfaces of social media platforms including but not limited to Dianping.com (https: / / www.dianping.com / ) and Ctrip.com (https: / / www.ctrip.com / ) were called to obtain review data of 130 parks in Guangzhou, Guangdong Province as the initial social media review data of the parks, including review date, text and images. The collection time was up to August 1, 2023, and a total of 62,106 valid reviews and 111,488 images were obtained.
[0071] S1012. Clean the initial comment data on the park’s social media and delete blank, meaningless, and duplicate comment data.
[0072] S1013. Randomly select a preset number of comment data from the crawled comment data as park social media comment data, where the preset number is 2500.
[0073] S1014. Combine the text and images in the park social media comment data to classify and annotate the comments and establish a park comment perception classification dataset.
[0074] Specifically, by combining the text and images (text-image pairs) in the park social media comment data, two experts in the field of landscape architecture were invited to classify and label from three aspects: accessibility, usability, and attractiveness, to establish a park comment perception classification dataset; in the process of data annotation, if there are disagreements, they need to discuss and modify each other to finally obtain consistent values, ensuring the consistency of data annotation and thus the scientific nature of model training; among them, "0" represents that this dimension is poor, "1" represents that this dimension is good, and "2" represents that there is no judgment related to this dimension. According to the scoring results, some comments were appropriately added or deleted to avoid, as much as possible, the imbalance of data affecting the performance of the deep learning model. The final dataset contains 3,139 comments, and the data distribution is shown in Table 1 below.
[0075] In this embodiment, the reasons for selecting accessibility, usability, and attractiveness are as follows:
[0076] (1) The benefits that a park can provide can only be obtained when residents can reasonably enter and use the park. Accessibility refers to the ease of reaching the park and is an assessment of the opportunity to access or use green spaces; usability refers to whether people can freely reach and enter the park and safely use it for entertainment purposes at any time; attractiveness refers to whether the park meets the personal needs, expectations, and preferences of users when a person is willing to use it and spend his or her time there. The three indicators are nested layer by layer. Before judging the attractiveness of a park, the park must be accessible and usable.
[0077] (2) Social media data is essentially subjective, but the first-person, location-based, and multi-dimensional attitudes it reflects are objective and can show emotions, opinions, views, and values related to landscapes and / or ecosystems driven by functionality (such as recreational activities), aesthetics, sense of place, sense of belonging, or the presence of biodiversity. Therefore, the subjective perception can be quantified by objectively scoring the comments. Taking the following comment as an example, the label for its accessibility is "0", "The transportation is not very convenient"; the label for usability is "0", "There are few small shops in the park, and it's even difficult to find a place to buy water"; the label for attractiveness is "0", and the overall emotion reflected is objective and negative.
[0078] The relationship diagram of the three dimensions of accessibility, usability, and attractiveness in this embodiment is as Figure 2 shown.
[0079] Comment example: "I came here by bus from home and transferred twice. Actually, the transportation is not very convenient. I took the bus to Gaodehui in Science City and it wasn't too far to walk directly. Because I came here with the intention of taking a walk and exercising, I still wanted to walk up the mountain at the entrance. There is no entrance fee. There aren't many people and the air is relatively fresh. However, there are few small shops in the park, and it's even difficult to find a place to buy water."
[0080] Data distribution status of Table 1
[0081]
[0082] S102. Fine-tune the CN-CLIP model based on the park review perception classification dataset to obtain the Park-CN-CLIP model.
[0083] Among them, the CN-CLIP (Chinese Contrastive Language Image Pretraining) model is trained using a large amount of Chinese data, approximately containing 200 million text-image pairs. The model architecture of the image encoder is ViT, and the model architecture of the text encoder is RoBERTa-wwm-Base.
[0084] Furthermore, this step S102 specifically includes:
[0085] S1021. Perform format preprocessing on the park review perception classification dataset to generate the following two formats of files:
[0086] TSV file: Each line of the file represents an image, including the image ID (int type) and the image content (stored in base64 format), separated by tabs. The format is as follows:
[0087] 506 / 9j / 4AAQSkZJ...YQj7314oA / / 2Q==
[0088] JSON file: Each line of the file contains the text ID, text content, and its corresponding image ID. The format is as follows:
[0089] {"text_id":606,"text":"I went to the Grand View Wetland Park when the kapok trees were in bloom...",
[0090] "image_ids":[1206,1207]}
[0091] S1022. Load the pre-trained CN-CLIP model, which can generate good feature representations for text and images.
[0092] S1023. Fine-tune the CN-CLIP model based on the park review perception classification dataset. Among them, for the images in the park review perception classification dataset, image augmentation strategies such as random cropping, random flipping, and color distortion are used to increase the generalization ability of the model, and the AdamW optimizer is used to update the model parameters. The contrast loss function is used to minimize the distance between positive pairs and maximize the distance between negative pairs. Its calculation formula is:
[0093]
[0094] Among them, N is the total number of samples, f i is the feature vector of the i-th image, t j is the feature vector of the j-th text, and α is the scaling factor.
[0095] S1024. Save the weight file generated after fine-tuning the CN-CLIP model as the Park-CN-CLIP model.
[0096] S103. Use the text and image encoders in the Park-CN-CLIP model to extract the features of the text and images in the park review perception classification dataset, use the Transformer encoder to fuse the two types of features, and use the fully connected layer to output predictions to complete the training of the park bimodal scoring model.
[0097] Furthermore, this step S103 specifically includes:
[0098] S1031. Use the text encoder in the Park-CN-CLIP model to extract the feature vectors of the text in the park review perception classification dataset.
[0099] S1032. Use the image encoder in the Park-CN-CLIP model to extract the feature vectors of the images in the park review perception classification dataset.
[0100] S1033. Use the Transformer encoder layer based on the Multi-head self-attention module to fuse the text and image features.
[0101] Among them, using the Dropout regularization technique helps to enhance the robustness of the model on the training data and the generalization ability on unseen data.
[0102] In this embodiment, using the Transformer encoder layer based on the Multi-head self-attention module to fuse the text and image features is specifically as follows:
[0103] In the traditional attention mechanism, the query (Q), key (K), and value (V) are obtained by multiplying the matrix multiplication with specific weight matrices. Multi-head self-attention is actually multiple parallel attention layers, and each layer has its own independent weights, capturing the dependencies of the data from different subspaces respectively.
[0104] Given a query matrix Q, a key matrix K, and a value matrix V, the multi-head attention is defined as:
[0105] MultiHead(Q, K, V) = Concat(head1, head2,..., head H )W O
[0106] where each head is defined as:
[0107] head i = Attention(QW Qi , KW Ki , VW Vi )
[0108] where W Qi , W Ki and W Vi are weight matrices associated with each head, and W O is the weight matrix used to integrate the outputs of all heads.
[0109] For the attention weights of each head, the calculation formula is:
[0110]
[0111] where d k is the dimension of the key vector.
[0112] The Softmax activation function is used to transform the original output into a probability distribution, and the calculation formula is:
[0113]
[0114] where x i is the i-th element of the output of the fully connected layer.
[0115] S1034. Use a fully connected layer to transform the fused features so that the features are suitable for the classification task, and output a feature vector corresponding to the categories of the park review perception classification dataset.
[0116] Specifically, use a fully connected linear layer to transform the extracted and fused features so that they can be suitable for the classification task, and output a feature vector with a size of "out_hidden_size = 3", corresponding to the three categories of the park review perception classification dataset; among them, use Dropout to prevent overfitting and increase its generalization ability on unseen data, and use the ReLU activation function to increase the non-linearity of the model and enable the model to capture more complex patterns in the data.
[0117] S1035. Evaluate the prediction performance of the park bimodal scoring model using the cross-entropy loss function and perform backpropagation as follows:
[0118]
[0119] where y ij is the label indicating whether the i-th sample belongs to class j, is the probability that the model predicts the i-th sample belongs to class j, and N is the number of samples
[0120] S1036. Evaluate the effectiveness of the park bimodal scoring model using accuracy as follows:
[0121]
[0122] where n is the total number of all comments, y i is the actual label of the i-th comment, is the predicted label of the model, is the number of comments with correct classification results.
[0123] S1037. Save the parameters of the park bimodal scoring model with the best performance according to the effectiveness of the park bimodal scoring model to complete the training of the park bimodal scoring model.
[0124] In this embodiment, during the training of the park bimodal scoring model, the AdamW optimizer is adopted, the maximum learning rate is initialized to 5×10 -5 , the learning rate strategy is cosine annealing, and the decay weight is set to 5×10 -4 ; Considering the size of the dataset and the GPU performance, the batch size is set to 32; To avoid model overfitting and increase time efficiency, an early stopping strategy is added, that is, if the accuracy and loss of the data on the validation set do not change in 20 training epochs, the existing best weights are saved and the park bimodal scoring model stops the training phase.
[0125] The model architectures involved in step S102 and step S103 of this embodiment are as Figure 3As shown below, a comparative experiment was conducted to further demonstrate the superiority of the model proposed in this embodiment. The experimental results are shown in Table 2. Among them, Park-CN-CLIP is the adjusted CN-CLIP model; BERT is a deep bidirectional natural language processing pre-training model based on the Transformer architecture, which can understand the context relationship in the text and is widely used in various NLP tasks. LSTM is a recursive neural network (RNN) architecture, specially designed to handle long-term dependencies and is widely used in sequence data such as time series and text. XGBoost is an optimized gradient boosting library, specially designed for efficient and flexible machine learning implementation, and is widely used in various classification and regression tasks. In this embodiment, the model will be trained up to 1000 rounds at most. To avoid overfitting, an early stopping strategy is used, that is, if the performance of the validation set does not improve further within 50 consecutive rounds, the training will stop early.
[0126] Table 2 Comparison of prediction accuracies of different methods
[0127]
[0128] The prediction accuracies of the model in this embodiment for reachability, usability, and attractiveness reached 85.29%, 84.89%, and 89.26% respectively, which are at a relatively high level, superior to the text or image encoders using Park-CN-CLIP alone, and are 2.18%, 7.19%, and 14.43% higher than BERT respectively, indicating that when only considering the semantic information of the review text, the algorithm performance is poor. At the same time, it also reflects that integrating image feature information into the perception prediction model is effective, and also proves the superiority of the proposed method of coupling text and image and constructing a deep learning model. Using the fine-tuned Park-CN-CLIP text encoder also has an accuracy 4.41% and 11.85% higher than the commonly used BERT model. The prediction effects of using the Park-CN-CLIP image encoder, LSTM, and XGBoost alone are poor.
[0129] In summary, the model proposed in this embodiment has achieved the best test results so far and can reliably automatically score the human perception in social media reviews.
[0130] S104. For the social media review data collected for the target park, use the trained park dual-modal scoring model to predict the perception score of the review to obtain the perception score of the park review.
[0131] Taking Guangzhou City as an example, this embodiment demonstrates the feasibility and practical application ability of the method proposed in this embodiment, including the following steps:
[0132] S1041. Automatically score 62,106 comments and 111,488 pictures from three dimensions of accessibility, usability, and attractiveness using the trained park dual-modal scoring model.
[0133] S1042. Use Python to calculate the scores of the three dimensions for 130 parks on a park-by-park basis.
[0134] S1043. Draw park points in ArcMap 10.6, connect and match them with the scores, and perform visual display according to the perceived scores of each point. As Figures 4 to 6 shown, it is found that the situation of the perceived distribution is highly consistent with the economic development level of the region, which proves the effectiveness of the park dual-modal scoring model proposed in this embodiment for urban perception evaluation.
[0135] It should be noted that although the method operations of the above embodiments are described in a specific order, this does not require or imply that these operations must be performed in that specific order, or that all the shown operations must be performed to achieve the desired result. On the contrary, the described steps can be changed in the order of execution. Additionally or alternatively, certain steps can be omitted, multiple steps can be combined into one step for execution, and / or one step can be decomposed into multiple steps for execution.
[0136] Embodiment 2:
[0137] As Figure 7 shown, this embodiment provides a dual-modal scoring system that couples park social media comment text and images. The system includes an establishment module 701, a fine-tuning module 702, a training module 703, and a scoring module 704. The specific descriptions of each module are as follows:
[0138] The establishment module 701 is used to obtain park social media comment data, perform classification and annotation, and establish a park comment perception classification data set. The park social media comment data includes comment date, text, and images.
[0139] The fine-tuning module 702 is used to fine-tune the CN-CLIP model based on the park comment perception classification data set to obtain the Park-CN-CLIP model.
[0140] The training module 703 is used to extract the features of text and images in the park comment perception classification data set using the text and image encoders in the Park-CN-CLIP model, fuse the two types of features using a Transformer encoder, and output predictions using a fully connected layer to complete the training of the park dual-modal scoring model.
[0141] A scoring module 704, which is used to predict the perception score of comments for the social media comment data collected for the target park by using the trained park bimodal scoring model, so as to obtain the perception score of the park comments.
[0142] It should be noted that the system provided in this embodiment is only illustrated by the above division of each functional module. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure is divided into different functional modules to complete all or part of the functions described above.
[0143] Embodiment 3:
[0144] This embodiment provides a computer device, as Figure 8 shown, which includes a processor 802, a memory, an input device 803, a display device 804, and a network interface 805 connected through a system bus 801. The processor is used to provide computing and control capabilities. The memory includes a non-volatile storage medium 806 and an internal memory 807. The non-volatile storage medium 806 stores an operating system, a computer program, and a database. The internal memory 807 provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. When the processor 802 executes the computer program stored in the memory, the bimodal scoring method of the above Embodiment 1 is implemented as follows:
[0145] Obtain park social media comment data, classify and label it, and establish a park comment perception classification data set. The park social media comment data includes comment date, text, and images;
[0146] Based on the park comment perception classification data set, fine-tune the CN-CLIP model to obtain the Park-CN-CLIP model;
[0147] Use the text and image encoders in the Park-CN-CLIP model to extract the features of the text and images in the park comment perception classification data set, use the Transformer encoder to fuse the two types of features, and use a fully connected layer to output predictions to complete the training of the park bimodal scoring model;
[0148] For the social media comment data collected for the target park, use the trained park bimodal scoring model to predict the perception score of the comments to obtain the perception score of the park comments.
[0149] Embodiment 4:
[0150] This embodiment provides a storage medium, which is a computer-readable storage medium and stores a computer program. When the computer program is executed by a processor, the bimodal scoring method of the above Embodiment 1 is implemented as follows:
[0151] Obtain park social media comment data, classify and label it, and establish a park comment perception classification dataset. The park social media comment data includes comment date, text, and images.
[0152] Fine-tune the CN-CLIP model based on the park comment perception classification dataset to obtain the Park-CN-CLIP model.
[0153] Use the text and image encoders in the Park-CN-CLIP model to extract the features of the text and images in the park comment perception classification dataset. Use the Transformer encoder to fuse the two types of features, and use the fully connected layer to output predictions to complete the training of the park bimodal scoring model.
[0154] For the social media comment data collected for the target park, use the trained park bimodal scoring model to predict the perception score of the comment to obtain the perception score of the park comment.
[0155] It should be noted that the computer-readable storage medium in this embodiment can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0156] In this embodiment, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. And in this embodiment, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable storage medium other than a computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable storage medium can be transmitted by any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0157] The above computer-readable storage medium can be written in one or more programming languages or combinations thereof for executing the computer program of this embodiment. The above programming languages include object-oriented programming languages such as Java, Python, C++, and also include conventional procedural programming languages such as C language or similar programming languages. The program can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0158] In summary, based on the landscape research theory of the landscape architecture discipline, the present invention uses deep learning technologies such as contrastive learning and multimodal fusion to be able to efficiently and accurately score the perceptions reflected in park reviews, which helps to improve the public's participation and satisfaction in urban green space planning and management, and provides a useful reference for the urban green space planning and management.
[0159] As described above, only the preferred embodiments of the present invention are provided. For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to include all changes falling within the meaning and scope of the equivalent elements of the claims in the present invention.
Claims
1. A dual-modal scoring method that couples park social media review text and images, characterized in that: The method comprises: Obtain park social media review data, and classify and annotate them from three dimensions: accessibility, usability, and attractiveness, to establish a park review perception classification dataset. The park social media review data includes review date, text, and image. In the classification annotation, "0" represents that this dimension is poor, "1" represents that this dimension is good, and "2" represents that this dimension is not involved in the judgment; Based on the park review perception classification dataset, the CN-CLIP model is fine-tuned to obtain the Park-CN-CLIP model; Use the text and image encoders in the Park-CN-CLIP model to extract the features of text and images in the park review perception classification dataset, use the Transformer encoder to fuse the two types of features, use the fully connected layer to output the prediction, and complete the training of the park bimodal rating model; For the social media review data collected from the target park, the trained bimodal rating model for the park is used to predict the perception score of the reviews and obtain the perception score of the park reviews. Based on the park review perception classification dataset, the CN-CLIP model is fine-tuned to obtain the Park-CN-CLIP model, which specifically includes: An image augmentation strategy is used for images in the park review perception classification dataset. The AdamW optimizer is used to update the CN-CLIP model parameters. The contrast loss function is used to minimize the distance between positive pairs and maximize the distance between negative pairs to fine-tune the CN-CLIP model and obtain the Park-CN-CLIP model. The image augmentation strategy includes random cropping, random flipping, and color distortion. The text and image encoder in the Park-CN-CLIP model is used to extract the features of text and image in the park review perception classification dataset, the two types of features are fused using the Transformer encoder, and the prediction is output using the fully connected layer to complete the training of the park bimodal rating model, which specifically includes: Use the text encoder in the Park-CN-CLIP model to extract the feature vector of the text in the park review perception classification dataset; use the image encoder in the Park-CN-CLIP model to extract the feature vector of the image in the park review perception classification dataset; use the Transformer encoder layer based on the Multi-head self-attention module to fuse the text and image features; use the fully connected layer to transform the fused features to make the features suitable for classification tasks, and output the feature vector corresponding to the category of the park review perception classification dataset; use the cross entropy loss function to evaluate the prediction performance of the park bimodal rating model and perform back propagation; use the accuracy rate to judge the effectiveness of the park bimodal rating model; according to the effectiveness of the park bimodal rating model, save the parameters of the park bimodal rating model with the best effect to complete the training of the park bimodal rating model.
2. The dual-modal scoring method according to claim 1, characterized in that: The acquisition of park social media comment data, classification and annotation, and establishment of a park comment perception classification dataset specifically include: Obtain initial park social media review data; Clean the initial review data of the park's social media and delete blank, meaningless and duplicate review data; Randomly select a preset number of comment data from the crawled comment data as the park social media comment data; Combining the text and images in the park social media comment data, the comments are classified and annotated to establish a park comment perception classification dataset.
3. The dual-modal scoring method according to claim 1, characterized in that: The contrast loss function is used to minimize the distance between positive pairs and maximize the distance between negative pairs, as follows: Where N is the total number of samples, f i is the feature vector of the i-th image, t j is the feature vector of the jth text, and α is the scaling factor.
4. The dual-modal scoring method according to claim 1, characterized in that: The cross entropy loss function is used to evaluate the predictive performance of the park bimodal rating model and perform back propagation, as shown in the following formula: Among them, y ij is the label of whether the i-th sample belongs to category j, The model predicts the probability that the i-th sample belongs to category j, and N is the number of samples.
5. The dual-modal scoring method according to claim 1, characterized in that: The accuracy rate is used to judge the effectiveness of the park bimodal scoring model, as follows: Where n is the total number of all comments, y i is the actual label of the i-th comment, is the predicted label of the model, The number of comments that are correctly classified.
6. A dual-modal rating system that couples park social media review text and images, characterized in that: The system comprises: Establish a module for obtaining park social media review data and classifying and annotating them from three dimensions: accessibility, usability, and attractiveness, to establish a park review perception classification dataset. The park social media review data includes review date, text, and image. In the classification annotation, "0" represents that this dimension is poor, "1" represents that this dimension is good, and "2" represents that no judgment is involved in this dimension. The fine-tuning module fine-tunes the CN-CLIP model based on the park review perception classification dataset to obtain the Park-CN-CLIP model; The training module is used to use the text and image encoders in the Park-CN-CLIP model to extract the features of the text and images in the park review perception classification dataset, fuse the two types of features using the Transformer encoder, and use the fully connected layer to output the prediction to complete the training of the park bimodal rating model; The scoring module is used to predict the perception scores of the social media review data collected from the target park using the trained bimodal scoring model of the park to obtain the perception scores of the park reviews; Based on the park review perception classification dataset, the CN-CLIP model is fine-tuned to obtain the Park-CN-CLIP model, which specifically includes: An image augmentation strategy is used for images in the park review perception classification dataset. The AdamW optimizer is used to update the CN-CLIP model parameters. The contrast loss function is used to minimize the distance between positive pairs and maximize the distance between negative pairs to fine-tune the CN-CLIP model and obtain the Park-CN-CLIP model. The image augmentation strategy includes random cropping, random flipping, and color distortion. The text and image encoder in the Park-CN-CLIP model is used to extract the features of text and image in the park review perception classification dataset, the two types of features are fused using the Transformer encoder, and the prediction is output using the fully connected layer to complete the training of the park bimodal rating model, which specifically includes: Use the text encoder in the Park-CN-CLIP model to extract the feature vector of the text in the park review perception classification dataset; use the image encoder in the Park-CN-CLIP model to extract the feature vector of the image in the park review perception classification dataset; use the Transformer encoder layer based on the Multi-head self-attention module to fuse the text and image features; use the fully connected layer to transform the fused features to make the features suitable for classification tasks, and output the feature vector corresponding to the category of the park review perception classification dataset; use the cross entropy loss function to evaluate the prediction performance of the park bimodal rating model and perform back propagation; use the accuracy rate to judge the effectiveness of the park bimodal rating model; according to the effectiveness of the park bimodal rating model, save the parameters of the park bimodal rating model with the best effect to complete the training of the park bimodal rating model.
7. A computer device comprising a processor and a memory for storing a program executable by the processor, characterized in that: When the processor executes the program stored in the memory, the dual-modal scoring method according to any one of claims 1 to 5 is implemented.
8. A storage medium storing a program, characterized in that: When the program is executed by a processor, the dual-modal scoring method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Network media multi-modal information extraction method based on Transform and data enhancement
CN117152573A
CLIP model-based vulnerable plaque identification method and system
CN117198514A