Village scene image subjective perception evaluation method and system based on man-machine cooperation, and product

Through the subjective perception evaluation method of village scene images based on human-computer collaboration, a high-precision village scene image comparison model is constructed using deep learning and TrueSkill algorithm, which solves the subjective cognitive quantification difficulties and high-cost problems in rural environmental evaluation, and realizes efficient and low-cost village scene perception evaluation, supporting rural planning and environmental improvement.

CN120298883APending Publication Date: 2025-07-11CHINA UNIV OF GEOSCIENCES (WUHAN)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510330287.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing urban perception model is difficult to directly apply to rural environments, and the lack of comprehensive review of local data sets leads to deviations between the model and human perception. In traditional landscape evaluation, subjective cognitive quantification is difficult, massive image processing is inefficient, and manual expert evaluation costs are too high.

Method used

The subjective perception evaluation method of village scene images based on human-computer collaboration is adopted, and the image processing is performed through deep learning comparison scoring network. Combined with TrueSkill algorithm and manual annotation, a high-precision and stable village scene image comparison model is constructed, and the image features are extracted using the ResNet50 skeleton, and the training process is optimized through the human-computer collaboration process.

Benefits of technology

It improves the stability and robustness of the model, reduces the demand for computing resources, reduces labor costs, and realizes efficient subjective perception evaluation of village landscapes, can intuitively reflect subjective perception differences, and supports rural planning and environmental improvement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298883A_ABST
    Figure CN120298883A_ABST
Patent Text Reader

Abstract

The invention discloses a human-machine cooperation-based village scene image subjective perception evaluation method, system and product, and the method comprises the steps: firstly obtaining a plurality of village scene image data, and obtaining the preprocessed village scene image data through the three rule processing of invalid data elimination, latitude and longitude information extraction and EXIF information deletion; then based on the processed village scene images, randomly pairing each village scene image with other village scene images, and inputting a comparison scoring network based on deep learning to obtain a village scene image comparison sequence with high precision and strong stability; and finally, based on a village scene image comparison sequence, converting the one-group two-group classification result into an absolute score of each village scene image. According to the method, machine intelligence and human cognition are creatively combined, high-efficiency, high-precision and low-cost subjective perception evaluation of mass village scene images is realized, and decision support with objective data support and subjective cognition basis is provided for rural planning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of deep learning and computer vision technology, and relates to a subjective perception evaluation method, system and product for rural scene images, in particular to a subjective perception evaluation method, system and product for rural scene images based on human-machine collaboration. Background Art

[0002] With the rapid development of high-performance computing systems, current deep learning technologies have been able to extract higher-dimensional information from images and have gradually become the mainstream method for human subjective perception research. However, in the rural area, there is still a lack of deep learning models for systematically perceiving and evaluating the rural environment. Due to the significant differences in physical characteristics and residents' lifestyles between the rural environment and the urban environment, existing urban perception models are difficult to directly apply to the evaluation of the rural environment. In addition, previous perception models lack a comprehensive review of local datasets, resulting in a deviation between the model and human perception. Summary of the Invention

[0003] The present invention aims to solve the core problems in traditional landscape evaluation at the present stage, such as the difficulty in quantifying subjective cognition, the low efficiency of processing massive images, and the high cost of artificial expert evaluation. A subjective perception evaluation method, system and product for rural scene images based on human-machine collaboration are proposed, providing a new method and new idea for subjective cognition quantification, rural landscape perception value and dynamic monitoring of the quality of the human settlement environment.

[0004] The technical solution adopted by the method of the present invention is: a subjective perception evaluation method for rural scene images based on human-machine collaboration, comprising the following steps:

[0005] Step 1: Obtain a number of rural scene image data, and obtain the preprocessed rural scene image data after processing according to three rules: elimination of invalid data, extraction of latitude and longitude information, and deletion of EXIF information;

[0006] Step 2: Based on the processed rural scene images, randomly pair each rural scene image with other rural scene images, and input them into a deep learning-based contrast scoring network to obtain a rural scene image contrast sequence with high accuracy and strong stability;

[0007] The deep learning-based contrast scoring network includes a data augmentation and feature extraction layer, a feature fusion layer, and a contrast scoring layer, which are set in parallel; the data augmentation and feature extraction layer consists of a first Convolution layer, a second Convolution layer, a Pooling layer, a first BottleNeck layer, a second BottleNeck layer, a third BottleNeck layer, and a first BottleNeck layer connected in sequence; the feature fusion layer consists of an early fusion layer, a third Convolution layer, a fourth Convolution layer, a fifth Convolution layer, and a Fully-Connected layer; the contrast scoring layer is a softmax layer;

[0008] Step 3: Based on the rural scene image contrast sequence, convert groups of binary classification results into the absolute scores of each rural scene image.

[0009] Preferably, in step 2, in the data augmentation and feature extraction layer, the steps of the BottleNeck layer include:

[0010]

[0011] When the input and output dimensions do not match, a projection matrix W will be introduced:

[0012]

[0013] where x is the input image; H(x) is the output feature; represents the residual function; W realizes dimension expansion through a 1×1 convolutional kernel; represents the residual function, which consists of three groups of convolutional operations:

[0014]

[0015] where is a 1×1 convolutional kernel with an output channel of 64 (dimension reduction); is a 3×3 convolutional kernel with an output channel of 64 (spatial feature extraction); is a 1×1 convolutional kernel with an output channel of 256 (restoring dimension); BN(·) is the normalization function; ReLU(·) is the activation function.

[0016] Preferably, in step 2, in the feature fusion layer, the operation process of two rural scene images in the early fusion layer is:

[0017]

[0018] Among them, F1 and F2 respectively represent the features extracted from the first image and the second image; H is the feature matching matrix; is an affine transformation function to match the feature sizes of the two images; Concat(·) is feature concatenation to achieve the fusion of the channel dimension.

[0019] Preferably, in step 2, the deep learning-based contrast scoring network outputs the weight matrix W predicted by the model through the following operation process fc :

[0020]

[0021] →Pool 3×3 (·)

[0022]

[0023] →Fusion(·)

[0024]

[0025] →FC(·);

[0026] Among them, is a 7×7 convolution kernel with an output channel of 3 (RGB space); is a 7×7 convolution kernel with an output channel of 64 (dimensionality increase); Pool 3×3 is a 3×3 pooling operator (sampling); are 3 feature extraction layers with an output channel of 256; are 4 feature extraction layers with an output channel of 512; are 6 feature extraction layers with an output channel of 1024; are 3 feature extraction layers with an output channel of 2048; Fusion(·) is a feature fusion layer that fuses the feature spaces of the two images and has an output channel of 4096; is a 3×3 convolution kernel with an output channel of 4096 (to enhance the feature fine-grainedness); ReLU(·) is an activation function; FC(·) is a fully connected layer.

[0027] Preferably, in step 2, the contrast scoring layer converts the predicted value into an interval probability Y in the range of [0,1] by outputting the softmax value; through the comparison of the probability with the intermediate value of 0.5, a binary classification task of 0 or 1 is achieved; when the output value is 0, it means the first image is better; when the output value is 1, it means the second image is better. This process is shown as follows:

[0028] Y = softmax(W fc +b fc )

[0029]

[0030] Among them, Y represents the probability value; W fc represents the weight matrix of the fully connected layer; b fc represents the bias vector; softmax(·) is the probability conversion function; img1 and img2 respectively represent two comparison images.

[0031] Preferably, in step 3, based on the TrueSkill algorithm, initially the scores of all images are the same; after each comparison, according to the comparison result, the scores of the two images are adjusted: the score of the winning image increases, and the score of the losing image decreases; through multiple comparisons, the scores of all images gradually converge to stable ranking values;

[0032] Adjusting the scores of the two images according to the comparison result, the adjustment rule is: the score range of each image is modeled as a random variable N(μ,σ 2 ):

[0033]

[0034] g(x) = f(x)·[f(x)+x];

[0036] Among them, winner is the dominant image in the comparison, loser is the inferior image in the comparison; μ is the score mean, σ 2 is the score variance, N(x) is the normal distribution probability density, Φ(x) is the cumulative normal distribution, β 2 is the performance variance, γ 2 is the dynamic variance; c 2 is the scaling factor; x is the mean difference between the two images; f(x) is the cumulative distribution function of the standard normal distribution;

[0037] Normalize the initial score to the range of 0 to 10:

[0038]

[0039] Among them, Score min is the lowest score within the perception dimension, Score max is the highest score within the perception dimension, and Score is the original score.

[0040] Preferably, the deep learning-based comparison scoring network described in step 2 is a trained network;

[0041] The training of the network relies on human cognition as semantic data and is divided into two parts: the front end and the back end. First, the front end manually compares the village scene perception indicators, and the back end records these data in real time, including two comparison images and the results of good and bad. Second, the network synchronously reads each record in the back end for training and outputs two indicators: the loss value and the accuracy. Finally, it is determined whether to terminate the manual comparison process by observing whether the loss value reaches stability and whether the accuracy meets the requirements.

[0042] In the manual comparison, let the number of manual annotation rounds be \(t\in[1,T]\), and annotation samples are generated in each round:

[0043]

[0044] Among them, represents the image pair in the comparison, and \(y\) t \(\in\{0,1\}\) represents the result of the good and bad judgment manually annotated;

[0045] In the network training, for the cumulative data set the network parameters \(\theta\) are updated by optimizing the objective function:

[0046]

[0047] Among them, represents the cross-entropy loss function, and \(f\) θ is the comparison scoring network based on deep learning, which is used to output the prediction probability; \((X\) (1) , \(X\) (2) , \(y)\) represents a group of samples in the annotated data set \(D\) t ;

[0048] Define the judgment basis for training stability:

[0049]

[0050] Among them, \(L\) t represents the loss value at the \(t\)-th round;

[0051] At the same time, the accuracy requirement needs to be met:

[0052]

[0053] Among them, \(A\) t represents the accuracy of the validation set at the \(t\)-th round, and \(A\) threshold is the accuracy threshold defined by the user. When the above conditions are met simultaneously, the process is terminated.

[0054] Preferably, in the training, the cross-entropy commonly used in the deep learning binary classification task is used as the loss function \(L\);

[0055]

[0056] In the formula, (i, j, y) represents a sample; y is the true label with a value of 0 or 1; K = 2 and g is the softmax of the final activation layer; x i represents the probability distribution that the model predicts a value of 0; x j represents the probability distribution that the model predicts a value of 1;

[0057] Use the accuracy A as an evaluation index to measure whether the classification result is correct;

[0058]

[0059] In the formula, TP is the number of correctly predicted positive samples, TN is the number of correctly predicted negative samples, FP is the number of incorrectly predicted positive samples, and FN is the number of incorrectly predicted negative samples.

[0060] The technical solution adopted by the system of the present invention is: a subjective perception evaluation system for village scene images based on human-machine collaboration, including:

[0061] One or more processors;

[0062] A storage device for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the subjective perception evaluation method for village scene images based on human-machine collaboration.

[0063] The technical solution adopted by the product of the present invention is: a subjective perception evaluation product for village scene images based on human-machine collaboration, including computer program instructions, which, when run on a computer, cause the computer to execute the subjective perception evaluation method for village scene images based on human-machine collaboration.

[0064] Compared with the prior art, the beneficial effects produced by the present invention are:

[0065] The present invention patent proposes a brand-new deep learning comparison model. By implementing two data augmentation methods, saturation increase and GridMask, on village scene images, it not only greatly increases the number of samples input into the model, but also improves the stability, robustness, and generalization ability of the model. Using ResNet50 as the backbone of the model, both low-dimensional and high-dimensional features of the images can be extracted. In addition, the features of two images are ingeniously fused in the model, converting the complex comparison task into a common classification task in deep learning, reducing the high resources required for model operation, and ensuring that the model can excellently complete the decision-making task. This method provides an innovative model and technology for the research of village subjective perception evaluation and has important application value.

[0066] In addition, the human-machine collaboration training process constructed in this invention patent for invention not only integrates the feature expression ability of deep neural networks but also takes into account the semantic association mechanism of human cognition, and is a typical example of large-scale data dynamic representation. Through the model training process of human-machine collaboration, on the one hand, the performance of the model can be observed in real time, such as changes in loss values and accuracy, to control the number of manually labeled samples and reduce waste of computing resources and labor costs. On the other hand, it is possible to train the model without relying on non-local data sets, avoiding perception biases caused by images or annotators, and enhancing the controllability of model quality. After being transformed by the TrueSkill algorithm, the practical applicability of this invention is further improved. Compared with the comparison results, decision-makers are more concerned about the score ranking among the overall samples. Therefore, this algorithm can meet the most basic needs of rural planners, intuitively reflect the differences between subjective perceptions, thereby carrying out landscape restoration and improvement for rural areas with poor quality, improving the living environment quality of villagers, and contributing to the construction of beautiful villages. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] The following uses embodiments and specific implementation manners to further illustrate the technical solutions of this invention. In addition, some drawings are also used in the process of illustrating the technical solutions. For those skilled in the art, without creative efforts, other drawings and the intention of this invention can also be obtained based on these drawings.

[0068] Figure 1 is the flowchart of the embodiment of this invention;

[0069] Figure 2 is the schematic diagram of the embodiment of this invention;

[0070] Figure 3 is the structure diagram of the contrast scoring network based on deep learning in the embodiment of this invention;

[0071] Figure 4 is the training flowchart of the contrast scoring network based on deep learning in the embodiment of this invention;

[0072] Figure 5 is the schematic diagram of the rural landscape evaluation scores in the cold regions of North China in the embodiment of this invention;

[0073] Figure 6 is the schematic diagram of the rural landscape evaluation scores in the regions of Jianghuai with warm summers and cold winters in the embodiment of this invention;

[0074] Figure 7 is the schematic diagram of the rural landscape evaluation scores in the warm regions of Southwest China in the embodiment of this invention;

[0075] Figure 8 is the schematic diagram of the rural landscape evaluation scores in the regions of South China with hot summers and warm winters in the embodiment of this invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0076] To facilitate the understanding and implementation of the present invention by those of ordinary skill in the art, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the embodiments described herein are only for the purpose of illustrating and explaining the present invention and are not intended to limit the present invention.

[0077] Please refer to Figure 1 and Figure 2 , a method for subjective perception evaluation of rural scene images based on human-machine collaboration provided in this embodiment includes the following steps:

[0078] Step 1: Obtain a number of rural scene image data, and after processing through three rules of invalid data elimination, latitude and longitude information extraction, and EXIF information deletion, obtain the preprocessed rural scene image data;

[0079] A high-quality rural scene database is a prerequisite for accurate process implementation. Since most of the sources of rural scene image data are mainly crowdsourcing, the quality of the data cannot be controlled. Therefore, blurred, overexposed, and distorted images will be regarded as invalid data and should be removed from the database. Also, EXIF information data with unknown acquisition time and location needs to be deleted. The latitude and longitude information hidden in the images can be used for area identification and will also be recorded in the database together. Finally, a rural scene database with controllable quality is formed.

[0080] Step 2: Based on the processed rural scene images, randomly pair each rural scene image with other rural scene images, and input them into a contrast scoring network based on deep learning to obtain a rural scene image contrast sequence with high accuracy and strong stability;

[0081] Please refer to Figure 3 , the contrast scoring network based on deep learning includes a data augmentation and feature extraction layer, a feature fusion layer, and a contrast scoring layer arranged in parallel; the data augmentation and feature extraction layer is composed of a first Convolution layer, a second Convolution layer, a Pooling layer, a first BottleNeck layer, a second BottleNeck layer, a third BottleNeck layer, and a first BottleNeck layer connected in sequence; the feature fusion layer is composed of an early fusion layer, a third Convolution layer, a fourth Convolution layer, a fifth Convolution layer, and a Fully-Connected; the contrast scoring layer is a softmax layer;

[0082] The deep learning-based contrast scoring network predicts the landscape quality of two village scene images by taking them as inputs and fully considering their image features. Its basic principle is to use deep learning image processing techniques to capture the texture, contour, and semantic structure features of the images under a hierarchical feature extraction mechanism. The features of the two images are fused so that the pooling layer can retain the significant information in the comparison process. Finally, the fully connected layer maps the features into measurable embedding vectors.

[0083] In one implementation, during the data augmentation process, the deep learning-based contrast scoring network model uses two methods: saturation enhancement and GridMask. The former can not only highlight the specific details of low-brightness images but also improve the fine-grainedness of landscape features. The latter randomly generates 20*20 rectangular grids to occlude the local details of the images, improving the generalization ability and robustness of the model.

[0084] In one implementation, during the feature extraction process, the deep learning-based contrast scoring network model uses the architecture of the Siamese network, and two pictures share an associated network with the same weights. The ResNet50 network has excellent performance in the problems of gradient disappearance and curse of dimensionality. Therefore, ResNet50 will be used as the feature extraction backbone of the model. The steps of the Bottleneck layer are as follows:

[0085]

[0086] When the input and output dimensions do not match, a projection matrix W is introduced:

[0087]

[0088] where x is the input image; H(x) is the output feature; represents the residual function; W realizes dimensionality expansion through a 1×1 convolutional kernel; represents the residual function, which consists of three groups of convolutional operations:

[0089]

[0090] where, is a 1×1 convolutional kernel with 64 output channels (dimensionality reduction); is a 3×3 convolutional kernel with 64 output channels (spatial feature extraction); is a 1×1 convolutional kernel with 256 output channels (restoring dimensionality); BN(·) is the normalization function; ReLU(·) is the activation function.

[0091] In one embodiment, during the feature fusion process, the model combines the features extracted from two village scene images by ResNet50 using the concatenation method in early fusion and forms an overall network mapping; in the feature fusion layer, the operation process of the two village scene images in the early fusion layer is as follows:

[0092]

[0093] where F1 and F2 respectively represent the features extracted from the first image and the second image; H is the feature matching matrix; is the affine transformation function to match the feature sizes of the two images; Concat(·) is the feature concatenation to achieve the fusion of the channel dimensions.

[0094] During the prediction process, the model converts the prediction value into an interval probability in the range of [0,1] by outputting the softmax value. The model regards the comparison as a binary classification of 0 or 1. When the output value is 0, it means the first image is better, and when the output value is 1, it means the second image is better. The deep learning-based comparison scoring network outputs the weight matrix W predicted by the model through the following operation process fc :

[0095]

[0096] →Pool 3×3 (·)

[0097]

[0098] →Fusion(·)

[0099]

[0100] →FC(·);

[0101] where, is a 7×7 convolution kernel with an output channel of 3 (RGB space); is a 7×7 convolution kernel with an output channel of 64 (dimension elevation); Pool 3×3 is a 3×3 pooling operator (sampling); is 3 feature extraction layers with an output channel of 256; is 4 feature extraction layers with an output channel of 512; is 6 feature extraction layers with an output channel of 1024; is 3 feature extraction layers with an output channel of 2048; Fusion(·) is the feature fusion layer that fuses the feature spaces of the two images and has an output channel of 4096; It is a 3×3 convolutional kernel with 4096 output channels (to enhance feature granularity); ReLU(·) is the activation function; FC(·) is the fully connected layer.

[0102] Through the above steps, the constructed deep learning-based contrast scoring network can effectively achieve the perceptual contrast of two village scene images, provide an efficient evaluation method, and also provide a model foundation for simulating the human brain in the subsequent human-machine collaboration process.

[0103] Step 3: Based on the village scene image contrast sequence, convert groups of binary classification results into the absolute scores of each village scene image.

[0104] In one implementation, based on the TrueSkill algorithm, process the village scene image contrast sequence generated by the human-machine collaboration process, and convert groups of binary classification results into the absolute scores of each image. This effectively improves the readability and analyzability of the perceptual results, laying a practical foundation for large-scale village scene perception work.

[0105] Among them, TrueSkill is a multi-player competitive ranking algorithm based on Bayesian theory, originally proposed by Microsoft and used in the ranking system of multi-player online games. This algorithm iteratively updates the ranking scores of contestants after each game, generating dynamically adjusted ranking results for the winners and losers in a two-player game. In this process, pairwise comparison is regarded as a two-player competition. The two images in each pair of comparison samples are defined as "contestants", and the winner is the image that the volunteer prefers more. Initially, all images have the same score. After each comparison, the TrueSkill algorithm adjusts the scores of the two images according to the comparison result: the score of the winning image increases, and the score of the losing image decreases. Through multiple comparisons, the scores of all images gradually converge to stable ranking values. The following are the update rules for the scores of the winners and losers in each game, where the skill distribution of each player is modeled as a random variable N(μ,σ 2 )

[0106]

[0107] g(x) = f(x)·[f(x) + x];

[0109] Among them, winner is the dominant image in the comparison, loser is the inferior image in the comparison; μ is the score mean, σ 2 is the score variance, N(x) is the normal distribution probability density, Φ(x) is the cumulative normal distribution, β 2 is the performance variance, γ 2 is the dynamic variance; c 2 is the scaling factor; x is the mean difference between the two images; f(x) is the cumulative distribution function of the standard normal distribution;

[0110] Normalize the initial score to the range of 0 to 10:

[0111]

[0112] where Score min is the lowest score within the perception dimension, Score max is the highest score within the perception dimension, and Score is the original score.

[0113] In one implementation, the deep learning-based contrast scoring network is a trained network;

[0114] The training of the network relies on human cognition as semantic data and is divided into two parts: the front end and the back end. First, the front end conducts a comparison of village scene perception indicators manually, and the back end records these data in real time, including two comparison images and the results of superiority and inferiority. Second, the network synchronously reads each record from the back end for training and outputs two indicators: the loss value and the accuracy. Finally, observe whether the loss value reaches stability and whether the accuracy meets the requirements to determine whether to terminate the manual comparison process;

[0115] In the manual comparison, let the manual annotation round be t ∈ [1, T], and annotation samples are generated in each round:

[0116]

[0117] where represents the image pair in the comparison, and y t ∈ {0, 1} represents the result of the superiority and inferiority determination manually annotated;

[0118] In the network training, for the cumulative data set the network parameters θ are updated through the optimization objective function:

[0119]

[0120] where represents the cross-entropy loss function, f θ is the deep learning-based contrast scoring network for outputting the prediction probability; (X (1) , X (2) , y) represents a group of samples in the annotated data set D t ;

[0121] Define the criterion for judging the training stability:

[0122]

[0123] where L t represents the loss value at the t-th round;

[0124] Meanwhile, the accuracy requirement needs to be met:

[0125]

[0126] Among them, A t represents the accuracy of the validation set in the t-th round, and A threshold is the accuracy threshold defined by the user. When the above conditions are met simultaneously, the process is terminated.

[0127] During training, the cross-entropy commonly used in deep learning binary classification tasks is used as the loss function L;

[0128]

[0129] In the formula, (i, j, y) represents a sample; y is the true label, with a value of 0 or 1; K = 2 and g is the softmax of the final activation layer; x i represents the probability distribution that the model predicts a value of 0; x j represents the probability distribution that the model predicts a value of 1;

[0130] The accuracy A is used as an evaluation index to measure whether the classification result is correct;

[0131]

[0132] In the formula, TP is the number of correctly predicted positive samples, TN is the number of correctly predicted negative samples, FP is the number of incorrectly predicted positive samples, and FN is the number of incorrectly predicted negative samples.

[0133] This embodiment also provides a subjective perception evaluation system for rural scene images based on human-machine collaboration, including:

[0134] One or more processors;

[0135] A storage device for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the subjective perception evaluation method for rural scene images based on human-machine collaboration.

[0136] This embodiment also provides a product for subjective perception evaluation of rural scene images based on human-machine collaboration, including computer program instructions, which, when running on a computer, cause the computer to execute the subjective perception evaluation method for rural scene images based on human-machine collaboration.

[0137] The following further elaborates on the present invention through specific experiments.

[0138] The experimental team collected more than 20,000 village scene images from more than 2,000 villages in 121 counties (county-level cities) of 28 provinces across the country during the 2023 National Rural Construction Evaluation, built an online village scene evaluation platform, invited 80 scoring volunteers from all over the country, conducted 84,131 image comparisons, and collected a massive village scene label sample library. Then, the technical solution of the present invention was used for perceptual evaluation. Please see Figure 5 , which is a schematic diagram of the evaluation scores of rural landscapes in the cold region of North China. Please see Figure 6 , which is a schematic diagram of the evaluation scores of rural landscapes in the region of Jianghuai with warm summers and cold winters. Please see Figure 7 , which is a schematic diagram of the evaluation scores of rural landscapes in the warm region of Southwest China. Please see Figure 8 , which is a schematic diagram of the evaluation scores of rural landscapes in the region of South China with hot summers and warm winters. The present invention realizes the automated evaluation from "massive village scenes" to "landscape quality". It depicts the local characteristics and regional differences of rural landscapes in national counties, and reflects the process of the country, collectives and farmers improving the village landscape through rural construction. It is a new exploration and attempt of "computable countryside" in the national rural construction evaluation, provides new ideas for the acquisition, interpretation and analysis of rural data, and promotes the rural construction evaluation to move towards a more convenient, efficient and intelligent direction.

[0139] The present invention combines human subjective cognition with artificial intelligence results, integrates deep learning models, human-machine collaboration processes and TrueSkill algorithms to achieve efficient and reasonable large-scale rural landscape evaluation. The results show that by means of machine intelligence, a great deal of manual labor can be liberated, unnecessary cost input can be reduced, and the results of the model trained through cognition are basically consistent with the human evaluation results, providing an effective new method for the subjective perceptual evaluation of village scene quality. In addition, it is recommended that rural planners pay more attention to high-performance computing means, rationally utilize the subjective perceptual evaluation method of village scene images based on human-machine collaboration, control the construction and improvement of the countryside from the perspective of humanism, improve the quality of the human settlement environment, and deeply explore the value of rural landscapes. The present invention innovatively combines machine intelligence and human cognition to achieve high-efficiency, high-precision and low-cost subjective perceptual evaluation of massive village scene images, providing decision-making support for rural planning with both objective data support and subjective cognitive basis. The invention can accurately identify the perceptual value differences of rural landscapes, which not only helps to formulate scientific rural planning schemes, but also can comprehensively promote the improvement of the human settlement environment, the optimization of the layout of villages and towns, and the development of industrial integration.

[0140] It should be understood that the above-described embodiments are some, but not all, of the embodiments of the present invention. In addition, the technical features in each of the embodiments or individual embodiments provided by the present invention can be combined with each other arbitrarily to form a feasible technical solution. Such combination is not restricted by the order of steps and / or the structural composition mode, but must be based on what can be achieved by those of ordinary skill in the art. When the combination of technical solutions results in contradictions or cannot be achieved, it should be considered that such combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

[0141] It should be understood that the above description of the preferred embodiments is relatively detailed and should not be considered as a limitation on the scope of protection of the present invention. Without departing from the scope of protection defined by the claims of the present invention, those of ordinary skill in the art can make substitutions or modifications under the inspiration of the present invention, and all of them fall within the scope of protection of the present invention. The scope of protection requested by the present invention shall be subject to the appended claims.

Claims

1. A subjective perception evaluation method for village scene images based on human-machine collaboration, characterized in that, It includes the following steps: Step 1: Obtain a number of village scene image data, and after processing by three rules of invalid data elimination, latitude and longitude information extraction, and EXIF information deletion, obtain the preprocessed village scene image data; Step 2: Based on the processed village scene images, randomly pair each village scene image with other village scene images, input them into a deep learning-based contrast scoring network, and obtain a village scene image contrast sequence with high precision and strong stability; The deep learning-based contrast scoring network includes a data augmentation and feature extraction layer, a feature fusion layer, and a contrast scoring layer set in parallel; the data augmentation and feature extraction layer is composed of a first Convolution layer, a second Convolution layer, a Pooling layer, a first BottleNeck layer, a second BottleNeck layer, a third BottleNeck layer, and a first BottleNeck layer connected in sequence; the feature fusion layer is composed of an early fusion layer, a third Convolution layer, a fourth Convolution layer, a fifth Convolution layer, and a Fully-Connected; the contrast scoring layer is a softmax layer; Step 3: Based on the village scene image contrast sequence, convert a group of binary classification results into the absolute scores of each village scene image.

2. The method for subjective perception evaluation of village scene images based on human-machine collaboration according to claim 1, wherein: In Step 2, in the data augmentation and feature extraction layer, the steps of the BottleNeck layer include: When the input and output dimensions do not match, a projection matrix W will be introduced: where x is the input image; H(x) is the output feature; denotes the residual function; W realizes dimensional expansion through a 1×1 convolutional kernel; denotes the residual function, which consists of three groups of convolutional operations: Among them, is a 1×1 convolutional kernel with an output channel of 64; is a 3×3 convolutional kernel with an output channel of 64; is a 1×1 convolutional kernel with an output channel of 256; BN(·) is a normalization function; ReLU(·) is an activation function.

3. The method for subjective perception evaluation of rural scene images based on human-machine collaboration according to claim 1, characterized in that: In Step 2, in the feature fusion layer, the operation process of two village scene images in the early fusion layer is: Among them, F1 and F2 respectively represent the features extracted from the first image and the second image; H is the feature matching matrix; is an affine transformation function to match the feature sizes of the two images; Concat(·) is feature concatenation to achieve fusion in the channel dimension.

4. The subjective perception evaluation method of village scene images based on human-machine collaboration according to claim 1, characterized in that: In step 2, the deep learning-based contrast scoring network outputs the weight matrix W predicted by the model through the following operation process fc : Among them, is a 7×7 convolution kernel with 3 output channels; is a 7×7 convolution kernel with 64 output channels; Pool 3×3 is a 3×3 pooling operator; is composed of 3 feature extraction layers with 256 output channels; is composed of 4 feature extraction layers with 512 output channels; is composed of 6 feature extraction layers with 1024 output channels; is composed of 3 feature extraction layers with 2048 output channels; Fusion(·) is the feature fusion layer that fuses the feature spaces of two images and outputs 4096 channels; It is a 3×3 convolutional kernel with 4096 output channels; ReLU(·) is the activation function; FC(·) is the fully connected layer.

5. The method for subjective perception evaluation of village scene images based on human-machine collaboration according to claim 1, characterized in that: In Step 2, the contrast scoring layer, by outputting the softmax value, converts the predicted value into an interval probability Y in the range of [0,1]; through the comparison of the probability with the intermediate value of 0.5, a binary classification task of 0 or 1 is realized; when the output value is 0, it means the first image is better; when the output value is 1, it means the second image is better. This process is shown as follows: Among them, Y represents the probability value; W fc represents the weight matrix of the fully connected layer; b fc represents the bias vector; softmax(·) is a probability conversion function; img1 and img2 respectively represent two comparison images.

6. The subjective perception evaluation method of village scene images based on human-machine collaboration according to claim 1, characterized in that: In Step 3, based on the TrueSkill algorithm, initially the scores of all images are the same; after each comparison, adjust the scores of the two images according to the comparison result: the score of the winning image increases, and the score of the losing image decreases; through multiple comparisons, the scores of all images gradually converge to stable ranking values; Adjust the scores of the two images according to the comparison result, and the adjustment rule is: the score range of each image is modeled as a random variable N(μ,σ 2 ): Among them, winner is the superior image in the comparison, and loser is the inferior image in the comparison; μ is the mean score, σ 2 is the score variance, N(x) is the probability density of the normal distribution, Φ(x) is the cumulative normal distribution, β 2 is the performance variance, γ 2 is the dynamic variance; c 2 is the scaling factor; x is the mean difference between the two images; f(x) is the cumulative distribution function of the standard normal distribution; Normalize the initial score to the range of 0 to 10: Among them, Score min is the lowest score within the perception dimension, and Score max is the highest score within the perception dimension. Score is the original score.

7. The subjective perception evaluation method of rural scene images based on human-machine collaboration according to any one of claims 1-6, characterized in that: The deep learning-based contrast scoring network described in Step 2 is a trained network; The training of the network depends on human cognition as semantic materials and is divided into two parts: the front end and the back end; first, the front end conducts a comparison of village scene perception indicators manually, and the back end records these data in real time, including two comparison images and the result of superiority and inferiority; secondly, the network synchronously reads each record in the back end for training and outputs two indicators: the loss value and the accuracy; finally, observe whether the loss value reaches stability and whether the accuracy meets the requirements to decide whether to terminate the manual comparison process; In manual comparison, let the manual annotation round be \(t\in[1,T]\), and annotation samples are generated in each round: Among them, represents the image pair in the comparison, and y t ∈ {0, 1} represents the determination result of the quality by manual annotation; During network training, for the cumulative dataset The network parameters θ are updated by optimizing the objective function: Among them, represents the cross-entropy loss function, and f θ is a deep learning-based contrast scoring network for outputting predicted probabilities; (X (1) , X (2) , y) represents a set of samples in the labeled dataset D t ; Define the evaluation criterion for training stability: Among them, L t represents the loss value at the t-th round; At the same time, the accuracy requirement needs to be met: Among them, A t represents the accuracy of the validation set in the t-th round, and A threshold is the accuracy threshold defined by the user. When the above conditions are met simultaneously, the process terminates.

8. The subjective perception evaluation method of village scene images based on human-machine collaboration according to claim 7, characterized in that: During training, the cross-entropy commonly used in deep learning binary classification tasks is used as the loss function \(L\); Where (i, j, y) represents a sample; y is the true label with values of 0 or 1; K = 2 and g is the softmax of the final activation layer; x i represents the probability distribution that the model's predicted value is 0; x j represents the probability distribution that the model's predicted value is 1; Use the accuracy \(A\) as an evaluation index to measure whether the classification result is correct; In the formula, \(TP\) is the number of correctly predicted positive samples, \(TN\) is the number of correctly predicted negative samples, \(FP\) is the number of incorrectly predicted positive samples, and \(FN\) is the number of incorrectly predicted negative samples.

9. A subjective perception evaluation system for village scene images based on human-machine collaboration, characterized in that, Including: One or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the method for subjective perception evaluation of rural scene images based on human-machine collaboration according to any one of claims 1 to 8.

10. A subjective perception evaluation product of village scene images based on human-machine collaboration, including computer program instructions, characterized in that: When the computer program instructions run on a computer, the computer is caused to execute the method for subjective perception evaluation of rural scene images based on human-machine collaboration according to any one of claims 1 to 8.