Street subjective perception evaluation method under support of multi-mode large model and streetscape image

By constructing a dynamic closed-loop system of geographic context database and multimodal large model, combined with computer vision data, the efficiency and accuracy problems of subjective perception evaluation of street scene environment are solved, realizing efficient and personalized street scene evaluation, which is applicable to different cities and street types.

CN121051184APending Publication Date: 2025-12-02TONGJI UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511597997.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-04
Publication Date
2025-12-02

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently and accurately evaluate subjective perceptions of street scene environments. Traditional methods are time-consuming, costly, and produce inconsistent results. Computer vision methods, on the other hand, cannot simulate human cognitive mechanisms, leading to significant discrepancies between evaluation results and human perceptions. Large models also exhibit inaccuracies in professional applications.

Method used

A geographic context database is constructed, and a dynamic closed-loop system combining a multimodal large model and the ELO algorithm, along with computer vision data, is used to achieve efficient and personalized evaluation of street scene environments.

Benefits of technology

It achieves efficient and accurate street view environment assessment, with results highly consistent with human perception, reduces labor costs, is applicable to assessment needs of different cities and street types, and has good scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121051184A_ABST
    Figure CN121051184A_ABST
Patent Text Reader

Abstract

The invention discloses a street subjective perception evaluation method under the support of a multi-mode large model and a streetscape image. The method comprises the following steps: firstly, constructing a geographic situation database through data integration; designing a multi-dimensional cognitive evaluation system, and automatically generating structured and personalized cue words; and finally, deeply combining an ELO algorithm with a multi-modal large model, and constructing a dynamically adaptive intelligent evaluation assembly line, namely, directly calling prompt word data in a database on the basis of dynamically generating an optimal pairing combination according to ELO scores through a mechanism of information interaction between a scoring module and the large model, so as to obtain an intelligent evaluation result. And a large model evaluation result is formed and fed back to an ELO system to realize score updating and then to an intelligent closed-loop system of next round of pairing optimization. According to the method, the illusion problem of a large model in professional evaluation can be effectively reduced, so that high personification intelligent evaluation of large-scale streetscape images is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the interdisciplinary field of geographic information and artificial intelligence in urban and rural planning, and specifically relates to a subjective perception evaluation method for streets based on large models and street view images. Background Technology

[0002] With urban development, quantitative evaluation of the perceived built environment has become an important way to improve the quality of urban street environments and enhance urban vitality. Traditional methods often rely on questionnaires, GIS spatial analysis, or manual ground surveys, but these often face problems such as long data acquisition cycles and significant subjective biases. In recent years, street view imagery has provided a high-resolution, widely covered visual data source for assessing the subjective perceived quality of streets. Recent advances in urban analysis methods, particularly the combination of street view image data and computer vision algorithms, such as semantic segmentation, have enriched the methods for evaluating the subjective perceived quality of streets by enabling automatic assessment of built environment characteristics from a subjective perspective. However, because computer vision algorithms typically assess built environment characteristics from a subjective perspective in isolation, they often fail to consider the objective built environment of the urban space corresponding to the street view imagery, thus limiting their overall ability to assess the subjective perceived quality of urban streets.

[0003] Recently, the emergence of multimodal large language models has introduced a revolutionary approach to image analysis by mimicking human perception. These models possess the ability to jointly understand images and text, generating image descriptions and completing instruction-driven image question-and-answer sessions, making it possible to directly assess the suitability of streetscape environments for specific groups (such as cyclists). These models significantly reduce the cost of manual annotation and possess strong generalization capabilities. While traditional methods for assessing streetscape environments can capture macroscopic built-up environment features based on streetscape images and GIS data, they struggle to integrate microscopic design features (such as bike path quality and the visual quality of green spaces) with macroscopic built-up environment features. By leveraging the visual reasoning capabilities of large language models and combining them with multi-source spatial data, accurate assessments of streetscape environments can be achieved. The assessment results of large language models can closely approximate human subjective perceptions of streetscape environments, identifying microscopic aesthetic and functional details of streets that traditional visual recognition models overlook.

[0004] However, existing technologies also have the following major drawbacks: 1. Limitations of Traditional Manual Evaluation Methods: Traditional subjective street view evaluations primarily rely on methods such as field surveys, questionnaires, and expert assessments. While these methods can obtain authentic human perception data, they suffer from significant technical bottlenecks: the evaluation process is time-consuming, labor-intensive, and has limited sample coverage, making it difficult to achieve systematic evaluations of large-scale urban streets. Furthermore, subjective differences among evaluators lead to poor consistency in results, and a standardized evaluation system is lacking.

[0005] 2. Cognitive Deficiencies of Computer Vision Methods: Existing computer vision-based street view evaluation methods mainly rely on techniques such as object detection and semantic segmentation to identify street view elements and derive evaluation results through statistical analysis. However, these methods have a fundamental flaw: they lack simulation of human cognitive mechanisms and cannot understand the influence mechanism of street view elements on human psychological perception. The evaluation results are often just simple feature statistics, which differ significantly from true human subjective perception.

[0006] 3. Specialization Issues in Large-Scale Model Applications: While large language models perform well in many domains, they exhibit significant shortcomings in specialized applications. The Transformer architecture relies on massive amounts of language data for content inference, lacking genuine logical reasoning capabilities and specialized knowledge accumulation. In tasks requiring expert judgment, such as street view evaluation, it is prone to "model illusion," resulting in inaccurate or inconsistent outputs.

[0007] These technological limitations make it difficult for existing methods to meet the needs of urban planning and transportation infrastructure optimization for efficient, accurate, and standardized street view evaluation. Summary of the Invention

[0008] To address the shortcomings of existing technologies, this invention provides a subjective street perception evaluation method based on a multimodal large model and street view imagery. By constructing a geographic context database and a multidimensional cognitive evaluation system, it compensates for the shortcomings of traditional survey methods in terms of efficiency and machine learning evaluation methods in terms of anthropomorphism, thus achieving a highly efficient and highly anthropomorphic street view perception evaluation.

[0009] Technical solution: A method for subjective street perception evaluation supported by a multimodal large model and street view imagery includes the following steps: Step 1: Convert the heterogeneous raw data from multiple sources into standardized data and establish a geographic context memory database; Step 2: Convert the technical data obtained from computer vision into natural language descriptions that the large model can understand; then, according to the thought chain format, assemble the converted data and prompt word templates, and provide them to Step 3 along with the images; Step 3: Construct a dynamic closed-loop system for the ELO algorithm and the large model, call the multimodal large model API, execute specific interpretable evaluation tasks, output in a fixed format, the ELO scoring system calculates scores and performs pair optimization based on the win and loss results of the large model, and automatically terminates the evaluation according to the stability criterion, thus completing the subjective perception evaluation of the street.

[0010] Beneficial effects: (1) Innovatively realized the dynamic closed-loop integration of ELO algorithm and large model. This invention is the first to deeply integrate the ELO ranking algorithm with large-model batch processing technology, constructing a dynamically feedback-adjusted intelligent evaluation pipeline. Unlike traditional static batch processing, the system achieves a closed-loop system that provides real-time feedback of evaluation results to the pairing strategy optimization through an alternating control mechanism between the control management module and the large model, significantly improving evaluation efficiency.

[0011] (2) A human-like cognitive evaluation mechanism was realized. By constructing a geographic context database and further processing the basic data and supplementing the background information, the large model was equipped with the cognitive background that simulates human evaluation, which significantly improved the consistency between the evaluation results and human real perception.

[0012] (3) A personalized evaluation system based on precise geographic coordinate indexing was constructed. The precise indexing mechanism based on GPS coordinates can generate personalized prompts containing the geographic context memories of both parties for each pairing, making the evaluation process highly targeted and accurate. This personalized mechanism makes the evaluation results closer to actual application scenarios and enhances the practical value of the system.

[0013] (4) It achieves efficient and intelligent batch evaluation processing. Compared with traditional manual survey methods, the processing efficiency is greatly improved, and the labor cost generated by manual evaluation is greatly reduced. By automatically determining the timing of evaluation termination through intelligent stability assessment, the problems of over-evaluation or under-evaluation are avoided, and the automation level of the system is greatly improved.

[0014] (5) Provides a scalable intelligent technology framework. The technical framework established by this invention has good scalability and is applicable to subjective perception evaluation of different cities and different types of streets. By adjusting the content of the memory database, the weight of the evaluation dimensions and the ELO algorithm parameters, it can adapt to the evaluation needs of different regions and application scenarios and has strong technology transfer capabilities. Attached Figure Description

[0015] Figure 1 Overall flowchart of the method of this invention; Figure 2 The processing results of this invention are shown in the figure. Detailed Implementation

[0016] The technical solutions provided in this application will be further described below with reference to specific embodiments and accompanying drawings.

[0017] A method for subjective perception evaluation of streets supported by a multimodal large model and street view imagery includes the following steps: (e.g.) Figure 1 ) Step 1: Convert the heterogeneous raw data from multiple sources into standardized data and establish a geographic context memory database.

[0018] Step 2: Convert the technical data obtained from computer vision into natural language descriptions that the large model can understand. Then, following the thought chain format, assemble the converted data and prompt word templates, and provide them along with the images to Step 3.

[0019] Step 3: Construct a dynamic closed-loop system for the ELO algorithm and the large model, call the multimodal large model API, execute specific interpretable evaluation tasks, output in a fixed format, the ELO scoring system calculates scores and performs pair optimization based on the win and loss results of the large model, and automatically terminates the evaluation according to the stability criterion, thus completing the subjective perception evaluation of the street.

[0020] Furthermore, in step 1, the multi-source heterogeneous raw data includes: street view image data (including GPS coordinate metadata), POI geographic information data, and semantic segmentation data; wherein, the semantic segmentation data is preprocessed data obtained by feature extraction of street view image data through a convolutional neural network model, and the POI geographic information data can be obtained from open-source map websites.

[0021] The method for converting the raw data into standardized data includes: semantic segmentation data preprocessing and spatial analysis algorithms. Specifically, a convolutional neural network architecture model is used to extract features from street view image data to obtain preprocessed semantic segmentation data, resulting in pixel proportion data for elements such as roads, buildings, greenery, and sky. The POI density distribution is calculated using a spatial analysis algorithm; all data are indexed using GPS coordinates. The spatial analysis algorithm described is existing technology.

[0022] The output standardized technical data (including semantic segmentation results, POI distribution statistics, and GPS coordinate indexes) is saved to a geographic context memory database. All data is indexed using image IDs and GPS coordinates.

[0023] Furthermore, the standardized data is stored in the system database, including: basic image information (ID, path, coordinates, upload time); semantic segmentation statistics (proportion of various elements); and five-dimensional rating data (rideability, accessibility, safety, comfort, and enjoyment).

[0024] Furthermore, the specific process of step 2 is as follows: Step 2.1 Transforms heterogeneous data into semantic descriptions through data embedding and rule transformation. The system first establishes a unified data model for the heterogeneous input data. After receiving the raw numerical values, the system performs necessary preprocessing and standardization, such as unit unification, missing value imputation, and boundary pruning, so as to enable calculations under the same rule system.

[0025] The heterogeneous data includes, but is not limited to, the proportion of scene elements based on semantic segmentation (such as the proportion of greenery, buildings, roads, sky, etc.).

[0026] In the rule transformation section, a set of configurable grading thresholds and judgment logic are preset for each type of key element. After judgment, the grading results of each element are obtained. The grading thresholds are used to define the boundaries of the grading, and the judgment logic is used to judge the element values ​​and complete the transformation from values ​​to grades.

[0027] Taking green coverage as an example, setting a high ,middle ,Low Three levels; taking building enclosure as an example, considering both building proportion and sky proportion, "tall buildings + low sky" is used to determine "closely enclosed," "medium buildings + medium sky" is used to determine "moderately enclosed," and otherwise it is determined as "open space." Threshold These values ​​are not fixed constants, but are determined based on historical samples, expert experience, or statistical distributions, and can be flexibly updated in the configuration file.

[0028] The above-mentioned conversion from numerical values ​​to levels is implemented in the system's threshold grading engine using conditional judgments or logical expressions. To avoid over-reliance on a specific implementation language, this invention expresses the algorithm through rule descriptions and parameterized configurations, ensuring that the algorithm is portable and scalable.

[0029] When generating semantic descriptions, the system follows the set priorities and combination logic to string together the hierarchical results of each element into natural language segments.

[0030] The priority settings reflect the semantic hierarchy. For example, the degree of greening usually takes precedence over the degree of building enclosure, then the road ratio, and finally, other elements such as vehicles, pedestrians, and fences are selectively added based on their prominence.

[0031] During the description process, the system adjusts the language emphasis according to the current target dimension (safety, comfort, rideability, accessibility, and enjoyment): if it is in the "safety" dimension, it strengthens the separation facilities and risk warnings; if it is in the "comfort" dimension, it highlights the greenery and spatial experience; if it is in the "rideability" dimension, it focuses on road conditions and the completeness of facilities.

[0032] To ensure the integrity and consistency of the output text, a description integrity check is performed after the semantic description is generated: the description is checked to see if it contains at least basic environmental elements (such as greenery, buildings or roads), safety information (such as separation facilities or risk warnings), and spatial information (such as enclosure or openness); if missing, supplementary rules or default short sentences are triggered.

[0033] For example, a combination of good greening and enclosure can improve the overall environmental experience, while a high safety weight and moderate greening can further enhance the perception of pleasure.

[0034] Finally, to ensure the reliability and traceability of semantic transformation, the input data undergoes a quality assessment, generating a quality score of 0 to 1 based on factors such as the completeness of key fields, the up-to-dateness of the sampling time, and the reliability of the data source. If the score is too low, the system can output a "data missing / insufficient quality" message or explicitly mark "uncertainty" in the description. All thresholds, weights, templates, and rules can be versioned through external configuration to support rapid migration and parameter tuning for different cities, projects, or evaluation dimensions.

[0035] Step 2.2 Structured prompt word assembly This step transforms the semantic segmentation data and geographic information data obtained in step 2.1 into prompt words that can be directly understood and output standardized results by a multimodal large model, according to preset thresholds and language rules. The assembly order is: role and task definition, data embedding and annotation, cognitive guidance injection, and output specification definition. Details are as follows: Data embedding and annotation: Standardize numerical or categorical indicators such as semantic segmentation ratio; set clear thresholds for key indicators and perform hierarchical mapping to generate corresponding semantic tag phrases; concatenate semantic phrases into natural language descriptions according to preset priority order, and retain explicit quantitative information such as original percentages and weights in parentheses; then select some quantitative values ​​to calculate the composite effect value according to a specific composite effect formula, and use it as a comprehensive quality indicator embedding prompt words.

[0036] Cognitive guidance prompt construction includes the following parts: Character setting: The first paragraph of the prompt specifies the professional or user identity of the large model; Task Framework: Clearly define the comparison between Scenario A and Scenario B across target evaluation dimensions, and list judgment principles such as data-driven approach, professional standards, user orientation, and comprehensive balance; where Scenario A and Scenario B are the scenarios to be compared. The thought chain design stipulates that the large model should output the reasoning process in four logical steps: "data interpretation → dimensional analysis → comparative evaluation → comprehensive judgment". Cognitive guidance: Inject the key points of attention for each evaluation dimension; Example guidance: Provides typical cases of high, medium and low scores and tie-breaking rules to constrain model scoring and conclusion standards.

[0037] Output specification definition and format control: A strict structural template is embedded in the prompt words, including fields such as winner information, judgment reason, key differences, and scenario analysis; the quality requirements of each field are clearly explained, and the model is prompted to output only according to the template; the backend performs regular expression validation on the model output, and if it does not conform to the specification, a secondary formatting request is triggered.

[0038] Below is an example of assembly prompts with security as the evaluation dimension: (Role creation): You are a professional urban planner and traffic safety expert.

[0039] (Task Framework) Now you need to compare and evaluate two street view scenarios in terms of security. Based entirely on the provided semantic analysis data, and in conjunction with professional standards and practical usage needs, determine which scenario performs better in terms of security. Your analysis must adhere to the principles of data-driven approach, professional standards, user orientation, and comprehensive balance.

[0040] (Mindset Design) [Data Interpretation] First, interpret the semantic data and safety facility information of Scenario A and Scenario B as a whole, understand the proportion of elements such as greenery, buildings, roads, vehicles, and pedestrians, as well as the type, level, and weight of motorized and non-motorized vehicle separation, and combine this with environmental parameters such as traffic volume and visibility conditions. [Dimensional Analysis] Then, reason around the core elements of safety, focusing on the physical protection effect of separation facilities, traffic conflict risks, and potential problems such as blind spots. [Comparative Evaluation] Next, compare the differences between the two scenarios in the above key points, identify the decisive factors, and reasonably weigh their importance. [Comprehensive Judgment] Finally, draw a conclusion under strict professional standards, give a clear judgment of winner or draw, indicate the confidence level, and present your reasoning process in its entirety.

[0041] The following is a data summary: Scenario A: High green coverage (32.4%), moderate enclosure, road coverage 21.7%, high vehicle volume (6.8%), active pedestrian traffic (4.2%). Motor vehicles and non-motor vehicles are separated by green belts, highest safety level (weight 1.0), lane width 3.0 m, moderate traffic volume.

[0042] Scenario B: Moderate green coverage (18.6%), open space, road coverage 28.3%, high vehicle volume (7.9%), low pedestrian volume (2.1%). Motor vehicle and non-motor vehicle separation is marked, basic safety level (weight 0.5), lane width 2.6 m, high traffic volume.

[0043] (Cognitive Guidance) In the cognitive guidance of safety assessment, the environmental memory layer should focus on the degree of physical protection of the separation of motor vehicles and non-motor vehicles; the spatial memory layer needs to analyze traffic conflict points and safe sight distances; and the cognitive network layer should assess the effectiveness of safety facilities in preventing risks.

[0044] (Example Guide) Please only output the results according to the following JSON template, and do not include any text other than JSON: { "winner": "A" | "B" | "tie", "confidence": 0.0-1.0, "reasoning": "A detailed professional analysis process, including a specific comparison of the two scenarios," "key_differences": [ "Key Difference 1", "Key Difference Point 2" Key Difference Point 3 ], "scene_analysis": { "scene_a": "Detailed analysis of scene A", "scene_b": "Detailed analysis of scene B" }, "dimension_score": { "scene_a": 1-5, "scene_b": 1-5 } } The winning side can only choose A or B; the confidence level must match the certainty of the reasoning; the reasons must be fully presented in contrast to the reasoning; the key differences must be specific and verifiable; the scenario analysis should describe two scenarios separately; the dimension scores should use a 1-5 point scale and reflect the subtle differences in professional judgment.

[0045] High-scoring case: Physical separation of green belts, complete separation of motor vehicles and non-motor vehicles, good visibility, and no blind spots; Mid-section case: Guardrail separation provides basic safety, but there are a few points of traffic conflict; Low-scoring cases: lack of physical separation or mixed traffic, posing obvious safety hazards.

[0046] Furthermore, in step 3, the dynamic closed-loop system of the ELO algorithm and the large model includes an ELO scoring system and a large model evaluation module. By calling the multimodal large model API, specific interpretability evaluation tasks are performed. The ELO scoring system calculates scores and performs pair optimization based on the win / loss results of the large model, and terminates the evaluation based on stability indicators.

[0047] The ELO scoring system, based on large-scale model evaluation results and historical scoring data, uses the ELO rating algorithm for dynamic scoring calculation and real-time updates of the five-dimensional scoring matrix. The ELO scoring system integrates an intelligent pairing algorithm, which dynamically generates optimal pairing combinations through score gap control, comparative demand calculation, and deduplication mechanisms, improving evaluation efficiency and accelerating convergence.

[0048] The large model evaluation module is responsible for performing specific intelligent evaluation tasks. The input is the paired street view images and personalized prompts selected from the ELO system. It performs pairing comparison and evaluation on various dimensions by calling the multimodal large model API, and outputs the comparison results (including the comparison results, reasoning process and result interpretation) using structured JSON response.

[0049] Specifically, the scoring process is as follows: Step 3.1: Initialize ELO scores; The initial ELO score for all street view images was set to 1000 across five evaluation dimensions, and the K-factor was set to 60.

[0050] Step 3.2: Intelligent pairing selection, generating the optimal pairing through multi-factor optimization; The pairing priority scoring function is defined as follows: in, The formula for calculating the score similarity is as follows: in, , These are the current ratings for the street view images corresponding to Scene A and Scene B, respectively. and These represent the maximum and minimum ELO scores for all objects in the real-time scoring results.

[0051] To compare demand: in This represents the current comparison count. To minimize the number of comparisons; To avoid heavy factors: in This represents the time interval since the last comparison.

[0052] , , As a weighting parameter, in an example, it is set to... .

[0053] Step 3.3: Evaluate and pair multimodal large models, and provide standardized evaluation results; Call the multimodal large model to perform paired comparison evaluation. The input includes structured cue words, two base64 encoded street view images, and evaluation dimension labels. Large model configuration parameters: temperature=0.7, topK=32, topP=0.9, maxOutputTokens=1024.

[0054] The process of analyzing and processing structured evaluation results is as follows: A standardized JSON response processing mechanism is adopted to ensure the accurate parsing and effective utilization of large model evaluation results. This system includes the following core technical components: (1) Standardized design of JSON response structure Define a unified JSON response format specification; the output of a large model should include four essential fields: { "winner": "A" | "B" | "tie", "confidence": 0.0-1.0, "reasoning": "detailed reasoning process" "key_differences": ["Key Difference 1", "Key Difference 2", ...] } The winner field represents the comparison result, the confidence field represents the confidence level, the reasoning field contains the complete logical reasoning chain, and the key_differences field lists the key differences.

[0055] (2) Multi-level response analysis algorithm Format validation algorithm: Performs format integrity checks on the JSON response, verifying the existence of necessary fields, the correctness of data types, and the integrity of necessary fields.

[0056] (3) Standardization of evaluation results ① Numerical processing of win / loss results: Convert the comparison results in string form into the numerical format required by the ELO algorithm: If winner = "A", then score_A = 1.0, score_B = 0.0; if winner = "tie", then score_A = 0.5, score_B = 0.5.

[0057] ② Confidence-weighted processing: The evaluation results are weighted and adjusted based on the confidence level output by the model. AdjustedScore = BaseScore × confidence + 0.5 × (1 - confidence) That is, when the confidence level is low, the result is adjusted towards a tie to reduce the impact of uncertain evaluation.

[0058] Step 3.4 Update the ELO score based on the evaluation results (calculate the expected win rate and update the existing score). The expected win rate is calculated using the standard ELO expected value formula, as follows: , in, , These are the current ratings for the street view images corresponding to Scene A and Scene B, respectively. The ELO parameter can be dynamically adjusted later to adapt to different evaluation scenarios.

[0059] The ELO rating system is used for dynamic scoring. The score update formula is as follows: in The actual score is 1 point for a win, 0 points for a loss, and 0.5 points for a draw.

[0060] Step 3.5 Calculate the stability index. Repeat steps 3.2 to 3.4 until the stability meets the threshold, thus completing the evaluation of all street scenes.

[0061] In this invention, whether a ranking can be considered "stable" no longer depends on the final ELO ranking result or a single statistic, but is determined by a combination of two process indicators. The specific calculation process of the stability indicator is as follows: Step 3.5.1: Calculate the stability score to evaluate whether the fluctuation of the score in the recent period has been small enough; Specifically, firstly, for each evaluation object i, its ELO value is established according to its time series record. To capture "recent" behavior, the N most recent ELO updates are selected (N is a configurable parameter, typically set to 8-15, and can be adaptively adjusted according to data size and update frequency), with the most recent N being the most recent ELO updates. Taking this as an example, calculate the change in ELO between two consecutive intervals. : Using the variance method, volatility is defined. .

[0062] Subsequently, the volatility is mapped to a stability score range of 0 to 1 using a preset benchmark constant. The denominator is then set. (e.g., 15000), calculate the stability score for each time point: .

[0063] For all objects The arithmetic mean is used to obtain the global stability score. in, This represents the total number of entities evaluated in this assessment.

[0064] Step 3.5.2: Calculate the comparison coverage, that is, whether each sample is fully and evenly involved in the comparison.

[0065] To ensure that the judgment of "stability" is based on sufficient data, this invention introduces comparison coverage as a second core indicator. A minimum threshold for the number of comparisons per object is set. This threshold can be determined based on the task size, quality requirements, and historical experience: common practices include a fixed lower limit (e.g., at least 6 times), or an adaptive approach that varies with the number of objects (e.g., ...). ), or defined through simulation and prior experience. The comparison function. For any object i, the actual number of comparisons is denoted as . This yields object-level coverage. Finally, the coverage of all objects is averaged to obtain the global coverage score: Step 3.5.3: Determine whether the ranking is stable based on the stability score and comparison coverage. Two strategies can be adopted for comprehensive judgment: The first method is to directly apply a linear weighting to the "stability score" and the "coverage score": ,in Typically, a value of 0.6 to 0.8 is used to ensure that the role of stability itself has a higher weight than coverage, but coverage is still considered a necessary constraint on credibility.

[0066] The second approach is to set tiered thresholds: when coverage falls below a certain threshold... When the score is 0.6, it is directly judged as "unstable or insufficient data," and there is no need to give a higher level of stability conclusion at this time; only when the coverage meets the requirements will the stability score S be used to proceed to the subsequent level classification. The level thresholds for stability (or comprehensive score) can be referenced as follows: ≥0.85 is judged as "very stable," 0.70~0.85 is "stable," 0.60~0.70 is "relatively stable," 0.50~0.60 is "moderately stable," 0.30~0.50 is "not stable enough," and less than 0.30 is "unstable." These dividing points can be determined based on historical backtesting data (minimizing parameters for stable / unstable samples with real labels), or by combining the business's requirements for error tolerance (e.g., allowing the top 5% to swap rankings) to infer the critical value of stability. In addition, these thresholds can be dynamically shifted or scaled according to the number of objects, noise level, and scoring parameters, and the results can be adjusted accordingly. Key parameters are externalized as configuration items to support online A / B testing and phased parameter tuning.

[0067] In addition, the system provides visualization and auditing interfaces, outputting time series of stability and coverage, allowing managers to intuitively see the evolution of indicators; it also reserves a manual intervention entry point to adjust thresholds, forcibly stop or continue sampling when necessary.

[0068] Example 1: Implementation of the data fusion processing module (establishing a geographic context database) (Step 1); 1. The characteristics of the basic input data in this example include: Street View Image Data: Panoramic street view images are collected using the Street View API, including metadata such as latitude and longitude, shooting time, and orientation angle. Batch uploading of street view images automatically generates unique identifiers using UUIDs. Geographic information data: POI data should include two core fields: a functional description and the corresponding geographic latitude and longitude coordinates. Functional types are classified according to urban planning standards. Semantic segmentation processing: A convolutional neural network-based semantic segmentation model is used to perform semantic segmentation on street view images, extracting environmental elements and calculating the pixel proportion of each type of element. This needs to cover the main environmental elements of the urban street view.

[0069] 2. Standardized data conversion and processing process A pre-trained convolutional neural network model is used to perform semantic segmentation on street view images, identify the pixel proportion of each feature element, calculate the POI density within the buffer area through spatial analysis algorithms, and calculate the comprehensive functional index according to urban planning weights.

[0070] The output of the above processing (including semantic segmentation results, POI distribution statistics, and GPS coordinate indexes) is saved to the geospatial memory database. All data is indexed using image IDs and GPS coordinates.

[0071] The standardized data is stored in the context database: basic image information (ID, path, coordinates, upload time); semantic segmentation statistics (proportion of various elements), the dominant function of the sampling point buffer, and POI density, etc.

[0072] Example 2: Specific implementation of structured prompt word assembly; The cue word engineering process converts standardized data from the geographic context database into natural language cue words that conform to the cognitive patterns of large models, including the implementation of two core sub-steps: The specific implementation of data embedding and rule transformation (corresponding to step 2.1) The system first establishes a unified data model for the input heterogeneous data. After receiving the raw values, the system performs necessary preprocessing and standardization, such as unit unification, missing value imputation, and boundary pruning, so as to perform calculations under the same rule system.

[0073] Threshold grading and determination The system pre-sets a set of configurable hierarchical thresholds and judgment logic for each type of key element. Taking green coverage as an example, a high threshold is set... ,middle ,Low Three levels. Threshold The parameters are determined based on historical samples, expert experience, or statistical distribution, and can be flexibly updated in the configuration file.

[0074] Semantic description generation When generating semantic descriptions, the system follows predetermined priorities and combination logic, stringing together the hierarchical results of each element into natural language segments. The priority settings reflect the semantic hierarchy; for example, the degree of greening usually takes precedence over the degree of building enclosure, then the road proportion, and finally, other elements such as vehicles, pedestrians, and fences are selectively added based on their prominence.

[0075] During the description process, the system will automatically adjust the language emphasis according to the current target dimension (safety, comfort, rideability, accessibility, enjoyment, etc.): if it is in the "safety" dimension, it will strengthen the separation facilities and risk warnings; if it is in the "comfort" dimension, it will highlight the greenery and spatial experience; if it is in the "rideability" dimension, it will focus on road conditions and the completeness of facilities.

[0076] Treatment of combined effects This technology allows for weighted superposition of relationships between multiple elements. For example, a combination of good greening and enclosure may improve the overall environmental experience, while a high safety weight and moderate greening can further enhance the perceived pleasantness. The system calculates the composite effect score by defining weights, element values ​​F_i, and correlation coefficients R_{ij} (such as "greening × enclosure" or "safety × greening"), and truncates it to the [0,1] interval.

[0077] The specific implementation of structured prompt word assembly (corresponding to step 2.2) This step transforms the semantic segmentation data, security facility identification data, and geographic information data obtained in step 2.1 into prompt words that can be directly understood and output standardized results by a multimodal large model, according to preset thresholds and language rules. The assembly order is: role and task definition, data embedding and annotation, cognitive guidance injection, and output specification definition.

[0078] Role and task definition implementation The prompt should specify the professional or user identity of the large model in the first paragraph; it should clearly state that scenario A and scenario B need to be compared in terms of target evaluation dimensions, and list judgment principles such as data-driven, professional standards, user orientation and comprehensive balance.

[0079] Data embedding and annotation implementation The numerical or categorical indicators such as semantic segmentation ratio are standardized; clear thresholds are set for key indicators and hierarchical mapping is performed to generate corresponding semantic tag phrases; the semantic phrases are concatenated into natural language descriptions according to a preset priority order, and the original percentage, weight and other explicit quantitative information are retained in parentheses.

[0080] Cognitive guidance construction and implementation The large model is designed to output a reasoning process in four logical steps: "data interpretation → dimensional analysis → comparative evaluation → comprehensive judgment". The model is designed to inject the key points of each evaluation dimension. The model provides typical cases of high, medium and low scores and rules for determining a tie, which constrain the model's scoring and conclusion standards.

[0081] Output specification definition implementation A strict structural template is embedded in the prompt, which includes fields such as the winner's information, the reason for the judgment, key differences, and scenario analysis; the quality requirements for each field are clearly explained, and the model is prompted to output only according to the template; the backend performs regular expression validation on the model output, and if it does not meet the specifications, a secondary formatting request is triggered.

[0082] Example 3: Implementation of ELO algorithm and dynamic closed-loop system of large model (corresponding to step 3); Step 3.1: ELO score initialization implementation This embodiment uses the ELO scoring system for perceptual evaluation and scoring. Details are as follows: In terms of five evaluation dimensions, the ELO score of each uploaded street view image is initialized to 1000 points, the K-factor is set to 60, and the expected win rate is calculated using the standard ELO expected value formula: , in, , These are the current ratings for the street view images corresponding to Scene A and Scene B, respectively. The ELO parameter can be dynamically adjusted later to adapt to different evaluation scenarios.

[0083] For example: Parsing the JSON result shows that A wins; obtaining the current ELO scores: A has 1205 points, B has 1195 points; calculating the expected score: Update the score using K=60: New score point.

[0084] Step 3.2 Implementation of the intelligent pairing algorithm The intelligent pairing algorithm achieves optimal pairing generation through a multi-factor optimization mechanism, thereby improving the overall evaluation speed and efficiency. This is extremely necessary in large-scale production applications and can save a significant amount of token call costs.

[0085] The pairing priority scoring function is defined as follows: The parameters are defined as follows: For score similarity: To compare demand: in This represents the current comparison count. To minimize the number of comparisons; To avoid heavy factors: in This represents the time interval since the last comparison.

[0086] , , As a weighting parameter, in an example, it is set to... Among them, the introduction This effectively avoids redundant comparisons in the short term. The weighting parameters have been optimized through extensive experiments to ensure a reasonable balance of the effects of each factor.

[0087] Step 3.3: Implementation of Paired Evaluation for Multimodal Large Models Multimodal large model pairing evaluation is the core execution component of the system, responsible for calling the multimodal large model to perform specific pairing comparison evaluation tasks. The system input includes three key components: structured cue words, two base64-encoded street view images, and evaluation dimension identifiers. To ensure the stability and reliability of the evaluation results, the system employs optimized model configuration parameters: temperature is set to 0.7, maintaining reasonable consistency while ensuring output diversity; topK is set to 32, topP to 0.9, and maxOutputTokens to 1024. These parameter settings have been verified through extensive experiments, achieving an optimal balance between evaluation quality and processing efficiency.

[0088] The system employs a standardized JSON response processing mechanism to ensure the accurate parsing and effective utilization of large-scale model evaluation results. The standardized JSON response structure includes four essential fields: the `winner` field represents the comparison result (values ​​can be "A", "B", or "tie"); the `confidence` field represents the confidence level of the judgment (value range 0.0-1.0); the `reasoning` field contains a complete logical reasoning chain, providing a detailed analysis of the evaluation process; and the `key_differences` field lists key differences, helping to understand the basis for the evaluation decisions. This structured output format not only facilitates subsequent processing but also improves the interpretability of the evaluation results.

[0089] Standardizing the evaluation results is a crucial step in ensuring the system's proper functioning. The system first quantifies the win / loss results, converting the string-based comparison results into the numerical format required by the ELO algorithm: when winner = "A", score_A = 1.0, score_B = 0.0; when winner = "B", score_A = 0.0, score_B = 1.0; when winner = "tie", score_A = 0.5, score_B = 0.5. Furthermore, the system implements a confidence-weighted processing mechanism, adjusting the evaluation results based on the confidence level of the model output: AdjustedScore = BaseScore × confidence + 0.5 × (1 - confidence). This mechanism is designed to adjust the result towards a tie when the confidence level is low, thereby reducing the impact of uncertain evaluations on the final result.

[0090] Figure 2The results of the large-scale model evaluation are shown for two images. Image 1 is a street view image corresponding to scene A, and image 2 is a street view image corresponding to scene B.

[0091] Step 3.4: Implementation of Real-time ELO Score Updates The rating update process uses transactional operations to ensure data consistency. 1. Get Current Score: Retrieves the current ELO scores of two images in the specified dimension from the database. 2. Calculate Expected Win Rate: Calculate the expected win rate for both players using the standard ELO formula. 3. Apply the K-factor: Calculate the score change based on the actual results (1 / 0 / 0.5) and the expected win rate. 4. Persistent Update: Write the new score back to the database and record the comparison history at the same time.

[0092] Tie Resolution: When the AI ​​cannot determine a clear winner, the system sets the winner's ID to null, and both players receive an actual score of 0.5 points.

[0093] Specific example: Images 1 and 2 are street view images corresponding to scene A and scene B, respectively. The expected score of image 1 is equal to 1 divided by (1 plus 10 of (image 2 score minus image 1 score) divided by 400), and the expected score of image 2 is 1 minus the expected score of image 1. The new score of image 1 is equal to the original score plus the K factor multiplied by (actual result minus expected score), and the new score of image 2 is calculated in a similar way.

[0094] Step 3.5 Implementation of Stability Index Calculation The calculation of stability metrics is a crucial step in ensuring evaluation quality and setting termination conditions. The system repeats steps 3.2 to 3.4 until stability meets a preset threshold, ultimately completing the evaluation of all street scenes. The system no longer relies on the final ELO ranking result or a single statistic, but instead makes a comprehensive judgment based on two process metrics: stability itself and comparative coverage.

[0095] The stability index is calculated by first establishing a time-series record of ELO value changes for each evaluated object. To accurately capture "recent" behavioral characteristics, the system selects the most recent N ELO updates (N is a configurable parameter, typically set to 8-15, and can be adaptively adjusted according to data size and update frequency) and calculates the change in ELO score between two consecutive updates. The system then maps this volatility to a stability score range of 0-1 using a preset baseline constant.

[0096] Comparison coverage, as the second core indicator, ensures that stability judgments are based on sufficient data. The system sets a minimum comparison threshold for a single object, which can be determined based on task scale, quality requirements, and historical experience. For any given object, the actual number of comparisons is recorded as follows: This allows us to calculate object-level coverage. The global coverage score is obtained by averaging the coverage of all objects.

[0097] The threshold for stability level is set according to the following range: when the overall score is ≥0.85, it can be basically judged as "stable". When both the global coverage value and the stability level reach the set value, the system terminates the evaluation.

[0098] The main terms in this invention are defined as follows: Geographic Context Database: Used to integrate and manage multi-source heterogeneous urban street view data.

[0099] Prompt word engineering system: An intelligent prompt word generation system designed based on the Anthropic thinking protocol, which converts structured data into natural language input that conforms to the cognitive patterns of large models.

[0100] Dynamic closed-loop control system: An adaptive control mechanism that realizes intelligent collaboration between the ELO scoring algorithm and the multimodal large model, and achieves automated management of the evaluation process through dynamic pairing optimization and real-time feedback adjustment.

[0101] ELO rating algorithm: derived from the dynamic rating system of chess competitions, which dynamically adjusts the participant's score through pairwise comparison results, and is used in this invention for relative quality assessment of street view images.

[0102] Semantic segmentation technology: a pixel-level classification technique in the field of computer vision, which classifies each pixel in an image into a specific semantic category. In this invention, it is used to extract features of street scene environmental elements.

[0103] Anthropic Mind Protocol: A structured AI interaction design framework that improves the reasoning ability and output quality of AI systems through elements such as role setting, task framework, and mind chain design.

[0104] Point of Interest (POI): In a geographic information system, a POI represents a geographical location that has a specific function or service, such as a shop, school, or hospital.

[0105] Geographic Context Layer: Primarily stores objective environmental characteristics and basic factual information, professional judgment rules and evaluation criteria related to specific scenarios, and maintains geographic location relationships and spatial index information.

[0106] The above description is merely a description of preferred embodiments of this application and is not intended to limit the scope of this application in any way. Any changes or modifications made by those skilled in the art based on the above-disclosed technical content should be considered as equivalent and valid embodiments and fall within the scope of protection of the technical solution of this application.

Claims

1. A method for subjective perception evaluation of streets supported by a multimodal large model and street view imagery, characterized in that, Includes the following steps: Step 1: Convert the heterogeneous raw data from multiple sources into standardized data and establish a geographic context memory database; Step 2: Convert the technical data obtained from computer vision into natural language descriptions that the large model can understand; then, according to the thought chain format, assemble the converted data and prompt word templates, and provide them to Step 3 along with the images; Step 3: Construct a dynamic closed-loop system for the ELO algorithm and the large model, call the multimodal large model API, execute specific interpretable evaluation tasks, output in a fixed format, the ELO scoring system calculates scores and performs pair optimization based on the win and loss results of the large model, and automatically terminates the evaluation according to the stability criterion, thus completing the subjective perception evaluation of the street.

2. The method for subjective perception evaluation of streets supported by a multimodal large model and street view imagery as described in claim 1, characterized in that, In step 1, the multi-source heterogeneous raw data includes: street view image data, POI geographic information data, and semantic segmentation data; wherein, the semantic segmentation data is preprocessed data obtained by feature extraction of street view image data through a convolutional neural network model, and the POI geographic information data is obtained through an open-source map website; The output standardized technical data is saved in a geographic context memory database, and all data is indexed by image ID and GPS coordinates.

3. The method for subjective perception evaluation of streets supported by a multimodal large model and street view imagery as described in claim 1, characterized in that, The specific process in step 2 is as follows: Step 2.1: Transform heterogeneous data into semantic descriptions through data embedding and rule transformation; First, a unified data model is established for the heterogeneous input data; after receiving the raw values, preprocessing and standardization are performed so that they can be computed under the same rule system. The heterogeneous data includes, but is not limited to, the proportion of scene elements based on semantic segmentation; In the rule conversion section, a set of configurable hierarchical thresholds and judgment logic are preset for each type of key element. After judgment, the hierarchical results of each element are obtained. The hierarchical thresholds are used to define the boundaries of the hierarchical classification, and the judgment logic is used to judge the element values ​​to complete the conversion from values ​​to levels. When generating semantic descriptions, the system follows the set priorities and combination logic to string together the hierarchical results of each element into natural language segments; the priority setting reflects the semantic hierarchy, and during the description process, the system adjusts the language focus according to the current target dimension. To ensure the completeness and consistency of the output text, a description completeness check is performed after the semantic description is generated; if it is missing, a supplementary rule or a default short sentence is triggered. Finally, to ensure the reliability and traceability of semantic transformation, the input data is evaluated for quality, and a quality score of 0 to 1 is generated based on the completeness of key fields, the up-to-dateness of sampling time, and the credibility of data source. If the score is too low, the system can output a "data missing / inadequate quality" message, or explicitly mark "uncertainty" in the description; Step 2.2: Structured prompt word assembly; The assembly sequence is as follows: role and task definition, data embedding and annotation, cognitive guidance injection, and output specification definition; specifically as follows: Data embedding and annotation: Standardize the semantic segmentation ratio or category indicators; set clear thresholds for key indicators and perform hierarchical mapping to generate corresponding semantic tag phrases; Semantic phrases are concatenated into natural language descriptions according to a preset priority order, and the original percentage and weight explicit quantitative information are retained in parentheses. Then, some quantitative values ​​are selected to calculate the composite effect value according to a specific composite effect formula, which is then used as a comprehensive quality indicator to embed prompt words. Cognitive guidance prompt construction includes the following parts: Character setting: The first paragraph of the prompt specifies the professional or user identity of the large model; Task framework: Clearly define the comparison between scenario A and scenario B across target evaluation dimensions, and list data-driven, professional standards, user-oriented, and comprehensive balanced judgment principles; The thought chain design specifies that the large model should follow a four-step logical output reasoning process: "data interpretation → dimensional analysis → comparative evaluation → comprehensive judgment". Cognitive guidance: Inject the key points of attention for each evaluation dimension; Example guidance: Provides typical cases of high, medium, and low scores, as well as tie-breaking rules, to constrain the model's scoring and conclusion criteria; Output specification definition and format control: A strict structural template is embedded in the prompt words, including the winner's information, the reason for the judgment, key differences, and scenario analysis fields; the quality requirements for each field are clearly explained, and the model is prompted to output only according to the template; the backend performs regular expression validation on the model output, and if it does not conform to the specification, a secondary formatting request is triggered.

4. The method for subjective perception evaluation of streets supported by a multimodal large model and street view imagery as described in claim 1, characterized in that, In step 3, the dynamic closed-loop system of the ELO algorithm and the large model includes an ELO scoring system and a large model evaluation module; by calling the multimodal large model API, specific interpretable evaluation tasks are performed. The ELO scoring system realizes the scoring calculation and pairing optimization functions based on the win and loss results of the large model, and terminates the evaluation based on the stability index. Among them, the ELO scoring system: based on the evaluation results of the large model and historical scoring data, it uses the ELO rating algorithm to perform dynamic scoring calculation and updates the five-dimensional scoring matrix in real time. The ELO scoring system integrates an intelligent pairing algorithm, which dynamically generates the optimal pairing combination through score gap control, comparison demand calculation, and deduplication mechanism, thereby improving evaluation efficiency and accelerating convergence. Large Model Evaluation Module: Responsible for performing specific intelligent evaluation tasks. The input consists of paired street view images and personalized prompts selected from the ELO system. It performs pairing comparison and evaluation across various dimensions by calling the multimodal large model API.

5. The method for subjective perception evaluation of streets supported by a multimodal large model and street view imagery as described in claim 4, characterized in that, The scoring process is as follows: Step 3.1: Initialize ELO scores; The initial ELO score of all street view images was set to 1000 across five evaluation dimensions, and the K factor was set to 60. Step 3.2: Intelligent pairing selection, generating the optimal pairing through multi-factor optimization; Step 3.3: Evaluate and pair multimodal large models, and provide standardized evaluation results; Step 3.4: Update the ELO score based on the evaluation results; The expected win rate is calculated using the standard ELO expected value formula, as follows: , in, , These are the current ratings for the street view images corresponding to Scene A and Scene B, respectively. The ELO rating system is used for dynamic scoring; the score update formula is as follows: in The actual score is 1 point for a win, 0 points for a loss, and 0.5 points for a draw. Step 3.5 Calculate the stability index. Repeat steps 3.2 to 3.4 until the stability meets the threshold, thus completing the evaluation of all street scenes.

6. The method for subjective perception evaluation of streets supported by a multimodal large model and street view imagery as described in claim 5, characterized in that, The pairing priority scoring function is defined as follows: in, The formula for calculating the score similarity is as follows: in, , These are the current ratings for the street view images corresponding to Scene A and Scene B, respectively. and These represent the maximum and minimum ELO scores for all objects in the real-time scoring results; To compare demand: in This represents the current comparison count. To minimize the number of comparisons; To avoid heavy factors: in This represents the time interval since the last comparison. , , These are the weight parameters.

7. The method for subjective perception evaluation of streets supported by a multimodal large model and street view imagery as described in claim 5, characterized in that, Step 3.3 specifically involves: The multimodal large model is invoked to perform paired comparison evaluation. The input includes structured prompts, two base64 encoded street view images, and evaluation dimension labels. The large model configuration parameters are: temperature=0.7, topK=32, topP=0.9, maxOutputTokens=1024. The structured evaluation results are analyzed and processed as follows: It employs a standardized JSON response processing mechanism to ensure the accurate parsing and effective utilization of large model evaluation results; it includes the following core technical components: (1) Standardized design of JSON response structure Define a unified JSON response format specification; the output of a large model should include four essential fields: { "winner": "A" | "B" | "tie", "confidence": 0.0-1.0, "reasoning": "detailed reasoning process" "key_differences": ["Key Difference 1", "Key Difference 2", ...] } The winner field represents the comparison result, the confidence field represents the confidence level, the reasoning field contains the complete logical reasoning chain, and the key_differences field lists the key differences. (2) Multi-level response analysis algorithm Format validation algorithm: Perform format integrity checks on the JSON response, verifying the existence of necessary fields and the correctness of data types, and checking the integrity of necessary fields; (3) Standardization of evaluation results ① Numerical processing of win / loss results: Convert the comparison results in string format into the numerical format required by the ELO algorithm: if winner = "A", then score_A = 1.0, score_B = 0.0; if winner = "tie", then score_A = 0.5, score_B = 0.5; ② Confidence-weighted processing: The evaluation results are weighted and adjusted based on the confidence level output by the model. AdjustedScore = BaseScore × confidence + 0.5 × (1 - confidence) That is, when the confidence level is low, the result is adjusted towards a tie to reduce the impact of uncertain evaluation.

8. The method for subjective perception evaluation of streets supported by a multimodal large model and street view imagery as described in claim 5, characterized in that, The specific calculation process for the stability index is as follows: Step 3.5.1: Calculate the stability score, which is used to evaluate the fluctuation of the score in the recent period; First, for each evaluation object i, its ELO value is established according to its time series record. To capture "recent" behavior, the N most recent ELO updates are selected, with the Nth update being the most recent. Taking this as an example, calculate the change in ELO between two consecutive intervals. : Using the variance method, volatility is defined. ; Subsequently, the volatility is mapped to a stability score range of 0 to 1 using a preset benchmark constant; Set the denominator Calculate the stability score at each time point: ; For all objects The arithmetic mean is used to obtain the global stability score. : in, This represents the total number of entities evaluated in this assessment. Step 3.5.2: Calculate the comparison coverage, i.e., whether each sample is fully and evenly involved in the comparison; Introduce a comparison coverage metric: Set a minimum threshold for the number of comparisons per object. For any object i, the actual number of comparisons is denoted as . This yields object-level coverage. Finally, the coverage of all objects is averaged to obtain the global coverage score. Step 3.5.3: Determine whether the ranking is stable based on the stability score and comparison coverage.

Citation Information

Patent Citations

  • Street space quality monitoring, evaluating and early warning method

    CN114331232A

  • Multi-stage content quality evaluation method and system based on large language model

    CN120448618A

  • Element feature measurement method and device for block style and appearance shaping

    CN120495859A

  • Street space quality intelligent evaluation and image generation method and system based on deep learning

    CN120724519A