Large model intelligent analysis and suggestion system and method for urban management field
By combining LoRA fine-tuning and RAG technologies with LangChain tools, the smart city management system has achieved multi-source data integration and intelligent analysis, solving the data silo problem, improving the scientific nature and efficiency of decision-making, and supporting rapid response and resource optimization.
Patent Information
- Application Number
- CN202510789623.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-11-11
AI Technical Summary
Existing smart urban management systems suffer from data silos and fragmented information, resulting in incomplete and untimely display of key data, making it difficult to support efficient decision analysis. Furthermore, a lack of understanding of urban management terminology leads to poor task performance and insufficient scientific rigor and efficiency in decision-making.
Employing LoRA fine-tuning technology and RAG retrieval enhancement generation technology, combined with LangChain's integration of external tools and data analysis tools, and using Mind Tree (ToT) reasoning technology, multi-source data is integrated and classified to construct a structured fine-tuning dataset. This supports real-time cross-departmental data retrieval and automates the analysis process through a timed trigger module.
The model's understanding of professional terminology and business processes in the urban management field has been improved, enabling rapid problem identification, scientific decision-making, and resource allocation. This has shortened response time, improved cross-departmental collaboration efficiency, and reduced manual review costs.
Smart Images

Figure CN120930773A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence and smart city technology, specifically a large-scale intelligent analysis and suggestion system and method for the field of urban management. Background Technology
[0002] Large Language Models (LLMs) are models built using deep learning techniques. Their core typically employs a neural network architecture based on Transformers and their variants to process and generate natural language. These models are trained on massive text datasets to master the structure, patterns, and contextual relationships of language, enabling them to handle language processing tasks such as text generation. LoRA (Low-Rank Adaptation) is an efficient optimization technique for large models. Its core idea is to decompose the parameters of a large model into low-rank matrices and then fine-tune them. This method significantly reduces training time and parameter size while maintaining model performance, making it particularly suitable for fine-tuning specific tasks in resource-constrained environments. LangChain is a development framework for building and deploying applications based on Large Language Models (LLMs). It helps developers design, connect, and extend language model-driven applications more efficiently by providing modular tools and interfaces. The core idea of LangChain is to decompose complex tasks into multiple steps or components and chain these components together to achieve end-to-end task processing. It supports integration with external data sources, tools, and services, such as Retrieval-Augmented Generation (RAG), API calls, and database queries, enabling models to generate more accurate output by combining real-time information and context. LangChain is widely used in question-answering systems, dialogue agents, document analysis, and other scenarios, providing developers with a flexible and powerful toolchain to fully leverage the potential of large language models. RAG (Retrieval-Augmented Generation) is a technique that combines retrieval and generation. By retrieving relevant information from external knowledge bases and combining it with the generative model, it enhances the output quality and accuracy of the model. It performs exceptionally well in tasks such as question answering and dialogue systems, effectively utilizing external knowledge to improve model performance. Chain-of-Thought is a technique that enhances the model's reasoning ability by simulating human thought processes. It guides the model to generate intermediate reasoning steps step by step, helping the model better understand and solve complex problems. It is particularly suitable for tasks requiring logical reasoning and step-by-step thinking, such as mathematical problems and complex question answering.
[0003] Through the use of existing smart urban management systems, it was found that when integrating data from multiple business systems, these systems often face problems such as data silos and information fragmentation. This results in insufficient and untimely display of key data and emergency events, making it difficult to support efficient decision analysis. Furthermore, general-purpose models lack sufficient understanding of urban management terminology (such as policies, regulations, and business processes), leading to poor performance. In practical applications, many systems employ single data analysis methods, lacking in-depth mining and intelligent analysis capabilities of multi-dimensional data. This results in inaccurate summaries of work strengths and weaknesses, failing to provide a scientific basis for adjusting work priorities and allocating resources. Simultaneously, existing systems rely heavily on human experience in resource allocation and decision support, lacking AI-based intelligent analysis capabilities, leading to insufficient scientific rigor and efficiency in decision-making. Summary of the Invention
[0004] The purpose of this invention is to provide a large-scale intelligent analysis and suggestion system and method for the field of urban management, so as to solve the problems of data silos, slow response and insufficient scientific decision-making in existing smart urban management systems mentioned in the background art.
[0005] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows: A large-scale intelligent analysis and suggestion system for urban management includes: The domain adaptation module is used to improve the understanding of professional terminology and business processes in the urban management field by using LoRA fine-tuning technology and RAG retrieval enhancement generation technology. The reasoning analysis module is used to integrate external tools and data analysis tools based on LangChain, and combine thinking tree (ToT) reasoning technology to generate decision suggestions. It includes an agent interaction submodule and a reasoning enhancement submodule. The timed trigger module is used to trigger execution at regular time intervals to automatically start the analysis process.
[0006] A method for intelligent analysis and suggestion of large-scale models in the field of urban management, comprising domain-adaptive enhancement functions and intelligent analysis and suggestion functions: Domain-adaptive enhancements include: multi-source data integration and classification to build structured fine-tuning datasets; LoRA lightweight fine-tuning with progressive training using a course-based learning strategy; and RAG knowledge base enhancement, which synchronously updates policy documents and supports real-time cross-departmental data retrieval. The intelligent analysis and suggestion functions include: external tool invocation, based on LangChain integration of data analysis tools and system reporting interfaces; mind tree (ToT) construction, which decomposes complex tasks into sub-tasks and optimizes decision-making logic; and scheduled analysis and real-time response, which collects data on a regular basis and inputs it into the model to trigger analysis.
[0007] According to the above technical solution, multi-source data integration and classification includes: Extract structured policy documents from government databases and collect cross-departmental business specification documents; The training and validation sets were divided in a 3:7 ratio. A course learning strategy was introduced, and the data was organized hierarchically according to the complexity of the cases, including single-law application scenarios, multi-law intersection scenarios, and extreme weather superimposed scenarios.
[0008] According to the above technical solution, LoRA lightweight fine-tuning includes: High-frequency scene data with a frequency of ≥5% were selected, and back-translation enhancement technology was used for long-tail scene data with a frequency of <1%. First, train for 3 epochs using high-frequency scene data, set the optimizer to AdamW (lr=5e-5, weight_decay=0.01), and the early stopping strategy is to terminate the test if the verification loss does not decrease for 2 consecutive times. When the F1-score is greater than or equal to 0.9 in high-frequency scenarios, continue training by mixing 60% high-frequency data with 40% long-tail data, and dynamically adjust the loss weights to Loss = 0.7 × CE_loss + 0.3 × Focal_loss.
[0009] According to the above technical solution, the enhancement of the RAG knowledge base includes: The API polling task is deployed to incrementally update the knowledge base daily at 03:00, and new case data is pulled via REST API every 15 minutes. The system employs a hybrid retrieval architecture. The BM25 index is built using Elasticsearch and configured with a Chinese word segmenter based on the jieba+ domain dictionary. The vector database uses the Faiss-HNSW structure, and the text encoder uses the text2vec-large model. Deploy an intent recognition model (BiLSTM-CRF) to optimize real-time retrieval.
[0010] According to the above technical solution, external tool invocation includes: Initialize the LangChain agent module and configure the tool call interface (such as PythonAgent or ToolChain). Define the tool call path and integrate data analysis tools such as Pandas and Matplotlib, as well as system reporting interfaces.
[0011] Based on the above technical solution, the construction of a mind tree (ToT) includes: Clearly define the objectives, inputs, and outputs of complex tasks, construct the root node of a mind tree, and associate it with the initial input data; The task is broken down into sub-tasks such as spatiotemporal distribution analysis and suggestion generation. Multiple reasoning paths are generated for each subtask, and the answer with the highest consistency is selected using self-consistency techniques.
[0012] According to the above technical solution, timed analysis and real-time response include: Configure Quartz core components (Scheduler, Job, Trigger) and define the analysis task cycle; Call environmental sensor APIs, sanitation vehicle GPS and other real-time data interfaces to clean noise data and standardize it into JSON format; Standardized data is input into the model to trigger the inference process.
[0013] According to the above technical solution, the agent interaction submodule realizes cross-departmental data scheduling and dynamic optimization of task decision chain based on LangChain, and the reasoning enhancement submodule explicitly models the law enforcement reasoning process through mind tree (ToT) technology to realize causal association and multi-step verification.
[0014] According to the above technical solution, the hybrid retrieval architecture combines the keyword retrieval of BM25 with the semantic retrieval of the vector database, and improves the query understanding ability through the intent recognition model, supporting real-time correlation retrieval of policy documents and case data.
[0015] Compared with the prior art, the present invention has the following beneficial effects: This invention provides a system and method for intelligent analysis and recommendations for large-scale models in the field of urban management. Through LoRA fine-tuning technology and RAG dynamic knowledge retrieval, combined with urban management policy documents and real-time data, the system enhances the model's understanding of professional terminology and business processes. Based on LangChain, it integrates external tools (such as GIS and environmental monitoring APIs) and data analysis tools (Pandas and Matplotlib), and combines ToT (Total Time) reasoning technology to decompose complex tasks into sub-tasks, optimizing decision paths through self-consistent reasoning. Based on the Quartz framework, it dynamically triggers the analysis process, and combined with agent modules to collect data in real time (such as GPS data from sanitation vehicles and sensor data), achieving automated early warning. Attached Figure Description
[0016] Figure 1 This is a schematic diagram of the system structure of the present invention; Figure 2 This is a flowchart of the method of the present invention; Figure 3 This is an activity diagram of the multi-source data integration and classification stage in this invention; Figure 4 This is an activity diagram of the fine-tuning stage in this invention; Figure 5 This is an activity diagram of the RAG knowledge base enhancement phase in this invention; Figure 6 This is an activity diagram of the external tool invocation phase in this invention; Figure 7 This is an activity diagram of the ToT (Mind Tree) construction phase in this invention; Figure 8 This is an activity diagram of the construction phase of the timed triggering module in this invention. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] Example 1 like Figure 1 As shown, a large-scale intelligent analysis and suggestion system for the urban management field includes: The domain adaptation module is used to improve the understanding of professional terminology and business processes in the urban management field by using LoRA fine-tuning technology and RAG retrieval enhancement generation technology. The reasoning analysis module is used to integrate external tools and data analysis tools based on LangChain, and combine thinking tree (ToT) reasoning technology to generate decision suggestions. It includes an agent interaction submodule and a reasoning enhancement submodule. The timed trigger module is used to trigger execution at regular time intervals to automatically start the analysis process.
[0019] This invention provides a system and method for intelligent analysis and recommendations for large-scale models in the field of urban management. Through LoRA fine-tuning technology and RAG dynamic knowledge retrieval, combined with urban management policy documents and real-time data, the system enhances the model's understanding of professional terminology and business processes. Based on LangChain, it integrates external tools (such as GIS and environmental monitoring APIs) and data analysis tools (Pandas and Matplotlib), and combines ToT (Total Time) reasoning technology to decompose complex tasks into sub-tasks, optimizing decision paths through self-consistent reasoning. Based on the Quartz framework, it dynamically triggers the analysis process, and combined with agent modules to collect data in real time (such as GPS data from sanitation vehicles and sensor data), achieving automated early warning.
[0020] Through the implementation of this system, the efficiency of urban management departments has been significantly optimized in multiple aspects. First, the time for problem location has been shortened and the cost of manual review has been saved. The system can quickly identify high-incidence areas through spatiotemporal analysis technology (such as "the frequency of garbage accumulation in a certain street during the morning rush hour has increased by 30%), eliminating the need for manual area-by-area investigation. Second, the scientific nature of decision-making has been improved. The system automatically generates resource allocation suggestions (such as "add 5 garbage trucks to XX area") and policy optimization plans (such as "adjust the emission standards for catering fumes"), reducing trial and error costs.
[0021] Furthermore, real-time response time has been reduced to the minute level. For example, when air quality exceeds standards, the system automatically triggers an alert and suggests "strengthening construction site dust control," significantly improving response speed compared to traditional processes. In cross-domain collaboration, the system enhances collaboration efficiency and avoids redundant surveys by multiple departments through domain correlation analysis (such as "construction site dust leads to increased complaints about surrounding restaurants").
[0022] Example 2 This embodiment also provides a method for intelligent analysis and suggestions based on large-scale models in the field of urban management.
[0023] The preset scenario for this embodiment is the field of smart urban management. The specific task involves periodically inputting real-time data into a large model to trigger model inference. The specific tasks include analyzing the current working situation based on the collected real-time data and knowledge of urban management, and providing improvement suggestions. This embodiment 2 includes the following steps: Step A: Multi-source data integration and classification. Collect data on policies, regulations, historical cases, and business standards in the urban management field to construct a structured, fine-tuned dataset.
[0024] Step B: LoRA Lightweight Fine-tuning. A learning strategy is designed to address local regulations and specific case patterns, progressively training students on high-frequency scenarios first, then moving to longer-tail scenarios. Step C: Enhance the RAG knowledge base. Policy documents are updated daily, and real-time retrieval of cross-departmental data (case management platform) is supported.
[0025] Step D: External Tool Invocation. Based on LangChain, integrate data analysis tools (Pandas, Matplotlib) and system reporting interfaces to help the model perform data analysis and report analysis and recommendations.
[0026] Step E: Constructing the Mind Tree (ToT). Complex tasks (such as "analyzing the operation of the smart city management system") are broken down into sub-tasks (spatiotemporal distribution analysis → suggestion generation), and the optimization decision-making logic is explored through a tree-like path.
[0027] Step F: Scheduled Analysis and Real-time Response. Data is collected periodically and input into the model to trigger model analysis.
[0028] like Figure 2 As shown, this embodiment includes the entire process of the illusion reduction scheme for the large text generation model, comprising the multi-source data integration and classification stage, the lightweight fine-tuning stage, the RAG knowledge base enhancement stage, the external tool invocation stage, the mind tree (ToT) construction stage, and the timed analysis and real-time response stage. In the multi-source data integration and classification stage, this embodiment collects data such as policies, regulations, historical cases, and business specifications in the urban management field to construct a structured fine-tuning dataset for subsequent fine-tuning training. In the lightweight fine-tuning stage, based on the obtained dataset, domain knowledge is learned using the LoRA fine-tuning method. RAG knowledge base enhancement is mainly used to update relevant policy documents in real time and support real-time retrieval of departmental data. In the external tool invocation stage, LangChain is mainly used to integrate relevant data analysis tools and system reporting interfaces to help the model perform data analysis and report analysis and suggestions. The mind tree (ToT) construction stage is mainly used to decompose complex tasks into executable sub-tasks and generate multiple reasoning paths from which the answer with the highest consistency is selected, thereby improving reasoning quality. The scheduled analysis and real-time response phase mainly involves collecting data periodically and inputting it into the model, triggering model analysis, and reporting the results.
[0029] like Figure 3 As shown, step A specifically includes the following steps: Step A1: Domain Data Collection. Extract structured policy documents from government databases (such as the Ministry of Housing and Urban-Rural Development's regulatory database), including core regulations such as the "Urban Management Enforcement Measures." Collect cross-departmental business standard documents (such as the Environmental Protection Bureau's pollution disposal standards and the Market Supervision Bureau's merchant management rules).
[0030] Step A2: Fine-tune dataset construction. Divide the training and validation sets in a 3:7 ratio. Introduce a course learning strategy, organizing the data hierarchically according to case complexity: Level 1: Single regulation application scenario; Level 2: Multiple regulation intersection scenario; Level 3: Extreme weather superimposed scenario.
[0031] like Figure 4 As shown, step B specifically includes the following steps: Step B1: High-frequency scenario data construction. Select case types with a frequency of ≥5% (such as street vending, demolition of illegal buildings) from the fine-tuning dataset in Step A.
[0032] Step B2: Long-tail scene data augmentation. Back-translation augmentation techniques are used for special cases with an occurrence frequency of <1% (such as damage to ancient and famous trees).
[0033] Step B3: Train for 3 epochs using high-frequency scene data. Optimizer: AdamW (lr=5e-5, weight_decay=0.01). Early stopping strategy: Terminate if the loss fails to decrease for two consecutive verification iterations.
[0034] Step B4: Continue training by mixing high-frequency data (60%) and long-tail data (40%). When the F1-score for the high-frequency scenario is ≥0.9, start long-tail training and dynamically adjust the loss weights. Loss = 0.7 × CE_loss + 0.3 × Focal_loss (focusing on long-tail samples).
[0035] like Figure 5 As shown, step C specifically includes the following steps: Step C1: Multi-source data access. Deploy a knowledge base API polling task to perform incremental updates daily at 03:00. Retrieve new case data (fields include case type, handling status, and relevant regulations) every 15 minutes via REST API.
[0036] Step C2: Knowledge Base Construction. A hybrid retrieval architecture is adopted, with the BM25 index built using Elasticsearch and a custom Chinese word segmenter (jieba + domain dictionary) configured. The vector database uses the Faiss-HNSW structure, and the text encoder uses the text2vec-large model (768 dimensions).
[0037] Step C3: Real-time retrieval optimization. To enhance query understanding, deploy an intent recognition model (BiLSTM-CRF).
[0038] like Figure 6 As shown, step D specifically includes the following steps: Step D1: Initialize the LangChain agent module and configure the tool call interface (such as PythonAgent or ToolChain).
[0039] Step D2: Define the tool call path, such as: tools = [ Tool(name="PandasAnalysis", func=pandas_ops, description="Data Cleaning and Statistical Analysis") ].
[0040] like Figure 7 As shown, step E specifically includes the following steps: Step E1: Define the complex task objectives (such as "analyzing the operation of the smart city management system"), define the inputs (real-time data, historical records) and outputs (analysis reports, decision-making suggestions), construct the root node of the mind tree, and link it to the initial input data.
[0041] Step E2: Decompose the subtasks. The first-level subtasks can be decomposed into analyzing data and generating suggestions from dimensions such as time and space.
[0042] Step E3: Generate multiple inference paths for each subtask (such as "Identify morning and evening peak hours → Suggest increasing the frequency of waste collection during peak hours" and "Locate high-incidence areas → Suggest optimizing waste collection routes").
[0043] Step E4: Select the answer with the highest consistency using the self-consistency technique.
[0044] like Figure 8 As shown, step F specifically includes the following steps: Step F1: Task scheduling initialization. Configure Quartz core components (Scheduler, Job, Trigger) and define the analysis task cycle (e.g., hourly execution or dynamic adjustment).
[0045] Step F2: Call the real-time data interface (such as environmental sensor API, sanitation vehicle GPS), clean the noisy data (such as removing sensor outliers), and standardize it into JSON format.
[0046] Step F3: Input the standardized JSON data into the model to trigger the inference process.
[0047] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0048] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A large-scale intelligent analysis and suggestion system for urban management, characterized in that, include: The domain adaptation module is used to improve the understanding of professional terminology and business processes in the urban management field by using LoRA fine-tuning technology and RAG retrieval enhancement generation technology. The reasoning analysis module is used to integrate external tools and data analysis tools based on LangChain, and combine the thinking tree (ToT) reasoning technology to generate decision suggestions. It includes the agent interaction submodule and the reasoning enhancement submodule. The timed trigger module is used to trigger execution at regular time intervals to automatically start the analysis process.
2. A large-scale intelligent analysis and suggestion method for the urban management field, characterized by: Using the intelligent analysis and suggestion system as described in claim 1, the method includes domain-adaptive enhancement functionality and intelligent analysis and suggestion functionality: Domain-adaptive enhancements include: multi-source data integration and classification to build structured fine-tuning datasets; LoRA lightweight fine-tuning with progressive training using a course-based learning strategy; and RAG knowledge base enhancement, which synchronously updates policy documents and supports real-time cross-departmental data retrieval. The intelligent analysis and suggestion functions include: external tool invocation, and integration of data analysis tools and system reporting interfaces based on LangChain; Mind Tree (ToT) construction breaks down complex tasks into sub-tasks and optimizes decision-making logic; Scheduled analysis and real-time response: Data is collected periodically and input into the model to trigger analysis.
3. The intelligent analysis and suggestion method for large-scale models in the field of urban management as described in claim 2, characterized in that: Multi-source data integration and classification includes: Structured policy documents were extracted from government databases, and cross-departmental business specification documents were collected. The training set and validation set were divided into a certain proportion, and a course learning strategy was introduced. The data was organized in layers according to the complexity of the cases, including single law application scenarios, multiple law intersection scenarios, and extreme weather superposition scenarios.
4. The intelligent analysis and suggestion method for large-scale models in the field of urban management as described in claim 3, characterized in that: LoRA lightweight tuning includes: High-frequency scene data with a frequency of ≥5% were selected, and back-translation enhancement technology was used for long-tail scene data with a frequency of <1%. First, train for 3 epochs using high-frequency scene data, set the optimizer to AdamW (lr=5e-5, weight_decay=0.01), and the early stopping strategy is to terminate the test if the verification loss does not decrease for 2 consecutive times. When the F1-score is greater than or equal to 0.9 in high-frequency scenarios, continue training by mixing 60% high-frequency data with 40% long-tail data, and dynamically adjust the loss weights to Loss = 0.7 × CE_loss + 0.3 × Focal_loss.
5. The intelligent analysis and suggestion method for large-scale models in the field of urban management as described in claim 4, characterized in that: Enhancements to the RAG knowledge base include: Deploy an API polling task to incrementally update the knowledge base daily, and pull new case data via REST API every 15 minutes; The system employs a hybrid retrieval architecture. The BM25 index is built using Elasticsearch and configured with a Chinese word segmenter based on the jieba+ domain dictionary. The vector database uses the Faiss-HNSW structure, and the text encoder uses the text2vec-large model. Deploy an intent recognition model (BiLSTM-CRF) to optimize real-time retrieval.
6. The intelligent analysis and suggestion method for large-scale models in the field of urban management as described in claim 5, characterized in that: External tool calls include: Initialize the LangChain proxy module and configure the tool to call the interface; Define the tool call path and integrate data analysis tools and system reporting interfaces.
7. The intelligent analysis and suggestion method for large-scale models in the field of urban management as described in claim 6, characterized in that: Mind Tree (ToT) construction includes: Clearly define the objectives, inputs, and outputs of complex tasks, construct the root node of a mind tree, and associate it with the initial input data; The task is broken down into spatiotemporal distribution analysis and suggestions for generating sub-tasks; Multiple reasoning paths are generated for each subtask, and the answer with the highest consistency is selected using self-consistency techniques.
8. The intelligent analysis and suggestion method for large-scale models in the field of urban management as described in claim 7, characterized in that: Timed analysis and real-time response include: Configure Quartz core components (Scheduler, Job, Trigger) and define the analysis task cycle; Call the environmental sensor API and the real-time GPS data interface of sanitation vehicles to clean the noise data and standardize it into JSON format; Standardized data is input into the model to trigger the inference process.
9. The intelligent analysis and suggestion method for large-scale models in the field of urban management as described in claim 8, characterized in that: The agent interaction submodule uses LangChain to achieve cross-departmental data scheduling and dynamic optimization of task decision chains. The reasoning enhancement submodule uses Mind Tree (ToT) technology to explicitly model the law enforcement reasoning process, realizing causal association and multi-step verification.
10. A method for intelligent analysis and suggestion of large-scale models in the field of urban management according to claim 9, characterized in that: The hybrid retrieval architecture combines BM25's keyword retrieval with the semantic retrieval of the vector database, and enhances query understanding capabilities through an intent recognition model, supporting real-time correlation retrieval of policy documents and case data.