A data collection method and system for urban and rural planning investigation

By combining an adaptive-scale isolated forest model and a convolutional neural network model, the problem of integrating environmental data with resident feedback data was solved, achieving deep integration and feature extraction of multi-source data, and improving the scientificity and accuracy of urban and rural planning.

CN120145295BActive Publication Date: 2026-04-24CHINA THREE GORGES UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA THREE GORGES UNIV
Filing Date
2025-02-12
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

In existing technologies, environmental data and resident feedback data are stored separately, which prevents deep integration, lacks effective data processing and feature extraction methods, and results in insufficient geospatial correction of data, affecting the scientific nature and comprehensiveness of planning decisions.

Method used

An adaptive-scale isolated forest model was constructed for environmental data preprocessing. A smart community connection platform was built to collect resident feedback data. A convolutional neural network model was used to fuse multi-source data and combine it with geographic information for correction, forming a unified multidimensional dataset.

Benefits of technology

Achieving deep integration of multi-source data will enhance the scientific rigor and comprehensiveness of planning decisions, ensure the accuracy and reliability of data, and improve the alignment of planning outcomes with actual needs and their geographical applicability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145295B_ABST
    Figure CN120145295B_ABST
Patent Text Reader

Abstract

The application discloses a kind of data collection method and system for urban and rural planning investigation, the method includes: collecting environmental data;Model preprocessing environmental data is constructed;Platform generates resident feedback data, and environmental data is combined into multi-source data and is stored in database;With neural network model, multi-source data is carried out feature extraction fusion, and forms unified multidimensional dataset.The application realizes the deep fusion of multi-source data, improves the comprehensiveness of planning decision, and breaks the situation of separate storage of environmental data and resident feedback data in traditional technology.This fusion mode can comprehensively reflect the actual situation of urban and rural areas, provide more comprehensive and related data support for urban and rural planning, so that planning decision is no longer dependent on single-dimensional data, significantly improves the scientificity and comprehensiveness of planning decision, and makes planning results more in line with the actual needs and environmental status of urban and rural residents.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of environmental monitoring technology, specifically relating to a data collection method and system for urban and rural planning surveys. Background Technology

[0002] In today's era of widespread adoption of smart city and smart community concepts, the importance of environmental data monitoring and resident feedback collection in urban and rural planning is becoming increasingly prominent. Accurate, comprehensive, and real-time data support has become a core element in formulating scientific and rational urban and rural planning schemes, playing a decisive role in improving residents' quality of life, optimizing urban resource allocation, and promoting sustainable urban and rural development.

[0003] With the promotion of smart city and smart community concepts, the importance of environmental data monitoring and resident feedback collection in urban and rural planning is becoming increasingly prominent. Traditional urban and rural planning survey methods mainly rely on manual surveys and statistical analysis. Manual surveys not only require a large investment of manpower, material resources, and time, but are also extremely inefficient, making it difficult to meet the timeliness requirements of planning data in the process of rapid urban and rural development; moreover, manual operation is highly susceptible to subjective interference, resulting in a significant reduction in data accuracy. At the same time, due to untimely data updates, its timeliness cannot be effectively guaranteed, causing planning decisions based on such data to become disconnected from the actual development situation of urban and rural areas, seriously affecting the scientific and rational nature of urban and rural development layout.

[0004] Current technologies still suffer from several significant shortcomings. First, environmental data and resident feedback data are typically stored separately, hindering deep integration between multiple data sources. Planners struggle to grasp the overall picture of urban and rural development from a macro perspective, lacking comprehensive and systematic data support for planning decisions, thus severely restricting the scientific rigor and comprehensiveness of planning decisions. Second, while multi-dimensional environmental sensors and smart community platforms improve data collection efficiency, effective methods for data processing and feature extraction are lacking. This significantly reduces the accuracy and reliability of data analysis results, failing to fully uncover the deeper value behind the data and providing precise and effective guidance for urban and rural planning practices. Furthermore, existing technologies are significantly inadequate in geospatial correction and integration of data, failing to fully consider the impact of geographical location on environmental data. Environmental conditions vary greatly across different geographical locations; neglecting this crucial factor during planning leads to poor adaptability of planning schemes to the actual geographical environment, thereby reducing the effectiveness and feasibility of the planning. Summary of the Invention

[0005] The purpose of this invention is to address the problems of separating and storing environmental data and resident feedback data in urban and rural planning surveys, making it difficult to integrate; lacking effective data processing and feature extraction methods, resulting in inaccurate and unreliable analysis results; and insufficient geospatial correction and integration of data, failing to consider the impact of geographical location on environmental data. This invention achieves deep integration, efficient processing, and accurate analysis of multi-source data by constructing an adaptive-scale isolated forest model to preprocess environmental data, building a smart community connection platform to collect resident feedback data, and using a convolutional neural network model to fuse multi-source data and combine it with geographic information correction. This provides comprehensive and reliable data support for urban and rural planning.

[0006] To solve the above-mentioned technical problems, the present invention provides a data collection method and system for urban and rural planning surveys, comprising the following steps:

[0007] Step 1: Collect environmental data;

[0008] Environmental data is collected using urban and rural multidimensional environmental sensors, which include air quality sensors, noise sensors, temperature sensors, humidity sensors, and air pressure sensors.

[0009] The environmental data includes PM2.5 concentration data, PM10 concentration data, CO2 concentration data, NO2 concentration data, environmental noise level data, ambient temperature data, air humidity data, and atmospheric pressure data.

[0010] Step 2: Construct an adaptive-scale isolated forest model and preprocess the environmental data obtained in Step 1;

[0011] Step 2.1: Perform preliminary cleaning of the environmental data, remove null values, and generate a clean dataset;

[0012] Step 2.2: Construct an adaptive-scale isolated forest model. Use the clean dataset obtained in Step 2.1 as the input of the adaptive-scale isolated forest model. The output of the adaptive-scale isolated forest model is the determination result of whether the environmental data is an outlier.

[0013] Step 2.2.1: Using an adaptive strategy, randomly extract time data samples and spatial dimension data samples for each data point from the clean dataset;

[0014] Step 2.2.2: Use multiple decision trees to capture abnormal feature data of environmental data, and generate anomaly scores based on the abnormal features;

[0015] The expression for abnormal scores is:

[0016] (1)

[0017] In the formula For the first The outlier score of each data point The detection function for the adaptive-scale isolated forest model is a named function. Part of For the first There are 1 environmental data point, where i is the index of the data point;

[0018] Step 2.2.3: Set an anomaly threshold. If the anomaly score obtained in Step 2.2.2 exceeds the anomaly threshold, it is determined to be an anomaly.

[0019] Set an anomaly threshold. If the anomaly score calculated by equation (1) exceeds the set anomaly threshold, then the data point is determined to be an anomaly.

[0020] Step 2.2.4: Remove environmental data with outlier results from the clean dataset;

[0021] Step 2.3: Standardize the clean dataset obtained by the method in Step 2.2.4 to generate standardized environmental data.

[0022] Standardized environmental data, expressed as:

[0023] (2)

[0024] (3)

[0025] In the formula, Standardized environmental data values, For environmental data values, At the current time point, To prevent constants with a denominator of zero, For position coordinates, For the measurement error of multi-dimensional environmental sensors in urban and rural areas, For correction functions, For sensor type, This refers to the accuracy attenuation coefficient of the sensor during long-term use. The coefficient represents the sensor type. This is the accuracy attenuation factor of the sensor during short-term calibration. This is a constant term.

[0026] Step 3: Collect resident feedback data corresponding to the environmental data in Step 1, and combine it with the data in Step 2 to form multi-source data;

[0027] Step 3.1: Develop a smart community connection platform using responsive web design and UX design. The front end of the smart community connection platform provides a resident questionnaire interface.

[0028] Step 3.2: Generate a dynamic form on the resident questionnaire survey interface and verify the environmental data. The backend of the smart community connection platform stores and manages the questionnaire survey data.

[0029] Step 3.3: Store the resident questionnaire data in a NoSQL database and encrypt it using a symmetric encryption algorithm;

[0030] Step 3.4: Based on the developed resident questionnaire interface, users fill out the questionnaire online;

[0031] Step 3.5: Set up various questionnaire types on the resident questionnaire survey interface to collect resident feedback data;

[0032] Step 3.6: Real-time communication between the front end and the smart community connection platform server to display feedback information and use the message queue RabbitMQ to process and transmit resident feedback data;

[0033] Step 3.7: Authorize using token authentication and encrypt resident feedback data transmission using HTTPS protocol;

[0034] Step 3.8: Develop a cross-platform application using Flutter;

[0035] Step 3.9: Use Discourse to build a community forum and generate community interaction data;

[0036] Step 3.10: Analyze community interaction data using Gephi;

[0037] Step 3.11: Store the preprocessed environmental data from Step 2 and the resident feedback data in the central database to generate multi-source data.

[0038] Step 4: Use the first neural network model to extract and fuse features from multi-source data to form a unified multidimensional dataset for assessing the degree of environmental pollution;

[0039] Step 4.1: Standardize the environmental data and resident feedback data from the multi-source data using the Z-score algorithm;

[0040] Step 4.2: Use a convolutional neural network model to extract features from environmental data and resident feedback data in the multi-source data. The input of the convolutional neural network model is multi-source data, and the output is a multi-dimensional feature dataset.

[0041] Step 4.2.1: Convolve the multi-source data using multiple convolutional kernels in the convolutional layer to extract local features and generate feature maps;

[0042] Step 4.2.2: Extract the main features from the feature map using a pooling layer and output a compressed feature map;

[0043] Step 4.2.3: Use a fully connected layer to perform linear combination and non-linear activation on the compressed feature map output by the pooling layer to generate high-dimensional feature data;

[0044] Step 4.2.4: The output layer of the convolutional neural network transforms and unifies the high-dimensional feature data through linear transformation and the Softmax activation function, outputting a multi-dimensional feature dataset;

[0045] Step 4.2.5: Perform location correction on the unified multidimensional dataset using geographic information. The expression is:

[0046] (4)

[0047] In the formula, For the location-corrected cube, The integral symbol is used. For environmental data values, For time variables, For geographic correction functions;

[0048] Step 4.3: Obtain the final unified multidimensional dataset.

[0049] The data collection method further includes step 5: using a second neural network model to assess the degree of environmental pollution, wherein the input of the second neural network model is a unified multidimensional dataset and the output is an assessment value of the degree of environmental pollution;

[0050] The second neural network model is a multilayer perceptron model.

[0051] This invention provides a data collection system for a data collection method in urban and rural planning surveys, comprising a collection module, a preprocessing module, a construction module, and a fusion module. The collection module collects environmental data through urban and rural multidimensional environmental sensors and performs preprocessing. The preprocessing module constructs an adaptive-scale isolated forest model to preprocess the environmental data. The construction module builds a smart community connection platform, generates resident feedback data, and combines the preprocessed environmental data and resident feedback data into multi-source data stored in a central database. The fusion module extracts and fuses features from the environmental data and resident feedback data from the multi-source data using a convolutional neural network model to form a unified multidimensional dataset.

[0052] The present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, it implements any step of the data collection method for urban and rural planning surveys described in the present invention.

[0053] The present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any step of a data collection method for urban and rural planning surveys as described in the present invention.

[0054] Compared with the prior art, the beneficial effects of the present invention include:

[0055] (1) This invention achieves deep fusion of multi-source data, enhancing the comprehensiveness of planning decisions. It breaks the traditional situation of separating and storing environmental data and resident feedback data. By combining preprocessed environmental data and resident feedback data into multi-source data and storing them in a central database, and using a convolutional neural network model for feature extraction and fusion, a unified multidimensional dataset is formed, effectively realizing the deep fusion of multi-source data. This fusion method can comprehensively reflect the actual situation in urban and rural areas, providing more comprehensive and relevant data support for urban and rural planning. It makes planning decisions no longer rely solely on single-dimensional data, significantly improving the scientific nature and comprehensiveness of planning decisions, and making planning results more in line with the actual needs of urban and rural residents and the current environmental situation.

[0056] (2) This invention uses a convolutional neural network model to extract and fuse features from environmental data and resident feedback data from multiple sources to form a unified multidimensional dataset. This achieves deep feature extraction and fusion of data. Local features are extracted through convolutional layers, main features are extracted through pooling layers, and high-dimensional feature data is generated through output layers. Linear combination and nonlinear activation are performed through fully connected layers to generate a unified multidimensional feature dataset. Location correction is performed in conjunction with geographic information to ensure the accuracy of data in geographic space. A high-dimensional feature dataset containing rich information is generated, providing a comprehensive high-dimensional feature representation for subsequent data analysis. Ultimately, the accuracy and comprehensiveness of data analysis are improved, providing solid data support for urban and rural planning.

[0057] (3) In contrast to the shortcomings of existing data processing and feature extraction methods, this invention constructs an adaptive-scale isolated forest model to preprocess environmental data. Through preliminary cleaning of the environmental data, the use of adaptive strategies to extract data samples, the use of multiple decision trees to capture abnormal features, the setting of anomaly thresholds and the processing of outliers, and standardization, null values ​​and outliers in the environmental data are effectively removed, improving the accuracy and reliability of the data. Simultaneously, the introduced correction function fully considers sensor type and accuracy degradation factors, further ensuring data quality and laying a solid foundation for subsequent data analysis and applications, making the analysis results more credible and valuable.

[0058] (4) This invention constructs a smart community connection platform, comprehensively utilizing advanced technologies. It employs Responsive Web Design and UX Design to develop the interface, ensuring a good display and operational experience on different mobile devices. The front-end framework React and the back-end framework Node.js are used for development, achieving efficient data processing and response. Real-time communication and data transmission are achieved through WebSocket and RabbitMQ, ensuring timely information updates. OAuth2.0 and HTTPS protocols are used to protect user account security and data privacy. The cross-platform application Flutter is developed, enabling the platform to run stably on different operating systems. Discourse is also used to establish a community forum function, encouraging resident interaction, generating rich community interaction data, and analyzing it with the help of Gephi. The application of these technologies not only improves the efficiency of collecting resident feedback data but also greatly enhances the user experience, strengthens residents' enthusiasm and initiative in participating in urban and rural planning surveys, and helps to obtain more comprehensive and authentic resident feedback information.

[0059] (5) This invention fully considers the impact of geographical location on environmental data. By combining geographic information correction, the location of the unified multidimensional dataset is corrected. The geographic correction function is used to process the data, so that the data can accurately reflect the actual situation of different geographical locations, effectively improving the applicability of the data in geographic space. This enables urban and rural planning analysis based on the data to more accurately consider the characteristics and differences of different regions, providing strong support for the formulation of planning schemes tailored to local conditions. It helps to improve the pertinence and effectiveness of urban and rural planning and promote the sustainable development of urban and rural areas.

[0060] (6) This invention employs a convolutional neural network (CNN) model for feature extraction and fusion of multi-source data. The CNN model used can deeply explore the complex features and intrinsic connections hidden in the multi-source data. Multiple convolutional kernels can extract local features from the multi-source data and generate feature maps, comprehensively reflecting information from different dimensions of the data. The max pooling operation of the pooling layer filters key features, reducing dimensionality while retaining core information. The fully connected layer deeply integrates environmental and resident feedback data features to generate high-dimensional feature data. The output layer finally transforms and unifies them into a multi-dimensional feature dataset. This approach overcomes the differences in data format and semantic gaps between data sources, achieves deep fusion, provides complete and accurate decision-making basis, improves the quality and efficiency of data fusion, enhances data usability, and powerfully promotes the intelligent and scientific development of urban and rural planning.

[0061] (7) This invention employs a multilayer perceptron model and, based on a unified multidimensional dataset containing resident feedback data, achieves a comprehensive assessment of the degree of environmental pollution. This comprehensive assessment method fully utilizes multi-source fusion data, encompassing both objective environmental monitoring data and residents' subjective feedback, making the assessment results more comprehensive and accurate in reflecting the actual pollution situation. The powerful learning ability of the multilayer perceptron model can uncover complex nonlinear relationships between data, accurately quantifying the degree of pollution. The assessment results provide crucial evidence for urban and rural planning, assisting planners in scientifically and rationally arranging urban functional areas, planning transportation routes, and optimizing the configuration of public facilities for different pollution areas and degrees. It can also assist in formulating effective pollution control and prevention strategies, promoting the green and sustainable development of urban and rural environments, and improving the quality of life for residents. Attached Figure Description

[0062] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0063] Figure 1 This is a flowchart of a data collection method for urban and rural planning surveys in an embodiment of the present invention.

[0064] Figure 2 This is a block diagram of a data collection system for urban and rural planning surveys in an embodiment of the present invention.

[0065] Figure 3 This is a flowchart for urban and rural planning data collection and processing in an embodiment of the present invention.

[0066] Figure 4 This is a system architecture diagram of the smart community connection platform in an embodiment of the present invention.

[0067] Figure 5 This is a schematic diagram of the interface of a data collection system for urban and rural planning surveys in an embodiment of the present invention.

[0068] Figure 6 This is a schematic diagram of the data processing flow in an embodiment of the present invention.

[0069] Figure 7 This is a line graph showing the change in PM2.5 concentration data for a single indicator in an embodiment of the present invention.

[0070] Figure 8 This is a line graph showing the combined data changes in an embodiment of the present invention. Detailed Implementation

[0071] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0072] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0073] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0074] like Figure 1 and Figure 2 and Figure 3 As shown, a data collection method for urban and rural planning surveys includes the following steps:

[0075] Step 1: Collect environmental data;

[0076] Define the objectives and scope of the investigation;

[0077] Design a specific implementation plan for the intelligent sensing network and smart community connection platform;

[0078] Assemble an investigation team, including technical personnel, investigators, and data analysts;

[0079] Develop a detailed implementation plan, including sensor deployment schemes and a development plan for a smart community connectivity platform;

[0080] According to the implementation plan, the technical team deployed air quality sensors, noise sensors, temperature sensors, humidity sensors, and air pressure sensors in the target area as the basis for urban and rural multi-dimensional environmental sensors, and collected environmental data such as PM2.5 concentration data, PM10 concentration data, CO2 concentration data, NO2 concentration data, environmental noise level data, ambient temperature data, air humidity data, and atmospheric pressure data.

[0081] Ensure the stable operation of the urban and rural multi-dimensional environmental sensor network and the real-time uploading of environmental data;

[0082] Urban and rural multi-dimensional environmental sensors transmit environmental data to a central database via a wireless network. Real-time uploading of environmental data ensures its freshness and accuracy. The environmental data includes timestamps, location coordinates, and environmental parameters.

[0083] Each urban and rural multi-dimensional environmental sensor can simultaneously monitor multiple environmental parameters, including PM2.5 concentration data, PM10 concentration data, CO2 concentration data, NO2 concentration data, environmental noise level, noise level in a specific frequency band, ambient temperature, air humidity, and atmospheric pressure.

[0084] High-precision sensor components are used to ensure the accuracy and reliability of environmental data;

[0085] The urban and rural multi-dimensional environmental sensor has a self-calibration function, which can perform self-testing and calibration regularly to ensure the accuracy of environmental data during long-term use;

[0086] All environmental data collected by urban and rural multi-dimensional environmental sensors will be transmitted in an encrypted manner to ensure the security of the environmental data;

[0087] The urban and rural multi-dimensional environment sensor adopts a low-power design to extend battery life and adapt to various outdoor environments.

[0088] Based on the environmental data collection requirements, set the monitoring parameters and upload frequency for urban and rural multi-dimensional environmental sensors;

[0089] Urban and rural multi-dimensional environmental sensors collect environmental data in real time;

[0090] It should be noted that by deploying multi-dimensional environmental sensors in urban and rural areas and designing intelligent sensing networks and smart community connection platforms, environmental data can be collected and transmitted efficiently. The sensors cover a variety of parameters such as air quality, noise, temperature, humidity, and air pressure, ensuring comprehensive and accurate data. Self-calibration function and encrypted transmission guarantee the long-term reliability and security of the data. The low-power design and real-time upload function of the sensors enhance practicality and stability, effectively supporting the dynamic monitoring and planning management of urban and rural environments.

[0091] Step 2: Construct an adaptive-scale isolated forest model and preprocess the environmental data obtained in Step 1;

[0092] Step 2.1: Perform preliminary cleaning of the environmental data, remove null values, and generate a clean dataset;

[0093] Step 2.2: Construct an adaptive-scale isolated forest model. Use the clean dataset obtained in Step 2.1 as the input of the adaptive-scale isolated forest model. The output of the adaptive-scale isolated forest model is the determination result of whether the environmental data is an outlier.

[0094] Step 2.2.1: Using an adaptive strategy, randomly extract time data samples and spatial dimension data samples for each data point from the clean dataset to ensure that the adaptive scale isolated forest model can capture the changing characteristics of the data at different time and spatial scales;

[0095] The adaptation strategy combines the features of the existing isolated forest algorithm and extends it to enable it to work effectively at different scales and adapt to the complexity and dynamic changes of environmental data.

[0096] Step 2.2.2: Use multiple decision trees to capture anomalous features of the environmental data, and generate anomaly scores based on these features. The expression for the anomaly score is as follows:

[0097] (1)

[0098] In the formula, For the first The outlier score of each data point This is the detection function for an adaptive-scale isolated forest model, used to calculate the outlier score for data points. For adaptive scaling, it is a named function. Part of it, used to indicate that this function has the ability to adaptively adjust the scale. For the first There are 1 environmental data point, where i is the index of the data point;

[0099] By calculating the outlier score for each data point, outliers in the data can be effectively identified.

[0100] Step 2.2.3: Set an anomaly threshold. If the anomaly score obtained in Step 2.2.2 exceeds the anomaly threshold, it is determined to be an anomaly.

[0101] Set an anomaly threshold. If the anomaly score calculated by equation (1) exceeds the set anomaly threshold, then the data point is determined to be an anomaly.

[0102] The threshold is set as follows:

[0103] when If the value is greater than 1.55, the data point is considered an outlier.

[0104] when If the value is ≤1.55, the data point is considered a normal value;

[0105] Step 2.2.4: Remove environmental data with outlier results from the clean dataset;

[0106] Detected outliers are removed and replaced with the mean of neighboring data points to ensure data continuity and rationality. The adaptive algorithm dynamically adjusts the processing strategy based on the current data distribution to ensure the rationality and effectiveness of outlier handling.

[0107] Specifically, multiple subsets of data are randomly selected from a clean dataset;

[0108] A decision tree is trained for each subset of data. These decision trees learn the distribution and characteristics of the data and are able to identify normal and abnormal patterns in the data.

[0109] Each decision tree evaluates the entire dataset and calculates the path length for each data point. Data points with shorter path lengths are usually considered outliers because they are more likely to be isolated by the decision tree.

[0110] The evaluation results of each decision tree are combined to generate a comprehensive anomaly score for each data point;

[0111] By combining the outlier scores generated by all decision trees, the final outlier score for each data point can be obtained by averaging or other weighting methods.

[0112] The outlier scores are standardized to ensure that the distribution of outlier scores is suitable for subsequent threshold settings;

[0113] Based on the standardized outlier score, an outlier threshold is set. When the outlier score calculated by equation (1) exceeds the set outlier threshold, the data point is determined to be an outlier.

[0114] Step 2.3: Standardize the clean dataset obtained by the method in Step 2.2.4 to generate standardized environmental data.

[0115] The clean dataset after detection by the adaptive-scale isolated forest model is standardized to generate standardized environmental data, expressed as:

[0116] (2)

[0117] In the formula, These are standardized environmental data values, representing data at a specific point in time and location, combining sensor type, health status, and environmental parameters. For environmental data values, From the initial time At the appointed time The time difference is used to characterize the dynamic changes of data over time. To prevent constants with a denominator of zero, These are location coordinates, used to represent the specific geographical location of the data collection point. The measurement error of urban and rural multi-dimensional environmental sensors is corrected by generating a correction coefficient by combining the sensor type and health status. For sensor types, such as air quality sensors, noise sensors, etc. This is a coefficient for the accuracy decay of a sensor during long-term use. For example, it can be the percentage of sensor performance degradation obtained through periodic calibration or the long-term deviation between the actual measured value and the standard value. A coefficient can be assigned to different accuracy decay situations.

[0118] Correction function The expression is:

[0119] (3)

[0120] In the formula, This is a coefficient representing the sensor type, indicating the impact of different sensor types on the correction value. For example, for an air quality sensor... The noise sensor is wait, The constant term represents the coefficient of accuracy degradation of the sensor during short-term calibration. Obtained by fitting experimental or historical data, it is used to balance and optimize the output of the correction function;

[0121] It should be noted that preprocessing environmental data using an adaptive-scale isolated forest model effectively removes null values ​​and outliers. Dynamic threshold settings accurately identify and handle outliers, standardized processing ensures data consistency and improves analytical reliability, and multiple decision trees capture data features, combined with correction functions to enhance accuracy. Efficient encrypted data transmission and real-time monitoring guarantee data security and stable equipment operation.

[0122] Step 3: Collect resident feedback data corresponding to the environmental data from Step 1, and combine it with the data from Step 2 to form multi-source data, such as... Figure 4 As shown.

[0123] Step 3.1: Develop a smart community connectivity platform using Responsive Web Design and UX Design, and display and operate it via mobile devices. The front-end framework React provides a resident questionnaire interface for the smart community connectivity platform, such as... Figure 5 As shown;

[0124] Responsive Web Design is a web design methodology that enables web pages to automatically adjust their layout and content display on different devices and screen sizes. By using flexible grid layouts, resizable images, and CSS media queries, web pages can be displayed correctly on various devices such as mobile phones, tablets, and computers. This ensures that web page content is clear, readable, and easy to operate regardless of the device the user uses to access the website, avoiding the need to create separate versions for different devices and reducing development and maintenance costs.

[0125] User Experience Design (UX Design) is a design process aimed at improving the overall experience of users when interacting with products or services. It involves understanding users' needs, behaviors, and pain points through research and analysis, developing design strategies, organizing and structuring content so that users can easily find the information they need, designing interfaces and interactive elements to ensure the convenience and enjoyment of user operations, and continuously optimizing the design through testing and feedback to improve user satisfaction.

[0126] Step 3.2: Generate a dynamic form on the resident questionnaire survey interface and verify the environmental data. The backend Node.js development framework of the smart community connection platform stores and manages the questionnaire survey data.

[0127] React is an open-source JavaScript library developed by Facebook for building user interfaces, especially single-page applications (SPAs). It allows developers to create reusable UI components to build complex user interfaces. React provides a component-based development model, enabling developers to break down the UI into independent, reusable components, improving development efficiency and code maintainability. React uses a virtual DOM to improve performance; each time the state changes, React generates a new virtual DOM tree and compares it with the old virtual DOM, updating only the parts of the actual DOM that need to be changed. React adopts a unidirectional data flow approach, with data flowing from parent components to child components, ensuring data controllability and ease of debugging. React ensures that residents can easily fill out questionnaires and submit feedback data.

[0128] Node.js is a JavaScript runtime environment based on the V8 engine, allowing developers to run JavaScript code on the server side. It uses an event-driven, non-blocking I / O model, making it suitable for building high-concurrency, high-performance network applications. Node.js can be used to build server-side applications, implementing functions such as APIs and web servers. Node.js supports asynchronous programming, and through event loops and callback mechanisms, it can handle a large number of concurrent requests without blocking threads. Node.js comes with the Node Package Manager, a rich package management tool and ecosystem hub that enables developers to easily integrate and use various third-party libraries and modules. Node.js achieves efficient data processing and response, ensuring stable operation and handling of a large number of user requests.

[0129] Step 3.3: Store the questionnaire data submitted by residents using a NoSQL database, and encrypt the questionnaire data using AES encryption to ensure data security;

[0130] Step 3.4: Based on the developed resident questionnaire interface, residents fill out the questionnaire online and submit their opinions and suggestions;

[0131] Step 3.5: The resident questionnaire survey interface is set up with multiple questionnaire survey types to comprehensively collect resident feedback data;

[0132] Step 3.6: The front end uses WebSocket to achieve real-time communication with the server of the smart community connection platform, displays the latest feedback information, and uses RabbitMQ message queue to process and transmit real-time resident feedback data to ensure timely updates of information;

[0133] WebSocket is a communication protocol that provides a full-duplex communication channel, allowing real-time data exchange between clients and servers. WebSocket allows a persistent connection between clients and servers, enabling real-time data transmission. Unlike the traditional HTTP request-response model, WebSocket supports bidirectional communication, allowing servers to proactively push data to clients instead of simply responding to client requests. Because the WebSocket connection is persistent, data transmission latency is greatly reduced, making it ideal for applications requiring real-time updates, such as online games, real-time chat, and real-time data monitoring. Through WebSocket, the front end can maintain a persistent connection with the server, receiving and displaying the latest feedback information in real time, ensuring the immediacy and interactivity of the data.

[0134] RabbitMQ is an open-source message queue center used for message transmission, processing, and distribution. It supports multiple message protocols, such as AMQP, STOMP, and MQTT. RabbitMQ can transmit and process messages between different applications, enabling asynchronous communication. RabbitMQ uses queues to cache and distribute messages, ensuring that messages are not lost even if the receiver is temporarily unavailable. Through the message queue mechanism, RabbitMQ can achieve load balancing, distributing messages evenly to multiple receivers, improving processing capacity and reliability. RabbitMQ supports message persistence and acknowledgment mechanisms, ensuring message reliability and data integrity during transmission. In the background, RabbitMQ is responsible for processing and transmitting real-time data, ensuring data transmission reliability and load balancing through the message queue mechanism, enabling efficient handling of a large number of concurrent requests.

[0135] Step 3.7: Protect user account security through the OAuth 2.0 authentication and authorization mechanism, and encrypt all resident feedback data transmissions using the HTTPS protocol to ensure the privacy of user data;

[0136] OAuth 2.0 is an authorization framework primarily used to enable third-party applications to obtain user access to resources on other services. It achieves this through interaction between an authorization server and a resource server. OAuth 2.0 allows applications to access resources on behalf of users without exposing the user's username and password. Through the authorization process, the application obtains an access token, which grants it access to the user's resources. By using access tokens, OAuth 2.0 reduces the risk of user credential leakage. Access tokens have an expiration date and can be refreshed as needed. Users can control and manage the access permissions of third-party applications, revoke authorization at any time, and ensure controllability of data access. By implementing OAuth 2.0, the smart community connection platform can effectively protect user account security, ensure the privacy and security of user data during transmission, thereby enhancing user trust and platform security.

[0137] Step 3.8: Develop the cross-platform application Flutter to enable the smart community connection platform to run smoothly on different mobile devices, ensuring that the platform can run smoothly on different operating systems such as Android and iOS;

[0138] Flutter allows developers to create applications that run on different platforms such as Android, iOS, Web, and desktop using a single codebase. It uses the Dart language and provides a rich set of components and tools to help developers quickly build beautiful user interfaces. Flutter allows developers to write a single codebase that can run on multiple platforms, significantly reducing development and maintenance costs. Through Flutter's Skia rendering engine, applications can achieve near-native performance and a smooth experience. Flutter provides a series of predefined, customizable UI components, enabling developers to easily build beautiful and feature-rich user interfaces. Flutter supports "hot reload," meaning developers can immediately see the effects of code modifications, greatly improving development efficiency. By using Flutter to develop a smart community connection platform, it ensures smooth operation on different operating systems such as Android and iOS, not only reducing development and maintenance costs but also providing high performance and a consistent user experience.

[0139] Step 3.9: Use the open-source forum software Discourse to build a community forum function, allowing residents to post and comment on posts, and add comment and like functions to encourage interaction among residents and generate community interaction data;

[0140] Step 3.10: Establish a community forum function using the open-source forum software Discourse. Residents can post and comment on posts and participate in community discussions. In addition, comment and like functions have been added to encourage resident interaction, increase community activity and participation. By collecting and analyzing this interaction data, valuable community interaction data is generated.

[0141] Discourse is a modern forum software written in Ruby on Rails with an Ember.js front-end. It aims to replace traditional forum software by enhancing online community interaction and engagement through a better user experience and functional design. Discourse provides a platform that allows users to create posts, replies, and comments, promoting communication and interaction among community members. Real-time notifications keep users informed of new posts, replies, and likes, further enhancing interactivity. Discourse supports category, tag, and topic management, helping users organize and find content of interest. Through trust levels, Discourse encourages active user participation, increasing privileges and functionality as levels rise. Discourse offers a rich set of plugins to extend functionality, such as polls, event management, and a points center. Its responsive interface ensures smooth display and operation on mobile devices. By using Discourse, smart community connection platforms can create highly interactive and engaging community forums, providing residents with a space for free communication and sharing, thereby enhancing community cohesion and resident satisfaction.

[0142] Use the social network analytics tool Gephi to analyze community interaction data and enhance community engagement;

[0143] Gephi is an interactive graphical visualization platform that allows users to import, manipulate, and analyze graph data to help discover hidden patterns and trends. It is widely used in social network analysis, network science, data mining, and information visualization. Gephi provides powerful visualization tools that can display complex network data graphically, enabling users to intuitively observe network structure and relationships between nodes. Through various analysis algorithms, Gephi can calculate and display key network metrics, such as node degree centrality, betweenness centrality, and clustering coefficients, helping users understand the properties and structure of the network. Gephi supports dynamic network analysis, showing how network structure changes over time, making it suitable for analyzing dynamic interactions in social networks. Gephi boasts a rich plugin ecosystem, allowing users to add functional modules as needed to expand its analytical and visualization capabilities. As a powerful social network analysis tool, Gephi helps smart community platforms enhance resident participation and interaction by analyzing and visualizing community interaction data, promoting community development and management.

[0144] Step 3.11: Store the preprocessed environmental data from Step 2 and the resident feedback data submitted through the smart community connection platform in the central database to generate multi-source data.

[0145] It should be noted that by building a smart community connectivity platform and combining Responsive Web Design and UX Design, a consistent user experience across various devices is ensured. React and Node.js are used for front-end and back-end development to ensure efficient data processing and response. WebSocket and RabbitMQ are used for real-time communication and data transmission, enhancing immediacy and reliability. OAuth 2.0 and HTTPS are used to protect user data privacy. Cross-platform applications are developed using Flutter to improve user experience. Discourse and Gephi are employed to promote community interaction and enhance community engagement. This comprehensive approach improves the comprehensiveness, security, and user participation of data collection.

[0146] Step 4: Extract and fuse features from environmental data and resident feedback data from multiple sources using a convolutional neural network model to form a unified multidimensional dataset for assessing the degree of environmental pollution, such as... Figure 6 As shown;

[0147] Organize and classify environmental data and resident feedback data from multiple sources to ensure data consistency and integrity;

[0148] Step 4.1: Standardize the environmental data and resident feedback data from the multi-source data using the Z-score standardization algorithm to ensure a uniform data format and eliminate differences between different data sources;

[0149] Step 4.2: Use a convolutional neural network model to extract features from environmental data and resident feedback data from multiple sources, generate high-dimensional feature data, and merge them into a unified multidimensional dataset;

[0150] The input to a convolutional neural network model is defined as multi-source data, and the output is a multi-dimensional feature dataset;

[0151] By convolving multiple convolutional kernels in a convolutional layer with multi-source data, local features are extracted and feature maps are generated. Convolutional layers can capture local patterns in data and generate feature maps.

[0152] Specifically, the convolution kernel is a small matrix, usually 3x3 or 5x5 in size, that slides across the input multi-source data;

[0153] The convolutional kernel slides gradually across the input multi-source data. Each slide calculates the dot product between the convolutional kernel and the corresponding region of the input multi-source data to generate a feature value.

[0154] By using a sliding window and convolution operations, a two-dimensional feature map can be generated, which contains local features of the input multi-source data;

[0155] After the convolution operation, an activation function is applied to perform a non-linear transformation on the feature map, increasing the expressive power of the convolutional neural network model.

[0156] Convolutional layers output multiple feature maps, each corresponding to a feature extracted by a convolutional kernel. These feature maps represent the local features of the input multi-source data.

[0157] Max pooling is used to extract air quality, temperature, humidity, and noise from environmental data and feedback frequency from resident feedback data from the feature map as the main features. The pooling layer can reduce the dimensionality of the feature map, reduce the amount of computation, and retain the main feature information.

[0158] Specifically, a fixed-size window is applied to the feature map, and the maximum value within each window is selected as the pooling result.

[0159] The window slides gradually across the feature map, and the maximum value within the window is calculated with each slide.

[0160] Pooling compression reduces the dimensionality of feature maps, decreases computational load, and preserves key feature information, such as air quality, temperature, humidity, and noise in environmental data, and feedback frequency in resident feedback data.

[0161] The pooling layer outputs a compressed feature map, which retains the main feature information from the input multi-source data;

[0162] A fully connected layer is used to linearly combine and non-linearly activate the main features output by the pooling layer to generate high-dimensional feature data.

[0163] Specifically, the feature map output by the pooling layer is flattened into a one-dimensional vector to facilitate subsequent processing.

[0164] The flattened one-dimensional vector is multiplied by the weight matrix of the fully connected layer to generate a linear combination result, and a bias term is added to enhance the expressive power of the convolutional neural network model.

[0165] Applying nonlinear activation functions to activate the linear combination results increases the nonlinear expressive power of the convolutional neural network model;

[0166] Based on the importance and relevance of the main features of environmental data characteristics and resident feedback data characteristics, different weights are assigned to different features to further optimize the feature fusion effect and generate high-dimensional feature data.

[0167] For high-weight features such as air quality PM2.5 and PM10 concentration features and feedback frequency in environmental data features and resident feedback data features, high-weight features are given to highlight their importance;

[0168] For features with medium weight, such as environmental data features and resident feedback data features, such as CO2 concentration, ambient temperature, and satisfaction rating, medium weight features are assigned to reflect their moderate importance.

[0169] Low-weight features, such as environmental data features and resident feedback data features like air humidity, suggestions, and opinions, are assigned low weight to reduce their impact on the final feature fusion.

[0170] The output layer transforms and unifies high-dimensional feature data through linear transformation and the Softmax activation function, and outputs a multi-dimensional feature dataset, which represents the deep feature representation of the input data.

[0171] The output layer performs linear transformations on high-dimensional feature data to adapt it to the final output form. For example, it uses the weighted summation operation of the fully connected layer to combine different features into a unified data representation.

[0172] The output layer can use appropriate activation functions to process the linearly transformed data to ensure the nonlinear characteristics and range of the output data. For example, the Softmax activation function, which is commonly used in classification tasks, can convert the output into a probability distribution.

[0173] The unified multidimensional feature dataset generated by the output layer represents the model's final understanding and feature representation of the input data. This feature data can be used for further analysis, decision-making, or other tasks.

[0174] By incorporating geographic information to perform location correction on the unified multidimensional dataset, the accuracy of the data in geospatial space is ensured. The expression is as follows:

[0175] (4)

[0176] In the formula, For the location-corrected cube, The integral symbol is used. These are environmental data values, i.e., the raw environmental data. For time variables, represent time variables A tiny increment, This is a geographic correction function, representing the relationship between location coordinates. The relevant geographic correction function is used to correct and adjust the data to ensure that the impact of geographic location on the data is fully considered.

[0177] It should be noted that a convolutional neural network model is used to extract and fuse features from multi-source data, forming a unified multidimensional dataset to achieve data consistency and integrity. The Z-score normalization algorithm is used to eliminate differences between data sources, ensuring a uniform data format. Convolutional and pooling layers extract deep features, outputting high-dimensional feature data, and fully connected layers further optimize the feature fusion effect. Combined with geographic information correction, the accuracy of the data in geospatial context is ensured, ultimately providing a comprehensive and accurate high-dimensional feature dataset, offering solid data support for urban and rural planning decisions.

[0178] In one embodiment, the unified multidimensional dataset uses a multilayer perceptron model to assess the degree of environmental pollution, defining the input as the unified multidimensional dataset and the output as the environmental pollution level assessment value.

[0179] Initialize the parameters of the multilayer perceptron model;

[0180] The input layer takes the fused unified multidimensional dataset as input, with n features.

[0181] Define two hidden layers: the first hidden layer has n1 nodes and the second hidden layer has n2 nodes.

[0182] The output layer has 1 node and is used to output continuous environmental pollution level assessment values.

[0183] The path from the input layer to the first hidden layer is as follows:

[0184] weight matrix The dimension is ;

[0185] weight value ;

[0186] In the formula, For the input layer to the first hidden layer The input feature to the first The weights of each hidden layer node;

[0187] bias vector Dimensions ;

[0188] The distance from the first hidden layer to the second hidden layer is:

[0189] weight matrix The dimension is ;

[0190] weight value ;

[0191] In the formula, From the first hidden layer to the second hidden layer The input feature to the first The weights of each hidden layer node;

[0192] bias vector Dimensions ;

[0193] The second hidden layer to the output layer is;

[0194] weight matrix The dimension is ;

[0195] weight value ;

[0196] In the formula, For the first Weights from each input feature to the output layer;

[0197] bias vector ;

[0198] The forward propagation input sample is x, and the dimension is n;

[0199] The input to the first hidden layer is In the formula for The transpose of; the output is = In the formula This is the ReLU activation function.

[0200] The input to the second hidden layer is In the formula for The transpose of the given value is: In the formula This is the ReLU activation function.

[0201] The input of the output layer is , for The transpose of the result; the output is the predicted value. ;

[0202] The mean squared error loss function is as follows:

[0203] ;

[0204] In the formula, For the sample size, Let be the true value of the i-th sample. Let be the predicted value for the i-th sample.

[0205] The error of the output layer calculated by backpropagation is ;

[0206] The error of the second hidden layer is ;

[0207] In the formula, The derivative of the ReLU activation function. This indicates element-wise multiplication.

[0208] The error of the first hidden layer is ;

[0209] In the formula, The derivative of the ReLU activation function. This indicates element-wise multiplication.

[0210] Calculate the gradients of the weights and biases:

[0211] for The gradient is ;

[0212] for The gradient is ;

[0213] for The gradient is ;

[0214] for The gradient is ;

[0215] for The gradient is ;

[0216] for The gradient is ;

[0217] Adam optimizer updates first-moment estimates Second-order moment estimation They are respectively:

[0218]

[0219]

[0220] In the formula, Let β1 be the gradient of iteration t, β2 be the decay rate estimated by the first moment, and β3 be the decay rate estimated by the second moment.

[0221] Corrected first-moment estimate Second-order moment estimation They are respectively:

[0222] ;

[0223] ;

[0224] In the formula, The corrected first-moment estimate, The corrected second-moment estimate;

[0225] The model parameters have been updated as follows:

[0226] ;

[0227] In the formula, θt is the parameter of iteration t, θt−1 is the parameter of iteration t-1, α is the learning rate, and ϵ is a small constant to prevent the denominator from being zero;

[0228] Evaluate the model;

[0229] The root mean square error (RMSE) is:

[0230] ;

[0231] The mean absolute error (MAE) is:

[0232] ;

[0233] Relevant data are accurately extracted from a unified multidimensional dataset based on the environmental indicators to be analyzed.

[0234] Choose the appropriate chart type to draw based on the characteristics of the data and the purpose of the analysis.

[0235] To display the trend of a single environmental indicator over time, a line chart can be drawn, with the x-axis representing the time point and the y-axis representing the corresponding environmental indicator value, clearly showing the trend of data change.

[0236] like Figure 7 The process involves using data query statements to filter out all data records related to PM2.5 concentration in the dataset, ensuring that the acquired data covers measurements from different time points. Simultaneously, the time point information corresponding to each PM2.5 concentration data point is extracted to construct a complete dataset structure. The same data filtering logic is applied to other indicators such as temperature, humidity, and noise levels to ensure the accuracy and completeness of the acquired data, providing a reliable foundation for subsequent chart creation.

[0237] For multiple environmental indicators that are related, the relationship between them can be shown by combining charts or graphs.

[0238] like Figure 8As shown, the maximum temperature is represented by a blue line, with the scale and label indicating the unit as "°C," visually reflecting the trend of temperature changes. The minimum temperature is represented by a light blue line, also labeled with the unit "°C," contrasting with the maximum temperature line to show the fluctuations in the temperature range. Sulfur dioxide and ozone concentrations are represented by green and pink lines, respectively, and displaying them in the same chart facilitates observation of their potential correlation with temperature changes. Furthermore, cyan lines represent wind direction data. Although wind direction data may differ from other indicators in terms of dimensions and meaning, displaying it in the same chart helps in comprehensively analyzing the synergistic changes of multiple environmental factors. Thus, this multi-line, dual-axis chart presentation method comprehensively and intuitively presents the changing trends of various environmental indicators and their potential interrelationships.

[0239] Observe the charts and analyze the patterns and trends in environmental data. From individual indicator charts, determine whether environmental indicators exceed normal ranges or standard values. For example, compare PM2.5 concentration data changes with values ​​specified by air quality standards. If PM2.5 concentrations consistently exceed standard values ​​over a certain period, it indicates air pollution during that time. Simultaneously, compare charts of different indicators to uncover potential correlations between data. Consult relevant environmental science literature to understand how changes in temperature and humidity affect the diffusion and chemical reactions of pollutants in the atmosphere, thus impacting air quality.

[0240] By comprehensively analyzing the charts, we can determine the severity of environmental pollution, the type of pollution, and the time and location of pollution occurrence, providing a strong basis for formulating subsequent governance strategies.

[0241] One embodiment of a data collection system for urban and rural planning surveys includes: a collection module, a preprocessing module, a construction module, and a fusion module. The collection module collects environmental data using urban and rural multidimensional environmental sensors and performs preprocessing. The preprocessing module constructs an adaptive-scale isolated forest model to preprocess the environmental data. The construction module builds a smart community connection platform, generates resident feedback data, and combines the preprocessed environmental data and resident feedback data into multi-source data stored in a central database. The fusion module extracts and fuses features from the environmental data and resident feedback data from the multi-source data using a convolutional neural network model to form a unified multidimensional dataset. It also includes using a multilayer perceptron model to assess the degree of environmental pollution.

[0242] One embodiment employs a computer device suitable for a data collection method for urban and rural planning surveys, comprising: a memory and a processor; the memory stores computer-executable instructions, and the processor executes the computer-executable instructions to implement the data collection method for urban and rural planning surveys as proposed in the above embodiment.

[0243] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, Near Field Communication (NFC), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.

[0244] One embodiment employs a storage medium on which a computer program is stored. When executed by a processor, the program implements the data collection method for urban and rural planning surveys as proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof.

[0245] In summary, this invention uses a convolutional neural network model to extract and fuse features from environmental data and resident feedback data from multiple sources, forming a unified multidimensional dataset. This achieves deep feature extraction and fusion, extracting local features through convolutional layers, extracting main features through pooling layers, generating high-dimensional feature data through the output layer, and performing linear combination and non-linear activation through fully connected layers to generate a unified multidimensional feature dataset. Location correction is performed using geographic information to ensure the accuracy of the data in geospatial space. Its function is to generate a high-dimensional feature dataset containing rich information, providing a comprehensive high-dimensional feature representation for subsequent data analysis, ultimately improving the accuracy and comprehensiveness of data analysis and providing solid data support for urban and rural planning.

[0246] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A data collection method for urban and rural planning surveys, characterized in that, Includes the following steps: Step 1: Deploy multi-dimensional environmental sensors in urban and rural areas to collect environmental data; Step 2: Construct an adaptive-scale isolated forest model and preprocess the environmental data obtained in Step 1; Step 3: Collect resident feedback data corresponding to the environmental data in Step 1, and combine it with the data in Step 2 to form multi-source data; Step 3.1: Develop a smart community connection platform. The front end of the smart community connection platform provides a resident questionnaire survey interface. Step 3.2: Generate a dynamic form on the resident questionnaire survey interface and verify the environmental data. The backend of the smart community connection platform stores and manages the questionnaire survey data. Step 3.3: Store the resident questionnaire data in a NoSQL database and encrypt it using a symmetric encryption algorithm; Step 3.4: Based on the developed resident questionnaire interface, users fill out the questionnaire. Step 3.5: Set up various questionnaire types on the resident questionnaire survey interface to collect resident feedback data; Step 3.6: The front end communicates with the server of the smart community connection platform in real time and displays feedback information. The RabbitMQ message queue is used to process and transmit resident feedback data. Step 3.7: Authorize using token authentication and encrypt resident feedback data transmission using HTTPS protocol; Step 3.8: Develop a cross-platform application using Flutter; Step 3.9: Use Discourse to build a community forum and generate community interaction data; Step 3.10: Analyze community interaction data using Gephi; Step 3.11: Store the preprocessed environmental data from Step 2 and resident feedback data in the central database to generate multi-source data; Step 4: Use the first neural network model to extract and fuse features from multi-source data to form a unified multidimensional dataset for assessing the degree of environmental pollution; The first neural network model is a convolutional neural network model, and step 4 includes the following sub-steps: Step 4.1: Standardize the environmental data and resident feedback data from the multi-source data using the Z-score algorithm; Step 4.2: Use a convolutional neural network model to extract features from environmental data and resident feedback data in the multi-source data. The input of the convolutional neural network model is multi-source data, and the output is a multi-dimensional feature dataset. Step 4.2.1: Convolve the multi-source data using multiple convolutional kernels in the convolutional layer to extract local features and generate feature maps; Step 4.2.2: Extract the main features from the feature map using a pooling layer and output a compressed feature map; Step 4.2.3: Use a fully connected layer to perform linear combination and non-linear activation on the compressed feature map output by the pooling layer to generate high-dimensional feature data; Step 4.2.4: The output layer of the convolutional neural network transforms and unifies the high-dimensional feature data through linear transformation and the Softmax activation function, outputting a multi-dimensional feature dataset; Step 4.2.5: Perform location correction on the unified multidimensional dataset using geographic information. The expression is: ;(4) In the formula, For the location-corrected cube, The integral symbol is used. Here are the environmental data values, and t is the time variable. For geographic correction functions; At the current time point, This is the initial time; These are the position coordinates; Step 4.3: Obtain the final unified multidimensional dataset.

2. The data collection method for urban and rural planning surveys according to claim 1, characterized in that, In step 1, the urban and rural multidimensional environmental sensor includes an air quality sensor, a noise sensor, a temperature sensor, a humidity sensor, and a barometric pressure sensor.

3. The data collection method for urban and rural planning surveys according to claim 2, characterized in that, In step 1, the environmental data includes PM2.5 concentration data, PM10 concentration data, CO2 concentration data, NO2 concentration data, environmental noise level data, noise level data, environmental temperature data, air humidity data, and atmospheric pressure data.

4. The data collection method for urban and rural planning surveys according to claim 3, characterized in that, Step 2 includes the following sub-steps: Step 2.1: Clean the environmental data to generate a clean dataset; Step 2.2: Construct an adaptive-scale isolated forest model. Use the clean dataset obtained in Step 2.1 as the input of the adaptive-scale isolated forest model. The output of the adaptive-scale isolated forest model is the determination result of whether the environmental data is an outlier. Step 2.2.1: Use an adaptive strategy to extract data samples from a clean dataset; Step 2.2.2: Use multiple decision trees to capture abnormal features of environmental data, and generate anomaly scores based on the abnormal features; Step 2.2.3: Set an anomaly threshold. If the anomaly score obtained in Step 2.2.2 exceeds the anomaly threshold, it is determined to be an anomaly. Step 2.2.4: Remove environmental data with outlier results from the clean dataset; Step 2.3: Standardize the clean dataset obtained by the method in Step 2.2.4 to generate standardized environmental data.

5. The data collection method for urban and rural planning surveys according to claim 4, characterized in that, In step 2, the adaptive scale isolated forest model randomly extracts time data samples and spatial dimension data samples for each data point from the clean dataset using an adaptive strategy. By using multiple decision trees, we can capture anomalous feature data from a clean dataset and generate anomaly scores. The expression for abnormal scores is: ;(1) In the formula, For the first The outlier score of each data point For the detection function of the adaptive-scale isolated forest model, For the first There are 1 environmental data point, where i is the index of the data point; Set an anomaly threshold. If the anomaly score calculated by equation (1) exceeds the set anomaly threshold, then the data point is determined to be an anomaly. The clean dataset after detection by the adaptive-scale isolated forest model is standardized to generate standardized environmental data, expressed as: ;(2) ;(3) In the formula For standardized environmental data values, For environmental data values, At the current time point, The initial time, To prevent constants with a denominator of zero, For position coordinates, For the measurement error of multi-dimensional environmental sensors in urban and rural areas, For correction functions, Indicates the sensor type, This refers to the accuracy attenuation coefficient of the sensor during long-term use. The coefficient represents the sensor type. This is the accuracy attenuation factor of the sensor during short-term calibration. This is a constant term.

6. The data collection method for urban and rural planning surveys according to claim 5, characterized in that, The data collection method further includes step 5: using a second neural network model to assess the degree of environmental pollution, wherein the input of the second neural network model is a unified multidimensional dataset and the output is an assessment value of the degree of environmental pollution; The second neural network model is a multilayer perceptron model.

Citation Information

Patent Citations

  • Low-carbon strategy optimization method and system for residential users

    CN116822733A

  • Municipal garden water pollution monitoring system and method based on big data analysis

    CN118886718A

  • Intelligent community resident health management method based on machine learning

    CN119274802A