Data collection method and system for urban and rural planning investigation
By building an adaptive scale isolated forest model and building a smart community connection platform, combining the convolutional neural network model to fuse multi-source data and perform geographic information correction, the problems of data separation storage, processing and fusion in urban and rural planning surveys are solved, and efficient and accurate data analysis and planning support are achieved.
Patent Information
- Application Number
- CN202510156074.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-12
AI Technical Summary
In the urban and rural planning survey, the existing technology has problems such as separate storage of environmental data from residents' feedback data, lack of effective data processing and feature extraction methods, and insufficient integration of data geospatial corrections, resulting in inaccurate and unreliable data analysis results.
By constructing an adaptive scale isolated forest model to preprocess environmental data, building a smart community connection platform to collect residents' feedback data, and using a convolutional neural network model to fuse multi-source data and modify it in combination with geographical information to achieve deep fusion and efficient processing of multi-source data.
It realizes deep integration and efficient processing of multi-source data, improves the accuracy and comprehensiveness of data analysis, provides solid data support for urban and rural planning, and enhances the scientificity and applicability of planning decisions.
Smart Images

Figure CN120145295A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of environmental monitoring, and particularly relates to a data collection method and system for urban and rural planning surveys. Background Art
[0002] In the current era of the widespread promotion of the concepts of smart cities and smart communities, the importance of environmental data monitoring and resident feedback collection in the field of urban and rural planning has become increasingly prominent. Precise, comprehensive, and real-time data support has become the core element for formulating scientific and reasonable urban and rural planning schemes, playing a decisive role in improving the quality of residents' lives, optimizing urban resource allocation, and promoting the sustainable development of urban and rural areas.
[0003] With the promotion of the concepts of smart cities and smart communities, the importance of environmental data monitoring and resident feedback collection in urban and rural planning has become increasingly prominent. Traditional urban and rural planning survey methods mainly rely on manual research and statistical analysis. Manual research not only requires a large amount of human, material, and time costs, but also has extremely low efficiency, making it difficult to meet the requirements for the timeliness of planning data in the rapid development process of urban and rural areas. Moreover, manual operations are extremely vulnerable to subjective factors, resulting in a significant reduction in the accuracy of data. At the same time, due to the lack of timely data updates, its timeliness cannot be effectively guaranteed, which makes the planning decisions based on such data out of touch with the actual development situation of urban and rural areas, seriously affecting the scientific and rational nature of urban and rural development layouts.
[0004] There are still a series of defects that cannot be ignored in the current related technologies. First, environmental data and resident feedback data are usually stored separately, which leads to the inability to achieve deep integration between multi-source data. It is difficult for planners to comprehensively grasp the actual situation of urban and rural development from a macro level, lacking comprehensive and systematic data support when making planning decisions, severely restricting the scientific and comprehensive nature of planning decisions. Second, although multi-dimensional environmental sensors and smart community platforms have improved the data collection efficiency, there is a lack of effective methods in data processing and feature extraction. This greatly reduces the accuracy and reliability of data analysis results, making it impossible to fully explore the deep value hidden behind the data and difficult to provide accurate and effective guidance for urban and rural planning practices. In addition, there are obvious deficiencies in the geographical space correction and integration of existing technologies, and the influence of geographical location on environmental data has not been fully considered. The environmental conditions vary greatly in different geographical locations. If this key factor is ignored during the planning process, it will lead to poor adaptability of the planning scheme to the actual geographical environment, thereby reducing the effectiveness and feasibility of the planning. Summary of the Invention
[0005] The object of the present invention is to solve the problems that the environmental data and the residents' feedback data in the urban and rural planning survey data collection are stored separately and difficult to integrate, the lack of effective data processing and feature extraction methods leads to inaccurate and unreliable analysis results, and the lack of sufficient geographical space correction and fusion of data cannot consider the influence of geographical location on environmental data. By constructing an adaptive-scale isolation forest model to preprocess environmental data, building a smart community connection platform to collect residents' feedback data, and using a convolutional neural network model to fuse multi-source data and combine geographical information correction, the deep fusion, efficient processing and accurate analysis of multi-source data are realized, providing comprehensive and reliable data support for urban and rural planning.
[0006] To solve the above technical problems, the technical solution provided by the present invention is a data collection method and system for urban and rural planning surveys, including the following steps: Step 1: Collect environmental data; The urban and rural multi-dimensional environmental sensors are used to collect environmental data, and the urban and rural multi-dimensional environmental sensors include an air quality sensor, a noise sensor, a temperature sensor, a humidity sensor and a barometric pressure sensor.
[0007] The environmental data includes PM2.5 concentration data, PM10 concentration data, CO2 concentration data, NO2 concentration data, environmental noise level data, noise level data, environmental temperature data, air humidity data and atmospheric pressure data.
[0008] Step 2: Construct an adaptive-scale isolation forest model to preprocess the environmental data obtained in Step 1; Step 2.1: Conduct a preliminary cleaning of the environmental data, remove null values, and generate a clean data set; Step 2.2: Construct an adaptive-scale isolation forest model, use the clean data set obtained in Step 2.1 as the input of the adaptive-scale isolation forest model, and the output of the adaptive-scale isolation forest model is the determination result of whether the environmental data is an outlier; Step 2.2.1: Through an adaptive strategy, randomly extract the time data samples and spatial dimension data samples of each data point from the clean data set; Step 2.2.2: Use multiple decision trees to capture the abnormal feature data of the environmental data and generate an abnormal score according to the abnormal features; The expression of the abnormal score is: ; (1) In the formula is the abnormal score of the th data point, is the detection function of the adaptive-scale isolation forest model, which is a part of the named function , is the An environmental data point, where i is the index of the data point; Step 2.2.3: Set the anomaly threshold. If the anomaly score obtained in Step 2.2.2 exceeds the anomaly threshold, it is determined as an outlier. Set the anomaly threshold. When the anomaly score calculated by Equation (1) exceeds the set anomaly threshold, it is determined that the data point is an outlier. Step 2.2.4: Remove the environmental data determined as outliers from the clean data set. Step 2.3: Standardize the clean data set obtained by the method in Step 2.2.4 to generate standardized environmental data.
[0009] The standardized environmental data, with the expression: ; (2) ; (3) In the formula, The value of the standardized environmental data, is the environmental data value, is the current time point, is a constant to prevent the denominator from being zero, is the position coordinate, is the measurement error of the urban-rural multi-dimensional environmental sensor, is the correction function, is the sensor type, is the accuracy decay coefficient during the long-term use of the sensor, is the coefficient of the sensor type, is the accuracy decay coefficient of the sensor calibrated in the short term, is the constant term.
[0010] Step 3: Collect the resident feedback data corresponding to the environmental data in Step 1 and combine it with the data in Step 2 to form multi-source data. Step 3.1: Develop a smart community connection platform using responsive web design and UX Design user experience design. The front end of the smart community connection platform provides a resident questionnaire interface. Step 3.2: Generate a dynamic form on the resident questionnaire interface and verify the environmental data. The back end of the smart community connection platform stores and manages the questionnaire data. Step 3.3: Store the resident questionnaire data in a NoSQL database and encrypt it using a symmetric encryption algorithm. Step 3.4: Based on the developed resident questionnaire interface, users fill in the questionnaire online. Step 3.5: Set multiple questionnaire types on the resident questionnaire interface to collect resident feedback data. Step 3.6: Establish real-time communication between the front end and the server of the smart community connection platform, display feedback information, and use the message queue RabbitMQ to process and transmit the residents' feedback data; Step 3.7: Adopt token authentication authorization and use the HTTPS protocol to encrypt the transmission of residents' feedback data; Step 3.8: Develop a cross-platform application Flutter; Step 3.9: Use Discourse to establish a community forum and generate community interaction data; Step 3.10: Use Gephi to analyze the community interaction data; Step 3.11: Store the preprocessed environmental data and residents' feedback data in step 2 in the central database to generate multi-source data.
[0011] Step 4: Use the first neural network model to extract and fuse features from the multi-source data to form a unified multi-dimensional data set for evaluating the degree of environmental pollution; Step 4.1: Standardize the environmental data and residents' feedback data in the multi-source data through the Z-score algorithm; Step 4.2: Use a convolutional neural network model to extract features from the environmental data and residents' feedback data in the multi-source data. The input of the convolutional neural network model is the multi-source data, and the output is a multi-dimensional feature data set; Step 4.2.1: Convolve the multi-source data through multiple convolutional kernels in the convolutional layer to extract local features and generate feature maps; Step 4.2.2: Extract the main features from the feature maps through the pooling layer and output compressed feature maps; Step 4.2.3: Use the fully connected layer to perform linear combination and non-linear activation on the compressed feature maps output by the pooling layer to generate high-dimensional feature data; Step 4.2.4: The output layer of the convolutional neural network converts and unifies the high-dimensional feature data through linear transformation and the Softmax activation function, and outputs a multi-dimensional feature data set; Step 4.2.5: Combine geographic information to correct the position of the unified multi-dimensional data set. The expression is: ; (4) In the formula, is the multi-dimensional data set after position correction, is the integral symbol, is the environmental data value, is the time variable, is the geographic correction function; Step 4.3: Obtain the final unified multi-dimensional data set.
[0012] The data collection method further includes step 5: using a second neural network model to evaluate the degree of environmental pollution, where the input of the second neural network model is a unified multi-dimensional data set and the output is an evaluation value of the degree of environmental pollution; The second neural network model is a multi-layer perceptron model.
[0013] The present invention provides a data collection system for a data collection method for urban and rural planning surveys, including a collection module, a preprocessing module, a construction module, and a fusion module; the collection module collects environmental data through urban and rural multi-dimensional environmental sensors and performs preprocessing; the preprocessing module is used to construct an adaptive-scale isolation forest model to preprocess the environmental data; the construction module is used to construct a smart community connection platform, generate resident feedback data, and combine the preprocessed environmental data and resident feedback data into multi-source data and store it in a central database; the fusion module is used to extract and fuse features of the environmental data and resident feedback data in the multi-source data through a convolutional neural network model to form a unified multi-dimensional data set.
[0014] The present invention provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the computer program is executed by the processor, any step of the data collection method for urban and rural planning surveys described in the present invention is implemented.
[0015] The present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by the processor, any step of the data collection method for urban and rural planning surveys described in the present invention is implemented.
[0016] Compared with the prior art, the beneficial effects of the present invention include: (1) The present invention realizes the deep fusion of multi-source data, improves the comprehensiveness of planning decisions, breaks the situation of separate storage of environmental data and resident feedback data in traditional technologies, combines the preprocessed environmental data and resident feedback data into multi-source data and stores it in a central database, and uses a convolutional neural network model to extract and fuse features to form a unified multi-dimensional data set, effectively realizing the deep fusion of multi-source data. This fusion method can comprehensively reflect the actual situation of urban and rural areas, provide more comprehensive and relevant data support for urban and rural planning, make planning decisions no longer rely solely on single-dimensional data, significantly improve the scientificity and comprehensiveness of planning decisions, and make the planning results more in line with the actual needs and environmental status of urban and rural residents.
[0017] (2) By using a convolutional neural network model to extract and fuse the environmental data and resident feedback data in the multi-source data, a unified multi-dimensional dataset is formed, realizing the deep feature extraction and fusion of the data. Local features are extracted through the convolutional layer, main features are extracted through the pooling layer, high-dimensional feature data is generated through the output layer, and linear combination and non-linear activation are performed through the fully connected layer to generate a unified multi-dimensional feature dataset. Combined with geographical information for position correction, the accuracy of the data in the geographical space is ensured, and a high-dimensional feature dataset containing rich information is generated, providing a comprehensive high-dimensional feature representation for subsequent data analysis, ultimately improving the accuracy and comprehensiveness of data analysis and providing solid data support for urban and rural planning.
[0018] (3) Different from the deficiencies in the data processing and feature extraction methods in the prior art, the present invention constructs an adaptive scale isolation forest model to preprocess the environmental data. By initially cleaning the environmental data, extracting data samples using an adaptive strategy, capturing abnormal features using multiple decision trees, setting abnormal thresholds and processing outliers, and performing standardization processing, null values and abnormal data in the environmental data are effectively removed, improving the accuracy and reliability of the data. At the same time, the introduced correction function can fully consider sensor types and accuracy attenuation factors, further ensuring the quality of the data, laying a solid foundation for subsequent data analysis and applications, and making the analysis results more credible and valuable for reference.
[0019] (4) The present invention constructs a smart community connection platform, comprehensively applying advanced technologies, developing the interface using Responsive Web Design and UX Design to ensure good display and operation experiences on different mobile devices; using the front-end framework React and the back-end framework Node.js for development to achieve efficient data processing and response; implementing real-time communication and data transmission through WebSocket and RabbitMQ to ensure timely information updates; protecting user account security and data privacy using the OAuth2.0 and HTTPS protocols; developing a cross-platform application Flutter to enable the platform to run stably on different operating systems; also establishing a community forum function using Discourse to encourage resident interaction, generating rich community interaction data, and analyzing it with the help of Gephi. The application of these technologies not only improves the collection efficiency of resident feedback data but also greatly enhances the user experience, increasing the enthusiasm and initiative of residents to participate in urban and rural planning surveys, and helping to obtain more comprehensive and authentic resident feedback information.
[0020] (5) The present invention fully considers the influence of geographical location on environmental data, and corrects the position of the unified multi-dimensional data set through combining geographical information correction. The geographical correction function is used to process the data, enabling the data to accurately reflect the actual situation of different geographical locations, effectively improving the applicability of the data in the geographical space, enabling the urban and rural planning analysis based on this data to more precisely consider the characteristics and differences of different regions, providing strong support for formulating planning schemes tailored to local conditions, helping to improve the pertinence and effectiveness of urban and rural planning, and promoting the sustainable development of urban and rural areas.
[0021] (6) The present invention uses a convolutional neural network model to extract and fuse features from multi-source data. The convolutional neural network model adopted can deeply excavate the complex features and internal relationships hidden in the multi-source data. Multiple convolutional kernels can extract local features of the multi-source data and generate feature maps, comprehensively reflecting information in different dimensions of the data; the max-pooling operation in the pooling layer screens key features, reducing the dimension while retaining core information; the fully connected layer deeply fuses the features of environmental and resident feedback data to generate high-dimensional feature data; the output layer finally converts it into a unified multi-dimensional feature data set. This method overcomes the data format differences and semantic gaps between data sources, realizes deep fusion, provides a complete and accurate decision-making basis, improves the quality and efficiency of data fusion, enhances the usability of the data, and strongly promotes the intelligent and scientific development of urban and rural planning.
[0022] (7) The present invention uses a multi-layer perceptron model to achieve a comprehensive assessment of the degree of environmental pollution based on the unified multi-dimensional data set containing resident feedback data. The comprehensive assessment method makes full use of the multi-source fusion data, covering the objective data of environmental monitoring and the subjective feedback of residents, enabling the assessment results to more comprehensively and accurately reflect the actual pollution situation. The powerful learning ability of the multi-layer perceptron model can excavate the complex non-linear relationships between data and accurately quantify the degree of pollution. The assessment results provide a key basis for urban and rural planning, helping planners scientifically and reasonably layout urban functional areas, plan traffic routes, and optimize the configuration of public facilities for different pollution regions and degrees, and can also assist in formulating effective pollution control and prevention strategies, promoting the urban and rural environment to develop in a green and sustainable direction and improving the quality of life of residents. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0024] Figure 1 It is a flowchart of the data collection method for urban and rural planning investigation in the embodiments of the present invention.
[0025] Figure 2 This is a module diagram of the data collection system for urban and rural planning surveys in the embodiments of the present invention.
[0026] Figure 3 This is a flowchart for data collection and processing in urban and rural planning in the embodiments of the present invention.
[0027] Figure 4 This is an architecture diagram of the intelligent community connection platform system in the embodiments of the present invention.
[0028] Figure 5 This is a schematic diagram of the interface of the data collection system for urban and rural planning surveys in the embodiments of the present invention.
[0029] Figure 6 This is a schematic diagram of the data processing flow in the embodiments of the present invention.
[0030] Figure 7 This is a line graph showing the change in PM2.5 concentration data of a single index in the embodiments of the present invention.
[0031] Figure 8 This is a combined line graph showing data changes in the embodiments of the present invention. Detailed implementation manners
[0032] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe the detailed implementation manners of the present invention with reference to the accompanying drawings of the specification.
[0033] In the following description, many specific details are set forth to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0034] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that can be included in at least one implementation manner of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment that excludes other embodiments.
[0035] Such as Figure 1 and Figure 2 and Figure 3 As shown, a data collection method for urban and rural planning surveys includes the following steps: Step 1: Collect environmental data; Determine the survey objectives and scope; Design specific implementation plans for the intelligent perception network and the intelligent community connection platform; Form an investigation team, including technical personnel, investigators, and data analysts; Develop a detailed implementation plan, including the sensor layout plan and the development plan of the intelligent community connection platform; According to the implementation plan, the technical team arranges air quality sensors, noise sensors, temperature sensors, humidity sensors, and barometric pressure sensors in the target area as the basis of the urban and rural multi-dimensional environmental sensors to collect environmental data, including PM2.5 concentration data, PM10 concentration data, CO2 concentration data, NO2 concentration data, environmental noise level data, noise level data, environmental temperature data, air humidity data, and atmospheric pressure data; Ensure the stable operation of the urban and rural multi-dimensional environmental sensor network and the real-time upload of environmental data; The urban and rural multi-dimensional environmental sensors transmit environmental data to the central database through a wireless network. The real-time upload of environmental data ensures freshness and accuracy. The environmental data includes timestamp, location coordinates, and environmental parameters; Each urban and rural multi-dimensional environmental sensor can simultaneously monitor multiple environmental parameters, including air quality PM2.5 concentration data, air quality PM10 concentration data, air quality CO2 concentration data, air quality NO2 concentration data, environmental noise level, noise level in a specific frequency band, environmental temperature, air humidity, and atmospheric pressure; Use high-precision sensor components to ensure the accuracy and reliability of environmental data; The urban and rural multi-dimensional environmental sensors have a self-calibration function and can perform self-detection and calibration regularly to ensure the accuracy of environmental data during long-term use; All environmental data collected by the urban and rural multi-dimensional environmental sensors will be transmitted through encryption to ensure the security of environmental data; The urban and rural multi-dimensional environmental sensors adopt a low-power design to extend the battery life and adapt to various outdoor environments; According to the environmental data collection requirements, set the monitoring parameters and upload frequency of the urban and rural multi-dimensional environmental sensors; The urban and rural multi-dimensional environmental sensors collect environmental data in real time; It should be noted that by arranging urban and rural multi-dimensional environmental sensors and designing an intelligent perception network and an intelligent community connection platform, environmental data can be efficiently collected and transmitted. The sensors cover multiple parameters such as air quality, noise, temperature, humidity, and barometric pressure, ensuring comprehensive and accurate data. The self-calibration function and encrypted transmission guarantee the long-term reliability and security of the data. The low-power design and real-time upload function of the sensors enhance the practicability and stability, effectively supporting the dynamic monitoring and planning management of urban and rural environments.
[0036] Step 2: Construct an adaptive-scale isolated forest model to preprocess the environmental data obtained in Step 1; Step 2.1: Conduct preliminary cleaning on the environmental data, remove null values, and generate a clean data set; Step 2.2: Construct an adaptive-scale isolation forest model, use the clean data set obtained in Step 2.1 as the input of the adaptive-scale isolation forest model, and the output of the adaptive-scale isolation forest model is the determination result of whether the environmental data is an outlier; Step 2.2.1: Through an adaptive strategy, randomly extract the time data samples and spatial dimension data samples of each data point from the clean data set to ensure that the adaptive-scale isolation forest model can capture the change characteristics of the data at different time and spatial scales; The adaptation strategy combines the characteristics of the existing isolation forest algorithm and expands it to enable it to work effectively at different scales and adapt to the complexity and dynamic changes of environmental data; Step 2.2.2: Use multiple decision trees to capture the abnormal characteristics of the environmental data and generate an abnormal score according to the abnormal characteristics. The expression of the abnormal score is: ; (1) In the formula, is the abnormal score of the th data point, is the detection function of the adaptive-scale isolation forest model, which is used to calculate the abnormal score of the data point, is the adaptive scale, which is a part of the naming function and is used to indicate that this function has the function of adaptively adjusting the scale, is the th environmental data point, and i is the index of the data point; By calculating the abnormal score of each data point, the outliers in the data can be effectively identified; Step 2.2.3: Set an abnormal threshold. If the abnormal score obtained in Step 2.2.2 exceeds the abnormal threshold, it is determined as an outlier; Set an abnormal threshold. When the abnormal score calculated by formula (1) exceeds the set abnormal threshold, it is determined that this data point is an outlier; The threshold is set as: When > 1.55, the data point is determined as an outlier; When ≤ 1.55, the data point is determined as a normal value; Step 2.2.4: Remove the environmental data determined as outliers from the clean data set; Remove the detected outliers and replace them with the mean value of adjacent data points to ensure the continuity and rationality of the data. The adaptive algorithm dynamically adjusts the processing strategy according to the current distribution of the data to ensure the rationality and effectiveness of outlier processing; Specifically, multiple data subsets are randomly selected from the clean dataset; For each data subset, a decision tree is trained. These decision trees can identify normal and abnormal patterns in the data by learning the data distribution and features; Each decision tree evaluates the entire dataset and calculates the path length of each data point. Data points with shorter path lengths are usually considered abnormal because they are more likely to be isolated by the decision tree; The evaluation results of each decision tree are combined to generate a comprehensive anomaly score for each data point; By combining the anomaly scores generated by all decision trees, the final anomaly score for each data point can be obtained through the average value or other weighted methods; The anomaly scores are standardized to ensure that the distribution of the anomaly scores is applicable to subsequent threshold setting; Based on the standardized anomaly scores, an anomaly threshold is set. When the anomaly score calculated by Equation (1) exceeds the set anomaly threshold, the data point is determined to be an outlier; Step 2.3: Standardize the clean dataset obtained by the method in Step 2.2.4 to generate standardized environmental data.
[0037] The clean dataset after being detected by the adaptive scale isolation forest model is standardized to generate standardized environmental data. The expression is: ; (2) In the formula, is the value of the standardized environmental data, representing the data at a certain time point and location, combining sensor type, health status, and environmental parameters, is the environmental data value, is the time difference from the initial time to the time point , used to characterize the time dynamic change of the data, is a constant to prevent the denominator from being zero, is the position coordinate, used to represent the specific geographical location of the data collection point, The measurement error of the urban-rural multi-dimensional environmental sensor generates a correction coefficient by combining the sensor type and health status, is the sensor type, such as air quality sensor, noise sensor, etc. is the accuracy attenuation coefficient during the long-term use of the sensor. For example, the proportion of sensor performance decline obtained through regular calibration or the long-term deviation between the actual measurement value and the standard value. A coefficient can be assigned for different accuracy attenuation situations; The correction function The expression of is: ; (3) In the formula, is the coefficient of the sensor type, which represents the influence of different sensor types on the correction value. For example, for the air quality sensor, it is , and for the noise sensor, it is etc., is the coefficient of the accuracy decay of the sensor calibration in the short term, and the constant term is obtained by fitting experimental data or historical data and is used to balance and optimize the output of the correction function; It should be noted that by preprocessing environmental data through the adaptive scale isolation forest model, null values and abnormal data can be effectively removed, and outliers can be accurately identified and processed through dynamic threshold setting. Standardization processing ensures data consistency, improves analysis reliability, uses multiple decision trees to capture data features, and combines with the correction function to improve accuracy. Efficient data encryption transmission and real-time monitoring ensure data security and stable operation of the device.
[0038] Step 3: Collect the resident feedback data corresponding to the environmental data in Step 1 and combine it with the data in Step 2 to form multi-source data, as Figure 4 shown.
[0039] Step 3.1: Develop a smart community connection platform using Responsive Web Design and User Experience Design (UX Design), and display and operate it through mobile devices. The front-end framework React of the smart community connection platform provides a resident questionnaire interface, as Figure 5 shown; Responsive Web Design is a web design method that enables web pages to automatically adjust the layout and content display on different devices and screen sizes. By using flexible grid layouts, resizable images, and CSS media query techniques, web pages can be correctly displayed on various devices such as mobile phones, tablets, and computers, ensuring that the web page content is clear, easy to read, and easy to operate regardless of the device used by the user, avoiding creating separate versions for different devices and reducing development and maintenance costs; User Experience Design (UX Design) is a design process aimed at enhancing the overall experience when users interact with a product or service. By conducting research and analysis to understand users' needs, behaviors, and pain points, formulating design strategies, organizing and structuring content, enabling users to easily find the information they need, designing interfaces and interaction elements to ensure the convenience and pleasure of user operations, and continuously optimizing the design through testing and feedback to improve user satisfaction; Step 3.2: Generate a dynamic form on the resident questionnaire survey interface and verify the environmental data. The back-end Node.js development framework of the smart community connection platform stores and manages the questionnaire survey data; React is an open-source JavaScript library developed by Facebook for building user interfaces, especially single-page applications (SPAs). It allows developers to create reusable UI components to build complex user interfaces. React provides a component-based development model that enables developers to split the UI into independent and reusable components, improving development efficiency and code maintainability. React uses virtual DOM to enhance performance. Every time the state changes, React generates a new virtual DOM tree and compares it with the old virtual DOM, only updating the parts of the actual DOM that need to be changed. React adopts a one-way data flow, where data flows from the parent component to the child component, ensuring data controllability and ease of debugging. React ensures that residents can conveniently fill out the questionnaire and submit feedback data; Node.js is a JavaScript runtime environment based on the V8 engine that allows developers to run JavaScript code on the server side. It uses an event-driven, non-blocking I / O model and is suitable for building high-concurrency, high-performance network applications. Node.js can be used to build server-side applications to implement functions such as APIs and web servers. Node.js supports asynchronous programming. Through the event loop and callback mechanism, it can handle a large number of concurrent requests without blocking threads. Node.js comes with Node Package Manager, which is a rich package management tool and ecosystem center that enables developers to easily integrate and use various third-party libraries and modules. Node.js achieves efficient data processing and response, ensuring stable operation and handling of a large number of user requests; Step 3.3: Use a NoSQL database to store the questionnaire survey data submitted by residents and encrypt the questionnaire survey data through AES encryption to ensure data security; Step 3.4: Based on the developed resident questionnaire survey interface, residents fill out the questionnaire online and submit opinions and suggestions; Step 3.5: Set multiple questionnaire survey types on the resident questionnaire survey interface to comprehensively collect resident feedback data; Step 3.6: The front end realizes real-time communication with the server of the smart community connection platform through WebSocket, displays the latest feedback information, and uses the message queue RabbitMQ to process and transmit real-time resident feedback data to ensure timely update of information; WebSocket is a communication protocol that provides a full-duplex communication channel, allowing for real-time data exchange between clients and servers. WebSocket enables a persistent connection between the client and the server, thus achieving real-time data transmission. Different from the traditional HTTP request-response mode, WebSocket supports two-way communication, enabling the server to actively push data to the client, rather than just responding to the client's requests. Since the WebSocket connection is persistent, the latency of data transmission is greatly reduced, making it very suitable for application scenarios that require real-time updates, such as online games, real-time chat, and real-time data monitoring. Through WebSocket, the front-end can maintain a persistent connection with the server, receiving and displaying the latest feedback information in real time, ensuring the immediacy and interactivity of data; RabbitMQ is an open-source message queue center used to implement message transmission, processing, and distribution. It supports multiple message protocols, such as AMQP, STOMP, and MQTT. RabbitMQ can transfer messages between different applications and process these messages to achieve asynchronous communication. RabbitMQ uses queues to cache and distribute messages, so that even if the receiver is temporarily unavailable, the messages will not be lost. Through the mechanism of message queues, RabbitMQ can achieve load balancing, evenly distributing messages to multiple receivers, improving processing capabilities and reliability. RabbitMQ supports message persistence and confirmation mechanisms to ensure the reliability and data integrity of messages during transmission. In the background, RabbitMQ is responsible for processing and transmitting real-time data. Through the message queue mechanism, it ensures the reliability of data transmission and load balancing, enabling efficient handling of a large number of concurrent requests; Step 3.7: Protect user account security through the authentication and authorization mechanism OAuth 2.0, and use the HTTPS protocol to encrypt all resident feedback data transmission to ensure the privacy of user data; OAuth 2.0 is an authorization framework mainly used to enable third-party applications to obtain resource permissions of users on other services. It achieves this function through the interaction between the authorization server and the resource server. OAuth 2.0 allows applications to access resources on behalf of users without exposing the users' usernames and passwords. Through the authorization process, the application obtains an access token, and with this token, it can access the users' resources. By using the access token, OAuth 2.0 reduces the risk of user credential leakage. The access token has an expiration date and can be refreshed as needed. Users can control and manage the access permissions of third-party applications and revoke the authorization at any time to ensure the controllability of data access. By implementing OAuth 2.0, the smart community connection platform can effectively protect the security of user accounts, ensure the privacy and security of user data during transmission, thereby enhancing users' trust and the security of the platform; Step 3.8: Develop a cross-platform application Flutter to enable the smart community connection platform to run smoothly on different mobile devices and ensure that the platform can run smoothly on different operations such as Android and iOS; Flutter allows developers to create applications that run on different platforms such as Android, iOS, Web, and desktops using a single codebase. It uses the Dart language and provides rich components and tools to help developers quickly build beautiful user interfaces. Flutter allows developers to write a single set of code and run it on multiple platforms, thus significantly reducing development and maintenance costs. Through Flutter's rendering engine Skia, applications can achieve high performance and a smooth experience close to native applications. Flutter provides a series of predefined and customizable UI components, enabling developers to easily build beautiful and feature-rich user interfaces. Flutter supports the "hot reload" function, that is, developers can immediately see the effect of changes after modifying the code, greatly improving development efficiency. By using Flutter to develop the smart community connection platform, it ensures that the platform can run smoothly on different operations such as Android and iOS, not only reducing development and maintenance costs but also providing high performance and a consistent user experience; Step 3.9: Use the open-source forum software Discourse to establish a community forum function, allowing residents to post and comment on posts, and add comment and like functions to encourage interaction among residents and generate community interaction data; Step 3.10: Establish a community forum function using the open-source forum software Discourse. Residents can post and comment on posts and participate in community discussions. In addition, comment and like functions are added to encourage resident interaction, increase the activity and sense of participation in the community, and generate valuable community interaction data by collecting and analyzing this interaction data; Discourse is a modern forum software written in Ruby on Rails with Ember.js used for the front end. It aims to replace traditional forum software and enhance the interaction and engagement of online communities through a better user experience and functional design. Discourse provides a platform that allows users to create posts, replies, and comments, promoting communication and interaction among community members. Through the real-time notification function, users can instantly learn about new posts, replies, and likes, enhancing interactivity. Discourse supports categorization, tagging, and topic management, helping users organize and find content they are interested in. Through trust levels, Discourse encourages users to actively participate and increases their permissions and function usage as they level up. Discourse offers a rich set of plugins that can extend functions such as voting, event management, and point centers. The Discourse interface has a responsive design, ensuring good display and operation on mobile devices. By using Discourse, the smart community connection platform can create a highly interactive and engaging community forum, providing a space for residents to freely communicate and share, thereby enhancing community cohesion and resident satisfaction; Use the social network analysis tool Gephi to analyze community interaction data and enhance the sense of community participation; Gephi is an interactive graph visualization platform that allows users to import, manipulate, and analyze graph data, helping to discover hidden patterns and trends. It is widely used in fields such as social network analysis, network science, data mining, and information visualization. Gephi provides powerful visualization tools that can display complex network data in graphical form, enabling users to intuitively observe the network structure and the relationships between nodes. Through various analysis algorithms, Gephi can calculate and display key metrics of the network, such as the degree centrality, betweenness centrality, and clustering coefficient of nodes, helping users understand the nature and structure of the network. Gephi supports dynamic network analysis and can show how the network structure changes over time, which is suitable for analyzing dynamic interactions in social networks. Gephi has a rich plugin ecosystem, and users can add functional modules according to their needs to expand its analysis and visualization capabilities. As a powerful social network analysis tool, Gephi helps the smart community connection platform enhance residents' sense of participation and interactivity and promote community development and management by analyzing and visualizing community interaction data; Step 3.11: Store the preprocessed environmental data in Step 2 and the resident feedback data submitted through the smart community connection platform in the central database to generate multi-source data.
[0040] It should be noted that by constructing a smart community connection platform and combining Responsive Web Design and UX Design, a consistent user experience on various devices is ensured. React and Node.js are used to develop the front and back ends to ensure efficient data processing and response. WebSocket and RabbitMQ are combined to achieve real-time communication and data transmission, enhancing immediacy and reliability. OAuth 2.0 and HTTPS are utilized to protect user data privacy. Cross-platform applications are developed using Flutter to enhance the user experience. Discourse and Gephi are adopted to promote community interaction and enhance the sense of community participation. This comprehensive approach improves the comprehensiveness, security, and user engagement of data collection.
[0041] Step 4: Extract and fuse the environmental data and resident feedback data in the multi-source data through a convolutional neural network model to form a unified multi-dimensional dataset for evaluating the degree of environmental pollution, as Figure 6 shown; Sort and classify the environmental data and resident feedback data in the multi-source data to ensure the consistency and integrity of the data; Step 4.1: Standardize the environmental data and resident feedback data in the multi-source data through the Z-score normalization algorithm to ensure a unified data format and eliminate differences between different data sources; Step 4.2: Use a convolutional neural network model to extract features from the environmental data and resident feedback data in the multi-source data, generate high-dimensional feature data, and fuse them into a unified multi-dimensional dataset; Define the input of the convolutional neural network model as the multi-source data and the output as the multi-dimensional feature dataset; Perform convolution on the multi-source data through multiple convolutional kernels in the convolutional layer to extract local features and generate feature maps. The convolutional layer can capture local patterns in the data and generate feature maps; Specifically, the convolutional kernel is a small matrix, usually of size 3x3 or 5x5, which slides on the input multi-source data; The convolutional kernel slides step by step on the input multi-source data. Each time it slides, the dot product of the convolutional kernel and the corresponding area of the input multi-source data is calculated to generate a feature value; Through the sliding window and convolution operations, a two-dimensional feature map can be generated, which contains the local features of the input multi-source data; After the convolution operation, an activation function is applied to perform a non-linear transformation on the feature map to increase the expressive power of the convolutional neural network model; The convolutional layer outputs multiple feature maps. Each feature map corresponds to the features extracted by a convolutional kernel, and these feature maps represent the local features of the input multi-source data; Extract the air quality, temperature, humidity, and noise in the environmental data, as well as the feedback frequency in the resident feedback data, from the feature map through max pooling in the pooling layer. The pooling layer can reduce the dimension of the feature map, reduce the computational amount, and retain the main feature information; Specifically, apply a window with a fixed size on the feature map, and select the maximum value within each window as the pooling result; The window slides step by step on the feature map, and the maximum value within the window is calculated each time it slides; Reduce the dimension of the feature map through pooling compression, reduce the computational amount, and retain the main feature information, such as the air quality, temperature, humidity, and noise in the environmental data, as well as the feedback frequency in the resident feedback data; The output of the pooling layer is the compressed feature map, which retains the main feature information in the input multi-source data; Use the fully connected layer to perform linear combination and non-linear activation on the main features output by the pooling layer to generate high-dimensional feature data; Specifically, flatten the feature map output by the pooling layer into a one-dimensional vector for subsequent processing; Take the dot product of the flattened one-dimensional vector and the weight matrix of the fully connected layer to generate the result of the linear combination, and add a bias term to enhance the expressive ability of the convolutional neural network model; Apply a non-linear activation function to activate the result of the linear combination to increase the non-linear expressive ability of the convolutional neural network model; According to the importance and relevance of the main features of the environmental data features and the resident feedback data features, assign different weights to different features to further optimize the feature fusion effect and generate high-dimensional feature data; For high-weight features such as the PM2.5 and PM10 concentration features and the feedback frequency in the environmental data features and the resident feedback data features, give high-weight features to highlight their importance; For medium-weight features such as the CO2 concentration, environmental temperature, and satisfaction score in the environmental data features and the resident feedback data features, give medium-weight features to reflect their medium importance; For low-weight features such as the air humidity, suggestions, and opinions in the environmental data features and the resident feedback data features, give low-weight features to reduce their impact on the final feature fusion; The output layer converts and unifies the high-dimensional feature data through linear transformation and the Softmax activation function, and outputs a multi-dimensional feature data set. These feature data represent the deep feature representation of the input data; The output layer performs a linear transformation on the high-dimensional feature data to make it adapt to the final output form. For example, use the weighted summation operation of the fully connected layer to combine different features into a unified data representation; The output layer can process the linearly transformed data using an appropriate activation function to ensure the non - linear characteristics and range of the output data. For example, the Softmax activation function commonly used in classification tasks can convert the output into a probability distribution; The unified multi - dimensional feature dataset generated by the output layer represents the model's final understanding and feature representation of the input data, and these feature data can be used for further analysis, decision - making, or other tasks; Combine geographical information to correct the position of the unified multi - dimensional dataset to ensure the accuracy of the data in the geographical space. The expression is: ; (4) In the formula, is the multi - dimensional dataset after position correction, is the integral symbol, is the environmental data value, that is, the original environmental data, is the time variable, representing a small increment of the time variable , is the geographical correction function, representing the geographical correction function related to the position coordinates . This function is used to correct and adjust the data to ensure that the influence of geographical location on the data is fully considered.
[0042] It should be noted that through the convolutional neural network model, feature extraction and fusion of multi - source data are carried out to form a unified multi - dimensional dataset, realizing the consistency and integrity of the data. The Z - score normalization algorithm is used to eliminate the differences between data sources to ensure the uniformity of the data format. The convolutional layer and pooling layer extract deep - level features, outputting high - dimensional feature data, and the fully - connected layer further optimizes the feature fusion effect. Combining geographical information correction to ensure the accuracy of the data in the geographical space, finally providing a comprehensive and accurate high - dimensional feature dataset, providing solid data support for urban and rural planning decision - making.
[0043] In one embodiment, the unified multi - dimensional dataset uses a multi - layer perceptron model to evaluate the degree of environmental pollution. The input is defined as the unified multi - dimensional dataset, and the output is the evaluation value of the degree of environmental pollution.
[0044] Initialize the parameters of the multi - layer perceptron model; The input layer inputs the unified multi - dimensional dataset after fusion, and the number of features is n; Define two hidden layers. The number of nodes in the first hidden layer is n 1 , and the number of nodes in the second hidden layer is n 2 ; The number of nodes in the output layer is 1, which is used to output the continuous evaluation value of the degree of environmental pollution; From the input layer to the first hidden layer is: Weight matrix The dimension is ; Weight value ; In the formula, is the weight from the th input feature of the input layer to the th hidden layer node of the first hidden layer; Bias vector The dimension is ; From the first hidden layer to the second hidden layer: Weight matrix The dimension is ; Weight value ; In the formula, is the weight from the th input feature of the first hidden layer to the th hidden layer node of the second hidden layer; Bias vector The dimension is ; From the second hidden layer to the output layer: Weight matrix The dimension is ; Weight value ; In the formula, is the weight from the th input feature to the output layer; Bias vector ; The input sample for forward propagation is x, with dimension n; The input of the first hidden layer is , where is the transpose; the output is = , where is the ReLU activation function.
[0045] The input of the second hidden layer is , where is the transpose, and the output is , where is the ReLU activation function.
[0046] The input of the output layer is , is the transpose; the output is the predicted value ; The mean squared error loss function is as follows: ; where is the number of samples, is the true value of the i-th sample. is the predicted value of the i-th sample.
[0047] The error of the output layer calculated by backpropagation is ; The error of the second hidden layer is ; where is the derivative of the ReLU activation function, denotes element-wise multiplication.
[0048] The error of the first hidden layer is ; where is the derivative of the ReLU activation function, denotes element-wise multiplication.
[0049] Calculate the gradients of the weights and biases: For the gradient is ; For the gradient is ; For the gradient is ; For the gradient is ; For the gradient is ; For the gradient is ; The Adam optimizer updates the first moment estimate , the second moment estimate respectively as:
[0050]
[0051] where is the gradient at iteration t, β 1 is the decay rate of the first moment estimate, β 2 is the decay rate of the second moment estimate; The corrected first moment estimate , the second moment estimate respectively as: ; ; wherein, the corrected first moment estimate, the corrected second moment estimate; Update the model parameters as follows: ; wherein, θt is the parameter at iteration t, θt−1 is the parameter at iteration t-1, α is the learning rate, and ϵ is a small constant to prevent the denominator from being zero; Evaluate the model; The root mean square error RMSE is: ; The mean absolute error MAE is: ; Precisely extract relevant data from the unified multi-dimensional dataset according to the environmental indicators required for analysis.
[0052] Select an appropriate chart type for plotting according to the data characteristics and analysis purposes.
[0053] For showing the change trend of a single environmental indicator over time, plot a line chart by setting the x-axis as the time point and the y-axis as the corresponding environmental indicator value to clearly present the change trend of the data; As Figure 7 described, with the help of data query statements, filter out all data records related to the PM2.5 concentration in the dataset to ensure that the obtained data covers the measurement values at different time points. At the same time, extract the time point information corresponding to each PM2.5 concentration data to construct a complete dataset structure. For other indicators such as temperature, humidity, and noise level, use the same data filtering logic to ensure that the obtained data is accurate and complete, providing a reliable basis for subsequent chart plotting.
[0054] For multiple environmental indicators with an association relationship, display their relationship through a combined chart or...
[0055] As Figure 8As shown, the maximum temperature is represented by a blue broken line, with the scale and label indicating the unit as "°C", which intuitively reflects the changing trend of the maximum temperature. The minimum temperature is represented by a light blue broken line, also labeled with the unit "°C", forming a contrast with the maximum temperature line to show the fluctuation of the temperature range. For the sulfur dioxide concentration and ozone concentration, they are represented by green and pink broken lines respectively. By presenting them in the same chart, it is convenient to observe the possible correlations between them and the temperature changes. In addition, the wind direction data is represented by a cyan broken line. Although the wind direction data may be different from other indicators in terms of dimension and meaning, presenting it in the same chart helps to comprehensively analyze the co-variation of multiple environmental factors. Thus, through this multi-broken line and dual-axis chart presentation method, the changing trends of various environmental indicators and their potential correlations with each other can be presented comprehensively and intuitively.
[0056] Observe the drawn chart and analyze the changing patterns and trends of the environmental data. From a single-indicator chart, determine whether the environmental indicator exceeds the normal range or standard value. For example, from the PM2.5 concentration data change chart, compare it with the value specified in the air quality standard. If it is found that the PM2.5 concentration continuously exceeds the standard value during a certain period, it indicates that the air quality is polluted during that period. At the same time, by comparing different indicator charts, explore the potential correlations between the data, and consult relevant environmental science materials to learn that such temperature and humidity changes affect the diffusion and chemical reactions of pollutants in the atmosphere, thereby affecting air quality.
[0057] By comprehensively analyzing the chart, determine the severity of environmental pollution, the type of pollution, as well as the time and area where the pollution occurs, providing a strong basis for formulating subsequent treatment strategies.
[0058] A data collection system for a data collection method used in urban and rural planning surveys in one embodiment includes: a collection module, a preprocessing module, a construction module, and a fusion module; the collection module collects environmental data through urban and rural multi-dimensional environmental sensors and performs preprocessing; the preprocessing module is used to construct an adaptive-scale isolation forest model to preprocess the environmental data; the construction module is used to construct an intelligent community connection platform, generate resident feedback data, and combine the preprocessed environmental data and resident feedback data into multi-source data and store it in the central database; the fusion module is used to extract and fuse the features of the environmental data and resident feedback data in the multi-source data through a convolutional neural network model to form a unified multi-dimensional data set, and also includes using a multi-layer perceptron model to evaluate the degree of environmental pollution.
[0059] A computer device used in one embodiment for a data collection method applicable to urban and rural planning surveys includes: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the data collection method for urban and rural planning surveys as proposed in the above embodiment.
[0060] The computer device may be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a carrier network, near-field communication NFC, or other technologies. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covering the display screen, or may also be a button, a trackball, or a touchpad provided on the housing of the computer device, or may also be an external keyboard, a touchpad, or a mouse, etc.
[0061] In one embodiment, a storage medium is adopted, on which a computer program is stored, and when the program is executed by a processor, it implements the data collection method for urban and rural planning surveys as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof.
[0062] In summary, the present invention extracts and fuses the environmental data and resident feedback data in multi-source data by using a convolutional neural network model to form a unified multi-dimensional data set, realizing the deep feature extraction and fusion of data. The local features are extracted through the convolutional layer, the main features are extracted through the pooling layer, the high-dimensional feature data is generated through the output layer, and the linear combination and non-linear activation are performed through the fully connected layer to generate a unified multi-dimensional feature data set. The position is corrected by combining geographical information to ensure the accuracy of the data in the geographical space. Its function is to generate a high-dimensional feature data set containing rich information, providing a comprehensive high-dimensional feature representation for subsequent data analysis, and ultimately improving the accuracy and comprehensiveness of data analysis, providing solid data support for urban and rural planning.
[0063] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not restrictive. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A data collection method for urban and rural planning survey, characterized in that: The following steps are involved: Step 1: Collect environmental data; Step 2: Construct an adaptive scale isolation forest model and preprocess the environmental data obtained in step 1; Step 3: Collect the resident feedback data corresponding to the environmental data in step 1 and combine it with the data in step 2 to form multi-source data; Step 4: Use the first neural network model to extract and fuse features from multi-source data to form a unified multidimensional data set for assessing the degree of environmental pollution.
2. The data collection method for urban and rural planning survey according to claim 1 is characterized in that: In the step 1, the environmental data is collected by using urban and rural multi-dimensional environmental sensors, which include air quality sensors, noise sensors, temperature sensors, humidity sensors and air pressure sensors.
3. The data collection method for urban and rural planning survey according to claim 2 is characterized in that: In step 1, the environmental data includes PM2.5 concentration data, PM10 concentration data, CO2 concentration data, NO2 concentration data, environmental noise level data, noise level data, environmental temperature data, air humidity data and atmospheric pressure data.
4. The data collection method for urban and rural planning survey according to claim 3 is characterized in that: The step 2 includes the following sub-steps: Step 2.1: Clean the environmental data and generate a clean data set; Step 2.2: construct an adaptive scale isolation forest model, and use the clean data set obtained in step 2.1 as the input of the adaptive scale isolation forest model. The output of the adaptive scale isolation forest model is the determination result of whether the environmental data is an outlier; Step 2.2.1: Extract data samples of the clean dataset using an adaptive strategy; Step 2.2.2: Use multiple decision trees to capture abnormal features of environmental data, and generate anomaly scores based on the abnormal features; Step 2.2.3: Set an abnormal threshold. If the abnormal score obtained in step 2.2.2 exceeds the abnormal threshold, it is judged as an abnormal value; Step 2.2.4: Remove the environmental data that are judged as outliers from the clean data set; Step 2.3: Standardize the clean data set obtained by the method in step 2.2.4 to generate standardized environmental data.
5. The data collection method for urban and rural planning survey according to claim 4 is characterized in that: In step 2, the adaptive scale isolation forest model randomly extracts a time data sample and a spatial dimension data sample of each data point from the clean data set through an adaptive strategy; Use multiple decision trees to capture abnormal feature data in clean data sets and generate anomaly scores; The anomaly score expression is: ;(1) In the formula, For the The anomaly score of the data point, is the detection function of the adaptive scale isolation forest model, For the environmental data points, i is the index of the data point; Set an anomaly threshold. When the anomaly score calculated by formula (1) exceeds the set anomaly threshold, the data point is determined to be an outlier. The clean data set after the adaptive scale isolation forest model detection is standardized to generate standardized environmental data, which is expressed as: ;(2) ;(3) In the formula is the standardized environmental data value, is the environmental data value, is the current time point, is the initial time, To prevent the denominator from being zero, is the position coordinate, is the measurement error of the urban and rural multi-dimensional environmental sensor, is the correction function, Indicates the sensor type, is the accuracy attenuation coefficient of the sensor during long-term use, is the coefficient of the sensor type, is the accuracy reduction coefficient of the sensor calibrated in the short term, is a constant term.
6. The data collection method for urban and rural planning survey according to claim 5 is characterized in that: The step 3 includes the following sub-steps: Step 3.1: Develop a smart community connection platform. The front end of the smart community connection platform provides a resident questionnaire survey interface; Step 3.2: Generate a dynamic form on the resident questionnaire interface and verify the environmental data. The back-end storage and management of the questionnaire data of the smart community connection platform; Step 3.3: Use NoSQL database to store the resident questionnaire data and encrypt it using symmetric encryption algorithm; Step 3.4: Based on the developed resident questionnaire interface, the user fills in the questionnaire; Step 3.5: Set up multiple questionnaire types on the resident questionnaire survey interface to collect resident feedback data; Step 3.6: The front end communicates with the server of the smart community connection platform in real time and displays feedback information, using the message queue RabbitMQ to process and transmit resident feedback data; Step 3.7: Use token authentication authorization and HTTPS protocol to encrypt the transmission of resident feedback data; Step 3.8: Develop cross-platform application Flutter; Step 3.9: Use Discourse to build a community forum and generate community interaction data; Step 3.10: Use Gephi to analyze community interaction data; Step 3.11: Store the environmental data preprocessed in step 2 and the resident feedback data in the central database to generate multi-source data.
7. The data collection method for urban and rural planning survey according to claim 6 is characterized in that: The first neural network model is a convolutional neural network model, and step 4 includes the following sub-steps: Step 4.1: Standardize the environmental data and resident feedback data in the multi-source data using the Z-score algorithm; Step 4.2: Use a convolutional neural network model to extract features from environmental data and resident feedback data in multi-source data. The input of the convolutional neural network model is multi-source data, and the output is a multi-dimensional feature data set; Step 4.2.1: Convolve the multi-source data through multiple convolution kernels of the convolution layer, extract local features, and generate feature maps; Step 4.2.2: Extract the main features from the feature map through the pooling layer and output the compressed feature map; Step 4.2.3: Use the fully connected layer to perform linear combination and nonlinear activation on the compressed feature map output by the pooling layer to generate high-dimensional feature data; Step 4.2.4: The output layer of the convolutional neural network transforms and unifies the high-dimensional feature data through linear transformation and Softmax activation function, and outputs a multi-dimensional feature data set; Step 4.2.5: Combine geographic information to correct the position of the unified multidimensional dataset. The expression is: ;(4) In the formula, is the position-corrected multidimensional dataset, is the integral symbol, is the environmental data value, t is the time variable, is the geographic correction function; Step 4.3: Get the final unified multidimensional cube.
8. The data collection method for urban and rural planning survey according to claim 7 is characterized in that: The data collection method further comprises step 5: using a second neural network model to evaluate the degree of environmental pollution, wherein the input of the second neural network model is a unified multidimensional data set, and the output is an evaluation value of the degree of environmental pollution; The second neural network model is a multi-layer perceptron model.
9. The data collection system for the data collection method for urban and rural planning survey according to any one of claims 1 to 7, characterized in that ,including collection module, preprocessing module, building module and fusion module; The collection module collects environmental data through urban and rural multi-dimensional environmental sensors and performs pre-processing; The preprocessing module is used to construct an isolation forest model and preprocess the environmental data; The construction module is used to construct a platform to generate resident feedback data, and combine the pre-processed environmental data and resident feedback data into multi-source data and store them in a database; The fusion module is used to extract and fuse the environmental data and resident feedback data in the multi-source data through the first neural network model to form a unified multidimensional data set.
10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the data collection method for urban and rural planning survey described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Low-carbon strategy optimization method and system for residential users
CN116822733A
Municipal garden water pollution monitoring system and method based on big data analysis
CN118886718A
Intelligent community resident health management method based on machine learning
CN119274802A
Social infrastructure control system
JP2024153964A