Urban governance data classification method, system, equipment and medium

By porting the MapReduce algorithm from MapReduce to the Dag platform and optimizing its execution, the efficiency problem of MapReduce when processing massive data in urban governance is solved, and more efficient and accurate data management and services are achieved.

CN119989120APending Publication Date: 2025-05-13INSPUR SOFTWARE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510148080.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

MapReduce is not good at real-time computing, streaming computing, DAG computing, and high consumption and low energy consumption when processing massive data in urban governance, which affects the efficiency of data classification.

Method used

By reconstructing the combined architecture of MapReduce algorithm and Dag, the MapReduce algorithm is ported from MapReduce to the Dag platform, the execution of MapReduce algorithm in Dag is optimized, and the support vector machine is used for data prediction and classification.

Benefits of technology

It improves the management and service efficiency of urban governance data, reduces the number of disk read and writes during the calculation process, improves the fault tolerance and resource utilization, and reduces energy consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989120A_ABST
    Figure CN119989120A_ABST
Patent Text Reader

Abstract

The invention discloses an urban governance data classification method, system and device and a medium, belongs to the technical field of urban management and governance, and aims to solve the technical problem of how to realize more accurate and efficient management and service of mass data in urban governance. Each row of the file comprises a city name field, a pollutant field, an air quality grade field, an AQI field, a monitoring time field, an outdoor temperature field, a sendible temperature field and a travel suggestion field; a pollutant data file is read through a pseudo entry, operation is executed in parallel through a tool function provided by MapReduce, and then key data extraction is achieved; training the extracted key data through a support vector machine, and predicting a business classification result of the urban governance data; and outputting a business classification result of the urban governance data through the pseudo-output node exit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of urban management and governance technology, and specifically to a method, system, device and medium for classifying urban governance data. Background Art

[0002] With the continuous advancement of smart city construction, the amount of data related to urban governance has increased dramatically, which has put forward higher requirements for the magnitude, accuracy and efficiency of data processing. As an efficient distributed computing model, MapReduce performs well in processing large-scale data sets. Because MapReduce is easy to program, has strong scalability and high fault tolerance, and can meet the offline processing of massive data above PB level, it is widely used in urban governance. MapReduce is the core framework for users to develop "hadoop-based data analysis applications". Its core function is to integrate the business logic code written by users and the default components into a complete distributed computing program, which runs concurrently on a hadoop cluster. Although MapReduce has the advantages of easy programming, strong scalability, high fault tolerance and massive data processing, its disadvantages are also significant:

[0003] ① Not good at real-time computing: MapReduce is not suitable for returning results within milliseconds or seconds, so it is not suitable for applications that require low latency.

[0004] ② Not good at streaming computing: The input data of streaming computing is dynamic, while the input data set of MapReduce is static and cannot change dynamically. This is because the design characteristics of MapReduce itself determine that the data source must be static.

[0005] ③ Not good at DAG (directed graph) calculation: Multiple applications have dependencies, and the input of the next application is the output of the previous one. In this case, MapReduce is not incapable of doing it, but after using it, the output of each MapReduce job will be written to the disk, which will cause a lot of disk IO and lead to relatively low performance.

[0006] ④ High consumption and low energy problem: Since the design inspiration of MapReduce comes from functional programming languages ​​​​(such as: Lisp, Scheme, ML, etc.), and the energy consumption of the system was not taken into consideration during the architecture design, the high consumption and low energy problem is more serious.

[0007] From this we can see that the defects of MapReduce also affect the efficiency of urban governance data classification.

[0008] Therefore, how to achieve more accurate and efficient management and service of massive data in urban governance is a technical problem that needs to be solved urgently. Summary of the invention

[0009] The technical task of the present invention is to provide a method, system, device and medium for urban governance data classification to solve the problem of how to achieve more accurate and efficient management and service of massive data in urban governance.

[0010] The technical task of the present invention is achieved in the following way: a method for classifying urban governance data, which is to reconstruct the combined architecture of MapReduce algorithm and Dag, transplant the MapReduce algorithm from MapReduce to the Dag platform, and realize the optimization of MapReduce algorithm in Dag; the details are as follows:

[0011] Collect pollutant data in the form of files. Each line of the file contains a city name field, a pollutant field, an air quality level field, an AQI field, a monitoring time field, an outdoor temperature field, a perceived temperature field, and a travel advice field.

[0012] Read the pollutant data file through the pseudo entry node, and execute the job in parallel through the tool function provided by MapReduce to extract key data;

[0013] The key data extracted through support vector machine training is used to predict the business classification results of urban governance data;

[0014] The business classification results of urban governance data are output through the pseudo output node exit, and the business classification results of urban governance data are evaluated and optimized.

[0015] As a preferred method, the tool functions provided by MapReduce are used to execute jobs in parallel, thereby realizing key data extraction as follows:

[0016] The split function encapsulated in MapReduce is used to split the pollutant data file into 16MB file blocks, and the Map function encapsulated in MapReduce is called on the data in the file blocks;

[0017] Map the data in the file block into<key,value> Key-value pairs;

[0018] According to the key-value pairs, six types of pollution, PM2.5, PM10, O3, CO, NO2 and SO2, the Air Quality Index (AQI) and the air quality grade data are extracted separately and saved in the file.

[0019] Preferably, the key data extracted through support vector machine training can predict the business classification results of urban governance data as follows:

[0020] The acquired data is divided into training set and test set in a ratio of 9:1;

[0021] Through the support vector machine (SVM) binary classification model, with air quality level data as labels, six types of pollution data including PM2.5, PM10, O3, CO, NO2 and SO2, and air quality index data as input, the support vector machine binary classification model is trained to realize the predictive classification of air quality levels. The support vector machine binary classification model is verified through the test set to optimize the business classification results of urban governance data.

[0022] Preferably, the Dag platform includes a map interface, a reduce interface, a filter interface, a flatmap interface, and a union interface.

[0023] Preferably, the MapReduce algorithm is ported from MapReduce to the Dag platform as follows:

[0024] Data reading: Dag data is stored in RDD, which is distributed and stored in each slave node. RDD also provides a cache mechanism. Specifically, the cache mechanism is as follows: RDD caches the calculation results of the Map process in MapReduce through the persist() function or the cache() function. However, these two methods are not cached immediately when they are called. Instead, when the subsequent action() function is triggered, the RDD will be cached in the memory of the computing node and used for subsequent calculations.

[0025] Resource application: After the Dag task is started, the required Executor resources are applied immediately. All Dag tasks run in a threaded manner and share Executor resources.

[0026] Task management: Optimize multiple tasks in MapReduce into one Dag task through the Dag task scheduling model.

[0027] More optimally, the MapReduce algorithm is optimized in Dag as follows:

[0028] The pollutant data collected in the form of files is used as input data, and the dependencies of multiple task patterns are divided into different stages, and each stage is responsible for executing the corresponding MapReduce algorithm; wherein the MapReduce algorithm includes Map operation, groupByKey operation and flatMap operation; Map operation, groupByKey operation and flatMap operation are method functions encapsulated by Map operation, groupByKey operation and flatMap operation tools;

[0029] After the stage is executed, the execution result of each stage is submitted to TaskScheduler for execution. The task is executed on multiple Task threads of the Executor process to complete the Task task.

[0030] More preferably, the optimized energy consumption of the MapReduce algorithm in Dag refers to the time period T in which the task is completed and the value range is [t s ,t e ], the formula is: E = E cpu +E mem +E disk +E net ;

[0031] Where E represents the optimized energy consumption of the MapReduce algorithm in Dag; E cpu Indicates the energy consumption of the CPU; E mem Indicates the energy consumption of the disk; E disk Represents the energy consumption of memory; E net Indicates the energy consumption of the network card.

[0032] A city governance data classification system, which is used to implement the above-mentioned city governance data classification method; the system comprises:

[0033] The collection module is used to collect pollutant data in the form of files. Each line of the file contains a city name field, a pollutant field, an air quality level field, an AQI field, a monitoring time field, an outdoor temperature field, a perceived temperature field, and a travel suggestion field.

[0034] The data extraction module is used to read the pollutant data file through the pseudo-entry node and execute jobs in parallel through the tool functions provided by MapReduce, thereby realizing key data extraction;

[0035] The prediction module is used to predict the business classification results of urban governance data through the key data extracted by support vector machine training;

[0036] The evaluation and optimization module is used to output the business classification results of urban governance data through the pseudo output node exit, and to evaluate and optimize the business classification results of urban governance data.

[0037] Among them, directed acyclic graphs (DAGs) play a vital role in representing workflows and data processing pipelines. For example, in the popular orchestration tool Apache Airflow, workflows are defined as DAGs, where each node represents a task and the edges represent the order of execution. This structure allows data scientists and engineers to visualize and manage complex workflows, ensuring that data is processed in the correct order. In addition, directed acyclic graph (DAG) technology promotes parallel processing with its flexible data dependency management, executing independent tasks simultaneously, optimizing the computing process, thereby improving resource utilization and reducing overall processing time.

[0038] An electronic device comprising: a memory and at least one processor;

[0039] Wherein, the memory stores a computer program;

[0040] The at least one processor executes the computer program stored in the memory, so that the at least one processor performs the urban governance data classification method as described above.

[0041] A computer-readable storage medium having a computer program stored therein, wherein the computer program can be executed by a processor to implement the urban governance data classification method as described above.

[0042] The urban governance data classification method, system, device and medium of the present invention have the following advantages:

[0043] (i) The present invention can effectively represent the execution order and dependency of tasks in the process of big data task processing using MapReduce technology; each MapReduce task is a node in Dag, and the dependency between tasks is represented by edges. The execution order of tasks can be determined based on DDag, and the order of tasks with dependencies can be ensured to be completed;

[0044] (II) The DAG technology of the present invention adopts the storage method of Resilient Distributed Datasets (RDD), and RDD also provides a cache mechanism, which can cache the calculation results of MapReduce tasks for subsequent use; therefore, on the one hand, the DAG technology effectively reduces the number of disk reads and writes during the MapReduce task calculation process, and improves the execution efficiency of the MapReduce task; on the other hand, the DAG technology caches the calculation results of the MapReduce task, which effectively improves the fault tolerance of the MapReduce task, that is, once the MapReduce task fails to execute, the DAG can automatically recover from the failed MapReduce task node and recalculate according to the RDD data cached by the previous node, which effectively improves the robustness of the MapReduce task execution;

[0045] (III) The DAG technology of the present invention can reuse the calculation results of MapReduce tasks. When multiple MapReduce tasks process the same data set with the same operation, they can only process it once and cache the calculation results through RDD, which can be directly provided to subsequent tasks, thus reducing repeated calculations and improving execution efficiency.

[0046] (IV) The present invention invents a new computing method by combining MapReduce with another efficient computing technology (DAG), and provides a big data computing and analysis solution with high feasibility, high concurrency and low power consumption;

[0047] (V) The invention uses a combination of MapReduce and Dag in data reading, resource application, task management, etc., which makes the management and service of massive data in urban governance more accurate and efficient;

[0048] (Six) The present invention has broad application prospects in the fields of smart cities, etc., and will greatly improve the operating efficiency of smart cities. The data from urban management and operation services can be aggregated, calculated, analyzed and managed in a more efficient manner, improving urban governance data services and feeding back to smart city construction, enriching smart city construction plans. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] The present invention is further described below in conjunction with the accompanying drawings.

[0050] Attached Figure 1 A flowchart of the urban governance data classification method;

[0051] Attached Figure 2 It is the task operation relationship diagram of MapReduce algorithm;

[0052] Attached Figure 3 This is a schematic diagram of the optimization of the MapReduce algorithm in Dag;

[0053] Attached Figure 4 This is a comparison chart of the execution time of the optimized algorithm and the original algorithm in the present invention;

[0054] Attached Figure 5 This is a comparison chart of energy consumption between the optimized algorithm and the original algorithm in the present invention. DETAILED DESCRIPTION

[0055] The urban governance data classification method, system, device and medium of the present invention are described in detail below with reference to the drawings and specific embodiments of the specification.

[0056] Embodiment 1:

[0057] As attached Figure 1 As shown, this embodiment provides a method for classifying urban governance data. The method is to reconstruct the combined architecture of the MapReduce algorithm and Dag, transplant the MapReduce algorithm from MapReduce to the Dag platform, and realize the optimization of the MapReduce algorithm in Dag; the details are as follows:

[0058] S1. Collect pollutant data in the form of files. Each line of the file contains a city name field, a pollutant field, an air quality level field, an AQI field, a monitoring time field, an outdoor temperature field, a perceived temperature field, and a travel advice field;

[0059] S2, read the pollutant data file through the pseudo entry node, and execute the job in parallel through the tool function provided by MapReduce, so as to realize the key data extraction;

[0060] S3, through the key data extracted by support vector machine training, predict the business classification results of urban governance data;

[0061] S4. Output the business classification results of urban governance data through the pseudo output node exit, and evaluate and optimize the business classification results of urban governance data.

[0062] In order to facilitate scheduling, this embodiment uses a pseudo entry node and a pseudo exit node to connect all entry nodes and exit nodes respectively.

[0063] In step S2 of this embodiment, the tool function provided by MapReduce is used to execute jobs in parallel, thereby realizing key data extraction as follows:

[0064] S201, using the split function encapsulated in MapReduce to split the pollutant data file into file blocks of 16MB in size, and calling the Map function encapsulated in MapReduce on the data in the file blocks;

[0065] S202: Map the data in the file block into<key,value> Key-value pairs;

[0066] S203. According to the key-value pairs, six types of pollution, namely PM2.5, PM10, O3, CO, NO2 and SO2, the Air Quality Index (AQI) and the air quality grade data are extracted separately and saved in the file.

[0067] The key data extracted by support vector machine training in step S3 of this embodiment predicts the business classification results of urban governance data as follows:

[0068] S301, dividing the acquired data into a training set and a test set according to a ratio of 9:1;

[0069] S302. Through the support vector machine (SVM) binary classification model, with air quality level data as labels, six types of pollution data including PM2.5, PM10, O3, CO, NO2 and SO2, and air quality index data as input, the support vector machine binary classification model is trained to achieve predictive classification of air quality levels, and the support vector machine binary classification model is verified through the test set to optimize the business classification results of urban governance data.

[0070] The Dag platform in this embodiment includes a map interface, a reduce interface, a filter interface, a flatmap interface, and a union interface.

[0071] In this embodiment, the MapReduce algorithm is transplanted from MapReduce to the Dag platform as follows:

[0072] ① Data reading: Dag data is stored in RDD, which is distributed and stored in each slave node, and RDD provides a cache mechanism; the cache mechanism is as follows: RDD caches the calculation results of the Map process in MapReduce through the persist() function or the cache() function, but it is not cached immediately when these two methods are called. Instead, when the subsequent action() function is triggered, RDD will be cached in the memory of the computing node and used for subsequent calculations;

[0073] ② Resource application: After the Dag task is started, the required Executor resources are applied immediately. All Dag tasks run in a threaded manner and share Executor resources;

[0074] ③Task management: Optimize multiple tasks in MapReduce into one Dag task through the Dag task scheduling model.

[0075] As attached Figure 3 As shown in the figure, the optimization of MapReduce algorithm in Dag is as follows:

[0076] ① The pollutant data collected in the form of files is used as input data, and the dependencies of multiple task patterns are divided into different stages. Each stage is responsible for executing the corresponding MapReduce algorithm; the MapReduce algorithm includes Map operation, groupByKey operation and flatMap operation; Map operation, groupByKey operation and flatMap operation are method functions encapsulated by Map operation, groupByKey operation and flatMap operation tools;

[0077] ②After executing the stage, the execution results of each stage are submitted to TaskScheduler for execution. The task is executed on multiple Task threads of the Executor process to complete the Task.

[0078] This embodiment sets up a comparative experiment to verify the improvement effect of the algorithm in terms of operating efficiency after transplantation, and obtains a comparison chart of the execution time of the optimized algorithm and the original algorithm, as shown in the attached figure. Figure 4 shown.

[0079] The optimized energy consumption of the MapReduce algorithm in the Dag in this embodiment refers to the time period T in which the task is completed and the value range is [t s ,t e ], the formula is: E = E cpu +E mem +E disk +E net ; Statistically calculate the energy consumption of the optimized algorithm and the original algorithm, and obtain the energy consumption comparison chart, as shown in the attached Figure 5 As shown;

[0080] Where E represents the optimized energy consumption of the MapReduce algorithm in Dag; E cpu Indicates the energy consumption of the CPU; E mem Indicates the energy consumption of the disk; E disk Represents the energy consumption of memory; E net Indicates the energy consumption of the network card.

[0081] Embodiment 2:

[0082] This embodiment provides a city governance data classification system, which is used to implement the city governance data classification method in Example 1; the system includes:

[0083] The collection module is used to collect pollutant data in the form of files. Each line of the file contains a city name field, a pollutant field, an air quality level field, an AQI field, a monitoring time field, an outdoor temperature field, a perceived temperature field, and a travel suggestion field.

[0084] The data extraction module is used to read the pollutant data file through the pseudo-entry node and execute jobs in parallel through the tool functions provided by MapReduce, thereby realizing key data extraction;

[0085] The prediction module is used to predict the business classification results of urban governance data through the key data extracted by support vector machine training;

[0086] The evaluation and optimization module is used to output the business classification results of urban governance data through the pseudo output node exit, and to evaluate and optimize the business classification results of urban governance data.

[0087] The working process of the system is as follows:

[0088] Step 1: Decompose the task steps of the MapReduce algorithm, analyze and find the defects that affect efficiency. After MapReduce became the standard in the field of massive data parallel processing in industry and academia, Spark conducted an in-depth analysis of it and found that the execution efficiency of the MapReduce algorithm was relatively low. The task steps of the MapReduce algorithm are decomposed as shown in the following figure. Figure 1 The task relationship of the MapReduce algorithm is shown in the attached Figure 2 shown.

[0089] Step 2: Migrate the algorithm from MapReduce to the Dag platform and optimize the implementation process and migration method. Compared with MapReduce, Dag provides a more flexible programming interface. In addition to the map and reduce interfaces, there are also filter, flatmap, union and other interfaces. Dag programs are more flexible and convenient. The specific optimization plan for migrating the algorithm from MapReduce to the Dag platform is as follows:

[0090] ① In terms of data reading, Dag data storage RDD is more efficient than MapReduce data storage HDFS. RDD is distributed and stored in the memory of each slave node, reducing the number of disk reads and writes during the calculation process. RDD also provides a cache mechanism.

[0091] ② In terms of resource application, after the Dag task is started, it will immediately apply for the required Executor resources. All Dag tasks run in a threaded manner and share Executor resources. MapReduce manages Slot resources in a heartbeat manner. Dag greatly reduces the number of resource applications and has higher resource management efficiency.

[0092] ③In terms of task management, multiple tasks in MapReduce can be optimized into one Dag task through the Dag model.

[0093] ④Through the above optimization strategies, the combined architecture of MapReduce algorithm and Dag is reconstructed to achieve the additional Figure 3 The figure shows the optimized implementation architecture diagram of the MapReduce algorithm in Dag.

[0094] Step 3: Set up a comparative experiment to verify the improvement in operating efficiency after the algorithm is transplanted, and obtain a comparison chart of the execution time of the optimized algorithm and the original algorithm, as shown in the attached figure. Figure 4 shown.

[0095] Step 4: Set up a comparative experiment to verify the improvement effect of the algorithm transplantation in terms of running energy consumption. The present invention redesigns the energy consumption comparison calculation formula, wherein the energy consumption E during the execution of the algorithm task is composed of the sum of the energy consumption of the CPU, disk, memory, and network card, wherein the energy consumption E during the execution of the algorithm task is composed of the sum of the energy consumption of the CPU, disk, memory, and network card, wherein the CPU energy consumption is E cpu , the energy consumption of memory is E mem , the energy consumption of the disk is E disk , the energy consumption of the network card is E net The time period T after the task is completed is in the range of [t s ,t e ], the specific calculation process is as follows:

[0096] E=E cpu +E mem +E disk +E net ;

[0097] Through the above formula, the energy consumption of the optimized algorithm and the original algorithm is statistically calculated to obtain an energy consumption comparison chart, as shown in the attached figure. Figure 5 shown.

[0098] Embodiment 3:

[0099] This embodiment also provides an electronic device, including: a memory and a processor;

[0100] Wherein, the memory stores computer-executable instructions;

[0101] The processor executes the computer execution instructions stored in the memory, so that the processor executes the urban governance data classification method in any embodiment of the present invention.

[0102] The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor may be a microprocessor or any conventional processor, etc.

[0103] The memory can be used to store computer programs and / or modules. The processor realizes various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function, etc.; the data storage area can store data created according to the use of the terminal, etc. In addition, the memory can also include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (SMC), a secure digital (SD) card, a flash memory card, at least one disk storage period, a flash memory device, or other volatile solid-state storage devices.

[0104] Embodiment 4:

[0105] This embodiment also provides a computer-readable storage medium, in which a plurality of instructions are stored, and the instructions are loaded by a processor, so that the processor executes the urban governance data classification method in any embodiment of the present invention. Specifically, a system or device equipped with a storage medium can be provided, on which a software program code that implements the functions of any of the above embodiments is stored, and a computer (or CPU or MPU) of the system or device reads and executes the program code stored in the storage medium.

[0106] In this case, the program code itself read from the storage medium can realize the function of any one of the above-mentioned embodiments, and thus the program code and the storage medium storing the program code constitute a part of the present invention.

[0107] The storage medium embodiments for providing the program code include a floppy disk, a hard disk, a magneto-optical disk, an optical disk (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RYM, DVD-RW, DVD+RW), a magnetic tape, a non-volatile memory card, and a ROM. Alternatively, the program code can be downloaded from a server computer via a communication network.

[0108] In addition, it should be clear that the functions of any of the above embodiments can be implemented not only by executing the program code read by the computer, but also by enabling an operating system operating on the computer to complete part or all of the actual operations based on instructions from the program code.

[0109] In addition, it can be understood that the program code read from the storage medium is written to a memory provided in an expansion board inserted into the computer or written to a memory provided in an expansion unit connected to the computer, and then based on the instructions of the program code, a CPU installed on the expansion board or the expansion unit is enabled to perform part or all of the actual operations, thereby realizing the functions of any of the above-mentioned embodiments.

[0110] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for classifying urban governance data, characterized in that: This method is to reconstruct the combined architecture of MapReduce algorithm and Dag, transplant the MapReduce algorithm from MapReduce to the Dag platform, and optimize the MapReduce algorithm in Dag; the details are as follows: Collect pollutant data in the form of files. Each line of the file contains a city name field, a pollutant field, an air quality level field, an AQI field, a monitoring time field, an outdoor temperature field, a perceived temperature field, and a travel advice field. Read the pollutant data file through the pseudo entry node, and execute the job in parallel through the tool function provided by MapReduce to extract key data; The key data extracted through support vector machine training is used to predict the business classification results of urban governance data; The business classification results of urban governance data are output through the pseudo output node exit, and the business classification results of urban governance data are evaluated and optimized.

2. The urban governance data classification method according to claim 1 is characterized in that: The tool functions provided by MapReduce are used to execute jobs in parallel, thereby extracting key data as follows: The split function encapsulated in MapReduce is used to split the pollutant data file into 16MB file blocks, and the Map function encapsulated in MapReduce is called on the data in the file blocks; Map the data in the file block into<key,value> Key-value pairs; According to the key-value pairs, the six types of pollution, PM2.5, PM10, O3, CO, NO2 and SO2, the air quality index and the air quality grade data are extracted separately and saved in the file.

3. The urban governance data classification method according to claim 1 or 2 is characterized in that: The key data extracted through support vector machine training and the business classification results of predicting urban governance data are as follows: The acquired data is divided into training set and test set in a ratio of 9:1; Through the support vector machine binary classification model, with air quality level data as labels, six types of pollution data such as PM2.5, PM10, O3, CO, NO2 and SO2, and air quality index data as input, the support vector machine binary classification model is trained to achieve predictive classification of air quality levels. The support vector machine binary classification model is verified through the test set to optimize the business classification results of urban governance data.

4. The urban governance data classification method according to claim 3 is characterized in that: The Dag platform includes map interface, reduce interface, filter interface, flatmap interface and union interface.

5. The urban governance data classification method according to claim 4 is characterized in that: The details of porting the MapReduce algorithm from MapReduce to the Dag platform are as follows: Data reading: Dag data is stored in RDD, which is distributed and stored in each slave node. RDD also provides a cache mechanism. Specifically, the cache mechanism is as follows: RDD caches the calculation results of the Map process in MapReduce through the persist() function or cache() function. When the subsequent action() function is triggered, the RDD will be cached in the memory of the computing node and used for subsequent calculations. Resource application: After the Dag task is started, the required Executor resources are immediately applied for. All Dag tasks run in a threaded manner and share Executor resources. Task management: Optimize multiple tasks in MapReduce into one Dag task through the Dag task scheduling model.

6. The urban governance data classification method according to claim 5 is characterized in that: The optimization of MapReduce algorithm in Dag is as follows: The pollutant data collected in the form of files is used as input data, and the dependencies of multiple task patterns are divided into different stages, and each stage is responsible for executing the corresponding MapReduce algorithm; wherein the MapReduce algorithm includes Map operation, groupByKey operation and flatMap operation; Map operation, groupByKey operation and flatMap operation are method functions encapsulated by Map operation, groupByKey operation and flatMap operation tools; After executing the stage, the execution result of each stage is submitted to TaskScheduler for execution. The task is executed on multiple Task threads of the Executor process to complete the Task task.

7. The urban governance data classification method according to claim 6 is characterized in that: The optimized energy consumption of the MapReduce algorithm in Dag refers to the time period T in which the task is completed. The value range is [t s ,t e ], the formula is: E = E cpu +E mem +E disk +E net ; Where E represents the optimized energy consumption of the MapReduce algorithm in Dag; E cpu Indicates the energy consumption of the CPU; E mem Indicates the energy consumption of the disk; E disk Represents the energy consumption of memory; E net Indicates the energy consumption of the network card.

8. A city governance data classification system, characterized in that: The system is used to implement the urban governance data classification method as described in any one of claims 1 to 7; the system comprises: The collection module is used to collect pollutant data in the form of files. Each line of the file contains a city name field, a pollutant field, an air quality level field, an AQI field, a monitoring time field, an outdoor temperature field, a perceived temperature field, and a travel suggestion field. The data extraction module is used to read the pollutant data file through the pseudo-entry node and execute the job in parallel through the tool function provided by MapReduce, thereby realizing the key data extraction; The prediction module is used to predict the business classification results of urban governance data through the key data extracted by support vector machine training; The evaluation and optimization module is used to output the business classification results of urban governance data through the pseudo output node exit, and to evaluate and optimize the business classification results of urban governance data.

9. An electronic device, characterized in that: include: memory and at least one processor; Wherein, the memory stores a computer program; The at least one processor executes the computer program stored in the memory, so that the at least one processor executes the urban governance data classification method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which can be executed by a processor to implement the urban governance data classification method as described in any one of claims 1 to 7.