A three-dimensional space PM2.5 concentration estimation method, device, medium and product
By stitching together haze images from multiple perspectives and training an improved VIT model, the robustness and resolution issues of PM2.5 concentration estimation in existing technologies have been resolved, achieving more accurate three-dimensional spatial PM2.5 concentration estimation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ARTIFICIAL INTELLIGENCE INNOVATION RES INST OF ZHEJIANG UNIV OF TECH BINJIANG DISTRICT HANGZHOU
- Filing Date
- 2024-12-27
- Publication Date
- 2026-06-30
Smart Images

Figure CN122312461A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of atmospheric environmental monitoring technology, and in particular to a three-dimensional spatial PM2.5 concentration estimation method, equipment, medium and product. Background Technology
[0002] In PM2.5 concentration monitoring, image-understanding-based PM2.5 concentration estimation methods can be broadly categorized into two types: image feature-based methods and deep learning-based methods. Image feature-based PM2.5 concentration estimation primarily infers PM2.5 levels by constructing a mapping relationship between the PM2.5 index and various image features. These features include physical characteristics such as transmission, traditional visual features like color, contrast, saturation, texture, and edges. Some methods also utilize additional information such as weather conditions and depth of field. Relevant algorithms for modeling include linear regression, random forests, support vector machines, and decision trees. However, because image features are heavily influenced by image content, this type of approach lacks robustness to different scenarios and can only provide a coarse-grained picture of haze distribution. Most deep learning-based PM2.5 concentration estimation methods use single two-dimensional visible light images as observation data, which is insufficient for extracting spatial information. In particular, the scale of haze scenes under the perspective of drones varies greatly, and current deep learning methods have difficulty capturing multi-scale features, resulting in insufficient feature information and the spatial resolution of the prediction results needs to be improved. Summary of the Invention
[0003] The purpose of this application is to provide a method, device, medium, and product for estimating PM2.5 concentration in three-dimensional space, which can obtain accurate estimated PM2.5 concentration values.
[0004] To achieve the above objectives, this application provides the following solution:
[0005] In a first aspect, this application provides a three-dimensional spatial PM2.5 concentration estimation method, including:
[0006] Acquire haze images from multiple perspectives within the sample area and the actual PM2.5 concentration values within the sample area;
[0007] The haze images from multiple perspectives are stitched together to obtain a three-dimensional spatial sample image; the three-dimensional spatial sample image and the corresponding actual PM2.5 concentration value constitute a training sample;
[0008] Multiple training samples are input into the improved VIT model for training to obtain a PM2.5 concentration prediction model; the improved VIT model includes a multi-scale image feature embedding module and a Transformer encoder module arranged sequentially.
[0009] The three-dimensional spatial sample image corresponding to the area to be detected is input into the PM2.5 concentration prediction model to obtain the corresponding estimated PM2.5 concentration value.
[0010] Secondly, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement a three-dimensional spatial PM2.5 concentration estimation method.
[0011] Thirdly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a three-dimensional spatial PM2.5 concentration estimation method.
[0012] Fourthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements a three-dimensional spatial PM2.5 concentration estimation method.
[0013] According to the specific embodiments provided in this application, this application has the following technical effects: This application provides a three-dimensional PM2.5 concentration estimation method, device, medium, and product. It uses haze images from multiple perspectives within a sample area and the corresponding actual PM2.5 concentration values to train an improved VIT (Vision Transformer) model. The use of haze images from multiple perspectives allows for the combination of multi-view information, resulting in more comprehensive and three-dimensional data. This enables fine-grained haze reconstruction in three-dimensional space, improving the ability to estimate air pollution and resulting in higher accuracy of the final estimated PM2.5 concentration value. The improved VIT model includes a multi-scale image feature embedding module and a Transformer encoder module arranged sequentially. The multi-scale image feature embedding module can better handle scale changes by acquiring features from different receptive fields, improving the model's adaptability at different scales, enhancing fine-grained estimation of air pollution, and also promoting the accuracy of the final estimated PM2.5 concentration value. Attached Figure Description
[0014] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0015] Figure 1 This is an application environment diagram of a three-dimensional spatial PM2.5 concentration estimation method according to an embodiment of this application;
[0016] Figure 2 A flowchart illustrating a three-dimensional spatial PM2.5 concentration estimation method provided in an embodiment of this application;
[0017] Figure 3 This is a schematic diagram of drone photography provided in an embodiment of this application;
[0018] Figure 4 This is a schematic diagram of the structure of an improved VIT model provided in an embodiment of this application;
[0019] Figure 5 This is a schematic diagram of the structure of a multi-scale image feature embedding module provided in an embodiment of this application;
[0020] Figure 6 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0021] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0022] The three-dimensional spatial PM2.5 concentration estimation method provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be set up separately, integrated into server 104, or placed in the cloud or on another server. Terminal 102 can send training samples or 3D spatial sample images corresponding to the area to be detected to server 104, and then train the improved VIT model in server 104, or input the data into a PM2.5 concentration prediction model to obtain the corresponding estimated PM2.5 concentration value. Server 104 can provide feedback on the estimated PM2.5 concentration value to terminal 102, or provide feedback to terminal 102 indicating that the model training is complete. Furthermore, in some embodiments, the 3D spatial PM2.5 concentration estimation method can be implemented separately by server 104 or terminal 102.
[0023] The terminal 102 can be, but is not limited to, various desktop computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. The server 104 can be implemented using a standalone server or a server cluster consisting of multiple servers, or it can be a cloud server.
[0024] In one exemplary embodiment, such as Figure 2 As shown, a three-dimensional spatial PM2.5 concentration estimation method is provided. This method is executed by a computer device, specifically by a terminal or server alone, or by both a terminal and a server. In this embodiment, the method is applied to... Figure 1 Taking server 104 as an example, the explanation includes the following steps 201 to 204.
[0025] Step 201: Obtain haze images from multiple perspectives within the sample area and the actual PM2.5 concentration values within the sample area.
[0026] In one application example, the process of acquiring haze images from multiple perspectives within the sample area includes: using a drone to simultaneously capture haze images of the sample area from different perspectives to obtain haze images from multiple perspectives. Specifically, the drone uniformly captures images of the sample area's center at a preset altitude and along a preset trajectory to obtain haze images from multiple perspectives. Figure 3 As shown, in a specific application, the drone's flight path corresponds to a preset altitude of 160 meters, and the image focus is set at the center of the area 100 meters above the ground. The drone takes 32 images evenly along the flight path to collect 32 haze images from different perspectives.
[0027] In another application example, the actual PM2.5 concentration value within the sample area was collected using a haze sensor. Furthermore, when drones capture haze images, they can obtain the data through ordinary surveillance cameras or other cameras (such as mobile phone cameras), without relying on a specific device model, thus possessing versatility.
[0028] Step 202 involves stitching the haze images from multiple perspectives together to obtain a three-dimensional spatial sample image. In another application example, the step of stitching the haze images from multiple perspectives involves stitching together multiple haze images from a set of data (corresponding to the 32 haze images from different perspectives acquired in step 201 above), with multi-channel input. Taking the stitching of two images as an example, the following formula is used to stitch together any two haze images from different perspectives:
[0029]
[0030] in, A first-view image of haze with dimensions C*H*W. A second-view haze image of size C*H*W, F 2C*H*W The image is a stitched image with dimensions of 2C*H*W, where C is the number of channels, H is the height, and W is the width.
[0031] After stitching together haze images from multiple perspectives using the above formula, the resulting three-dimensional spatial sample image is matched with the corresponding actual PM2.5 concentration value to obtain a training sample. Multiple training samples can be obtained based on different sample regions and different times, such as obtaining 1500 sets of image data (i.e., 1500 training samples). Then, they can be divided into a training subset and a validation subset in an 8:2 ratio.
[0032] In one application example, based on the regional extent of the sample area, the 3D spatial sample image is divided into an n*n*n voxel grid. For example, the space of the sample area is divided into 250m... 3 A cell is divided into voxel grids, resulting in 4*4*4 voxel grids. Correspondingly, in the construction of a training sample, each voxel grid corresponds to an actual PM2.5 concentration value. That is, a training sample includes multiple voxel grids of a three-dimensional spatial sample image and the actual PM2.5 concentration value corresponding to each voxel grid.
[0033] Step 203: Input multiple training samples into the improved VIT model for training to obtain a PM2.5 concentration prediction model; the improved VIT model includes a multi-scale image feature embedding module and a Transformer encoder module set sequentially, such as... Figure 4 As shown.
[0034] In another application example, such as Figure 5 As shown, the multi-scale image feature embedding module (corresponding to the multi-layer image feature fusion embedding module in the figure) includes a block layer, multiple convolutional layers, and a chimera layer. The block layer is used to divide the received 3D spatial sample image into multiple image blocks according to different scales, such as 16*16, 8*8, and 4*4, thereby capturing feature information at different scales. Each convolutional layer corresponds to a scale and is used to convolve multiple image blocks corresponding to any scale to generate block embedding features; that is, a separate convolutional layer is used for each image block corresponding to a scale to generate block embedding features at different scales. The chimera layer is connected to all the convolutional layers and is used to add the embedding features corresponding to different scales to obtain multi-scale image features, which are then input to the Transformer encoder module.
[0035] The Transformer encoder module includes a LayerNorm layer, a multi-head attention mechanism layer, a Norm layer, and an MLP layer arranged in sequence; wherein, the input of the LayerNorm layer is residually connected to the output of the multi-head attention mechanism layer, and the input of the Norm layer is residually connected to the output of the MLP layer.
[0036] During the training of the improved VIT model, the root mean square error is used as the loss function, as shown below:
[0037]
[0038] Where L is the root mean square error loss value, y' i and y i , respectively, represent the estimated PM2.5 concentration value and the actual PM2.5 concentration value corresponding to the i-th three-dimensional spatial sample image, and N is the number of three-dimensional spatial sample images.
[0039] Based on the above loss function, iterative training is performed until the loss function value is minimized during model iteration training. At this point, the optimal weight parameters and the final PM2.5 concentration prediction model are obtained.
[0040] Step 204: Input the 3D spatial sample image corresponding to the area to be detected into the PM2.5 concentration prediction model to obtain the corresponding estimated PM2.5 concentration value. Corresponding to the application example in step 202 above, the 3D spatial sample image corresponding to the area to be detected is divided into multiple n*n*n voxel grids, and then the 3D spatial sample image corresponding to the area to be detected is input into the PM2.5 concentration prediction model to obtain the estimated PM2.5 concentration value corresponding to each voxel grid, thereby achieving a detailed estimation of the 3D spatial sample image.
[0041] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 6 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media to run. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements a three-dimensional spatial PM2.5 concentration estimation method.
[0042] Those skilled in the art will understand that Figure 6The structures shown are merely block diagrams of some structures related to the present application and do not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than shown in the figures, or combine certain components, or have different component arrangements. In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0043] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0044] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0045] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0046] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).
[0047] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0048] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0049] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A three-dimensional space PM2.5 concentration estimation method, characterized in that, The three-dimensional spatial PM2.5 concentration estimation method includes: Acquire haze images from multiple perspectives within the sample area and the actual PM2.5 concentration values within the sample area; The haze images from multiple perspectives are stitched together to obtain a three-dimensional spatial sample image; the three-dimensional spatial sample image and the corresponding actual PM2.5 concentration value constitute a training sample; Multiple training samples are input into the improved VIT model for training to obtain a PM2.5 concentration prediction model; the improved VIT model includes a multi-scale image feature embedding module and a Transformer encoder module arranged sequentially. The three-dimensional spatial sample image corresponding to the area to be detected is input into the PM2.5 concentration prediction model to obtain the corresponding estimated PM2.5 concentration value.
2. The three-dimensional spatial PM2.5 concentration estimation method according to claim 1, characterized in that, The process of acquiring haze images from multiple perspectives within the sample area includes: The sample area was photographed simultaneously from different perspectives using drones to obtain haze images from multiple viewpoints.
3. The three-dimensional spatial PM2.5 concentration estimation method according to claim 1, characterized in that, The multi-scale image feature embedding module includes a block layer, multiple convolutional layers, and a chirp layer; The block layer is used to divide the received three-dimensional spatial sample image into multiple image blocks according to different scales; Each of the convolutional layers corresponds to a scale, and the convolutional layer is used to: convolve multiple image blocks corresponding to any scale to generate block embedding features; The fusion layer is connected to all the convolutional layers, and the fusion layer is used to: add the embedding features corresponding to different scales to obtain multi-scale features of the image, and input them into the Transformer encoder module.
4. The three-dimensional spatial PM2.5 concentration estimation method according to claim 1, characterized in that, The Transformer encoder module includes a LayerNorm layer, a multi-head attention mechanism layer, a Norm layer, and an MLP layer arranged sequentially. The input of the LayerNorm layer is residually connected to the output of the multi-head attention mechanism layer, and the input of the Norm layer is residually connected to the output of the MLP layer.
5. The three-dimensional spatial PM2.5 concentration estimation method according to claim 1, characterized in that, During the training of the improved VIT model, the root mean square error is used as the loss function, as shown below: Wherein, L is the root mean square error loss value, y i and y i are the estimated PM2.5 concentration value and the actual PM2.5 concentration value corresponding to the i th three-dimensional space sample image respectively, and N is the number of three-dimensional space sample images.
6. The three-dimensional spatial PM2.5 concentration estimation method according to claim 1, characterized in that, In the step of stitching together the haze images from multiple perspectives, the following formula is used to stitch together any two haze images from different perspectives: in, A first-view image of haze with dimensions C*H*W. A second-view haze image of size C*H*W, F 2C*H*W The image is a stitched image with dimensions of 2C*H*W, where C is the number of channels, H is the height, and W is the width.
7. The three-dimensional spatial PM2.5 concentration estimation method according to claim 2, characterized in that, When using a drone to capture haze images of the sample area from different perspectives at the same time, the drone uniformly captures images of the center of the sample area at a preset altitude and along a preset trajectory to obtain haze images from multiple perspectives.
8. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the three-dimensional spatial PM2.5 concentration estimation method according to any one of claims 1-7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the three-dimensional spatial PM2.5 concentration estimation method according to any one of claims 1-7.
10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the three-dimensional spatial PM2.5 concentration estimation method according to any one of claims 1-7.