Air particulate matter concentration prediction method and system based on machine vision and hearing
By combining panoramic cameras and sensors, visual and acoustic features are extracted, and a predictive model is used to achieve efficient monitoring of particulate matter concentration on urban sidewalks. This solves the problem of limited coverage of air pollution data and improves monitoring accuracy and coverage.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-09
- Publication Date
- 2026-03-13
AI Technical Summary
The existing air pollution monitoring network is sparsely distributed, making it difficult to fully capture the spatial dynamics of road pollution. This results in limited coverage of air pollution data and an inability to effectively monitor the concentration of particulate matter in the air on urban sidewalks.
By acquiring panoramic street videos using panoramic cameras, combining them with meteorological data collected by sensors, extracting visual and acoustic features, and using a pre-trained air particulate matter concentration prediction model to predict short-term particulate matter concentration on urban sidewalks, a hybrid prediction model that integrates audiovisual variables and geospatial variables is used.
It has improved the monitoring coverage of air pollution data, accurately predicted the concentration of air particulate matter on urban sidewalks, and contributed to air quality management and public health protection.
Smart Images

Figure CN121661561A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of environmental monitoring and data analysis technology, specifically to a method and system for predicting air particulate matter concentration based on machine vision and hearing. Background Technology
[0002] Air pollution has become one of the most serious public health challenges globally, posing a severe threat to urban environments and human health. Road traffic emissions are widely recognized as a major source of urban air pollution. Pedestrian exposure to air pollution is particularly concerning, as pedestrians walking along sidewalks are chronically exposed to traffic emissions, increasing health risks and exacerbating the effects of air pollution. To monitor urban air pollution, governments and environmental agencies worldwide have established networks of fixed monitoring stations, including roadside sensors, for real-time air quality monitoring. These stations provide high-precision pollutant concentration data, but due to high installation and maintenance costs, these networks are often sparsely distributed, resulting in limited spatial coverage. Because air pollutant concentrations can fluctuate significantly over short distances and time periods, scattered and fixed monitoring networks struggle to fully capture the spatial dynamics of road pollution. Therefore, effectively monitoring roadside air pollution data is crucial for protecting public health. Summary of the Invention
[0003] This application aims to provide a method and system for predicting air particulate matter concentration based on machine vision and hearing, which can improve the coverage of air pollution monitoring data and accurately predict the air particulate matter concentration on urban sidewalks.
[0004] The technical solution of this application is implemented as follows: In a first aspect, embodiments of this application provide a method for predicting air particulate matter concentration based on machine vision and hearing, the method comprising: The system uses panoramic cameras to capture panoramic street videos and sensors to collect meteorological data. Feature extraction is performed on the panoramic street video to obtain audiovisual features; wherein, the audiovisual features include: visual features and sound features; Based on the soundscape features and the visual features, target soundscape features and target visual features are selected; wherein, the target soundscape features represent soundscape features that directly affect the air pollution level; and the target visual features represent visual features that directly affect the air pollution level. Based on the meteorological data, the target acoustic features, and the target visual features, a pre-determined air particulate matter concentration prediction model is used to predict the short-term particulate matter concentration of urban sidewalks, and the air particulate matter concentration prediction result is obtained; wherein, the air particulate matter concentration prediction model is trained based on historical street panoramic videos of sidewalks, historical meteorological data, and historical air particulate matter concentrations.
[0005] In the above scheme, the street panoramic video undergoes feature extraction to obtain audiovisual features, including: Based on the street panoramic video, video frames are extracted at a preset frame sampling rate to obtain multiple images; and visual features are extracted from the multiple images to obtain the visual features. Based on the panoramic street video, audio segments are extracted in preset time segments to obtain multiple audio segments; and sound scene features are extracted from the multiple audio segments using a pre-determined sound scene classification model to obtain the sound scene features.
[0006] In the above scheme, before extracting soundscape features from the multiple audio segments using a pre-determined soundscape classification model to obtain the soundscape features, the method further includes: Acquire a sound dataset and determine the research objective; wherein, the sound dataset includes multiple audio samples corresponding to multiple different sounds; the research objective is the air pollution situation in a preset area; Based on the research objectives, the sound dataset is classified to obtain multiple sound scene types; wherein, the multiple sound scene types include ambient sound, animal sound, car horn sound, car engine sound, vehicle traffic sound, construction noise, and pedestrian sound. Based on the multiple soundscape types and the sound dataset, the audio spectrogram converter model is trained to determine the soundscape classification model.
[0007] In the above scheme, the step of extracting soundscape features from the multiple audio segments using a pre-determined soundscape classification model to obtain the soundscape features includes: The soundscape classification model is used to convert the multiple audio segments into multiple spectrogram images. Based on the multiple spectrogram images, sound scene classification is performed using the sound scene classification model to obtain multiple sound classification values for each audio segment; The soundscape features are determined based on multiple sound classification values for each audio segment.
[0008] In the above scheme, the step of extracting visual features from the multiple images to obtain the visual features includes: Image segmentation and target detection are performed simultaneously on the multiple images to obtain multiple segmentation regions and multiple objects corresponding to each of the multiple images; The visual features are determined based on the multiple segmented regions and the multiple objects.
[0009] In the above scheme, the step of predicting short-term particulate matter concentration on urban sidewalks based on the meteorological data, the target acoustic features, and the target visual features, using a pre-determined air particulate matter concentration prediction model, to obtain the air particulate matter concentration prediction result, includes: The target visual features, the meteorological data, and the target soundscape features are matched to obtain a matched feature data set; wherein, the feature data set consists of target visual features, meteorological data, and target soundscape features within the same time period; Based on the aforementioned feature data set, the air particulate matter concentration prediction model is used to predict the short-term particulate matter concentration on urban sidewalks, resulting in the predicted air particulate matter concentration.
[0010] In the above scheme, after filtering based on the soundscape features and the visual features respectively, and selecting the target soundscape features and target visual features, the method further includes: Obtain geospatial data corresponding to the panoramic street video; wherein, the geospatial data includes static feature data characterizing the urban built environment; The target visual features, the meteorological data, and the target soundscape features are matched to obtain a set of matched feature data. Based on the feature data set and the geospatial data, a short-term particulate matter concentration prediction for urban sidewalks is performed using a pre-determined mixed prediction model for air particulate matter concentration, resulting in the air particulate matter concentration prediction result; wherein, the mixed prediction model for air particulate matter concentration represents a mixed prediction model that integrates audiovisual variables and geospatial variables.
[0011] Secondly, embodiments of this application provide an air particulate matter concentration prediction system based on machine vision and hearing. The system includes: an acquisition module, an extraction module, and a prediction module. The acquisition module is used to acquire panoramic street video through a panoramic camera and to collect meteorological data through a sensor. The extraction module is used to extract features from the street panoramic video to obtain audiovisual features; wherein, the audiovisual features include: visual features and sound features; based on the sound features and the visual features, target sound features and target visual features are selected respectively; wherein, the target sound features represent sound features that directly affect the air pollution level; the target visual features represent visual features that directly affect the air pollution level; The prediction module is used to predict the short-term particulate matter concentration of urban sidewalks based on the meteorological data, the target soundscape features, and the target visual features, using a pre-determined air particulate matter concentration prediction model, and to obtain the air particulate matter concentration prediction result; wherein, the air particulate matter concentration prediction model is trained based on historical street panoramic videos of the sidewalks, historical meteorological data, and historical air particulate matter concentrations.
[0012] Thirdly, embodiments of this application provide an air particulate matter concentration prediction device based on machine vision and hearing, comprising: a processor and a memory; wherein, The memory is used to store computer programs; The processor is configured to call and run the computer program from the memory to perform the method as described in the first aspect.
[0013] Fourthly, embodiments of this application provide a computer-readable storage medium storing executable instructions for causing a processor to perform the method described in the first aspect.
[0014] This application provides a method and system for predicting air particulate matter concentration based on machine vision and hearing. The method includes: acquiring panoramic street video using a panoramic camera and collecting meteorological data using sensors; extracting features from the panoramic street video to obtain audiovisual features; wherein the audiovisual features include visual features and soundscape features; filtering based on the soundscape features and visual features to select target soundscape features and target visual features; wherein the target soundscape features represent soundscape features that directly affect the air pollution level; the target visual features represent visual features that directly affect the air pollution level; and based on the meteorological data, the target soundscape features, and the target visual features, predicting short-term particulate matter concentration on urban sidewalks using a pre-determined air particulate matter concentration prediction model to obtain air particulate matter concentration prediction results; wherein the air particulate matter concentration prediction model is trained based on historical panoramic street video of sidewalks, historical meteorological data, and historical air particulate matter concentrations. In the above scheme, visual and acoustic features are extracted from panoramic street videos. These features are then filtered to obtain target acoustic and visual features. Based on meteorological data, the target acoustic and visual features, a short-term particulate matter concentration prediction model is used to predict the concentration of particulate matter on urban sidewalks. The acoustic features represent the acoustic features that directly affect the air pollution level, and the target visual features represent the visual features that directly affect the air pollution level. Therefore, the predicted particulate matter concentration is relatively accurate. In addition, since panoramic street videos are easy to collect, the coverage of air pollution monitoring data can be improved, and the air particulate matter concentration on urban sidewalks can be accurately predicted. This also helps in air quality management and public health protection. Attached Figure Description
[0015] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the specification, serve to explain the technical solutions of this application. Obviously, the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0016] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.
[0017] Figure 1 An optional flowchart illustrating an air particulate matter concentration prediction method based on machine vision and hearing provided in this application embodiment. Figure 1; Figure 2 An optional flowchart illustrating an air particulate matter concentration prediction method based on machine vision and hearing provided in this application embodiment. Figure 2 ; Figure 3 A schematic diagram of the structure of an air particulate matter concentration prediction system based on machine vision and hearing, provided for an embodiment of this application; Figure 4 This is a schematic diagram of the structure of an air particulate matter concentration prediction device based on machine vision and hearing, provided in an embodiment of this application. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the specific technical solutions of this application will be further described in detail below with reference to the accompanying drawings of the embodiments of this application. The following embodiments are used to illustrate this application, but are not intended to limit the scope of this application.
[0019] Unless otherwise defined, all technical and scientific terms used in this application have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. The terminology used in this application is for the purpose of describing embodiments of this application only and is not intended to be limiting of this application.
[0020] In the following description, references to "some embodiments," "this embodiment," "this application embodiment," and examples, etc., describe a subset of all possible embodiments. However, it is understood that "some embodiments" may be the same subset or different subset of all possible embodiments and may be combined with each other without conflict.
[0021] If the application documents contain similar descriptions such as "first / second", the following explanation shall be added: In the following description, the terms "first / second / third" are used only to distinguish similar objects and do not represent a specific order of objects. It is understood that "first / second / third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0022] This application provides a method for predicting air particulate matter concentration based on machine vision and hearing. Figure 1 An optional flowchart illustrating an air particulate matter concentration prediction method based on machine vision and hearing provided in this application embodiment. Figure 1 , will combine Figure 1 The steps shown are explained.
[0023] S101: Acquire panoramic street video through a panoramic camera and collect meteorological data through sensors.
[0024] In some embodiments of this application, the street panoramic video is video data collected from urban sidewalks.
[0025] In some embodiments of this application, the air particulate matter concentration prediction method based on machine vision and hearing is adapted to air pollution prediction scenarios.
[0026] In some embodiments of this application, the machine vision and hearing-based method for predicting air particulate matter concentration is adapted to a machine vision and hearing-based system for predicting air particulate matter concentration.
[0027] In some embodiments of this application, panoramic street video is captured by a panoramic camera along the sidewalk, and meteorological data is collected by an air pollution monitor. The meteorological data includes temperature and humidity.
[0028] For example, the portable air quality monitor AirBeam3 was used to monitor PM1 and PM2.5 concentrations. AirBeam has been widely used in previous mobile air pollution monitoring studies, demonstrating high efficiency and reliability in accurately collecting air pollution data. The device integrates a Plantower PMS7003 sensor to measure PM concentrations and collects meteorological and location data, including relative humidity, temperature, and geographic latitude and longitude, via meteorological sensors and GPS. Simultaneously, an Insta360 ONE RS panoramic camera was mounted on the helmet of each assistant to record street view video, i.e., panoramic street video, throughout the event.
[0029] Researchers collected PM concentration (μg / m³) at 5-second intervals. 3 The data collected included temperature (Fahrenheit), relative humidity (%), and GPS location data (longitude and latitude). Data collection was conducted during peak commuting hours—specifically 7:30 AM to 9:30 AM and 5:30 PM to 7:30 PM—to capture pollution levels during periods of high pedestrian and vehicular activity. After data cleaning, over 107,000 valid data points were retained. To ensure accuracy, all pollution data points underwent calibration. The calibration process began with pre-monitoring equipment calibration, comparing PM measurements from the AirBeam3 device with data from regulatory-grade instruments.
[0030] In addition to PM data, temperature and relative humidity were measured using an AirBeam3 device. This device incorporates a Texas Instruments HDC1080DMBR weather sensor. Audio and video variables were extracted from video footage captured by four panoramic cameras, yielding approximately 50 hours of video data related to air pollution measurements, i.e., street panoramic video.
[0031] S102. Extract features from the panoramic street video to obtain audiovisual features; among which, audiovisual features include: visual features and sound features.
[0032] In some embodiments of this application, based on the panoramic street video, video frames are extracted at a preset frame sampling rate to obtain multiple images; visual features are extracted from the multiple images to obtain visual features; audio segments are extracted based on the panoramic street video at a preset time segmentation to obtain multiple audio segments; and sound scene features are extracted from the multiple audio segments using a pre-determined sound scene classification model to obtain sound scene features.
[0033] S103. Based on sound scene features and visual features, respectively, select target sound scene features and target visual features; among them, target sound scene features represent sound scene features that directly affect the air pollution level; target visual features represent visual features that directly affect the air pollution level.
[0034] In some embodiments of this application, target acoustic features represent acoustic features that directly affect air pollution levels; target visual features represent visual features that directly affect air pollution levels.
[0035] For example, target soundscape features include vehicle traffic soundscape, pedestrian soundscape, and ambient soundscape; target visual features include road area, pedestrian count, building area, and vehicle area. Based on data from historical experiments, key variables contributing most to PM concentration prediction were selected by ranking feature importance. In terms of soundscape type, vehicle traffic soundscape and pedestrian soundscape are positively correlated with air pollution levels. These findings confirm that road dust from vehicle exhaust and pedestrian activity is a significant source of PM pollution from sidewalks. Conversely, ambient soundscape is negatively correlated with air pollution levels, indicating lower pollution levels in quieter, less urbanized areas. Visual variables continue to play a crucial role in model predictions. Larger road areas (road regions in the image) result in higher predicted PM2.5 and PM1 values, reflecting the contribution of road traffic activity to air pollution. Similarly, pedestrian count (the number of pedestrians on the sidewalk in the image) is positively correlated with sidewalk PM concentration levels. In PM2.5 predictions, built-up areas (i.e., the area occupied by buildings in the image) show a moderate positive impact, indicating that high-density urban environments or street canyon effects help capture or accumulate fine particulate matter. For PM1 predictions, vehicle areas (i.e., the area occupied by vehicles) become an important feature, but their relationship with PM1 concentration levels is more complex.
[0036] S104. Based on meteorological data, target soundscape features, and target visual features, a pre-determined air particulate matter concentration prediction model is used to predict the short-term particulate matter concentration of urban sidewalks, and the air particulate matter concentration prediction results are obtained. The air particulate matter concentration prediction model is trained based on historical street panoramic videos of sidewalks, historical meteorological data, and historical air particulate matter concentrations.
[0037] In some embodiments of this application, target visual features, meteorological data, and target soundscape features are matched to obtain a matched feature data set; wherein, the feature data set consists of target visual features, meteorological data, and target soundscape features within the same time period; based on the feature data set, a short-term particulate matter concentration prediction model is used to predict the particulate matter concentration of urban sidewalks to obtain the particulate matter concentration prediction result.
[0038] It is understandable that by extracting visual and acoustic features from panoramic street videos, and then filtering these features to obtain target acoustic and visual features, short-term particulate matter concentration prediction for urban sidewalks can be performed based on meteorological data, these target acoustic and visual features, and an air particulate matter concentration prediction model. Since acoustic features directly influence air pollution levels, and target visual features directly influence visual levels, the predicted air particulate matter concentration is relatively accurate. Furthermore, because panoramic street videos are easy to collect, they can improve the coverage of air pollution monitoring data and accurately predict air particulate matter concentrations on urban sidewalks, thus contributing to air quality management and public health protection efforts.
[0039] In some embodiments of this application, before S102, the following is also included: The research involves acquiring a sound dataset and defining the research objective. The sound dataset includes multiple audio samples corresponding to various sounds. The research objective is to determine the air pollution situation in a predefined area. Based on the research objectives, the sound dataset was classified to obtain multiple sound scene types; these multiple sound scene types include ambient sound, animal sound, car horn sound, car engine sound, vehicle traffic sound, construction noise, and pedestrian sound. Based on multiple sound scene types and sound datasets, an audio spectrogram converter model is trained to determine the sound scene classification model.
[0040] Exemplary examples exist, but dedicated tools for analyzing urban traffic and pedestrian soundscapes are still lacking. To comprehensively represent the diverse characteristics of urban pedestrian soundscapes, a deep learning-based soundscape prediction model was developed, drawing inspiration from the research of Verma et al. The framework is as follows: Figure 2 As shown. First, audio clips representing urban sidewalk soundscapes were selected and downloaded from the well-known audio dataset AudioSet. This database includes various human and animal sounds, as well as common environmental sound effects. The definitions of these sound categories can be queried using the AudioSet ontology. This application selected 15 sound categories, containing 2,103 audio samples, and reclassified them into 7 soundscape types relevant to the research objectives: 1. Ambient sounds: Sounds from natural, non-biological elements, such as the sound of wind. These sounds are usually associated with lower levels of air pollution.
[0041] 2. Animals: Sounds of urban wildlife, such as birds and insects, often associated with vegetation and the natural environment.
[0042] 3. Car horn: Reflects the sound of vehicles honking, indicating traffic congestion.
[0043] 4. Car engine sound: refers to the sound made by a vehicle engine, representing urban traffic.
[0044] 5. Vehicle traffic: The noise generated by moving vehicles constitutes the traffic soundscape.
[0045] 6. Construction noise: Urban construction noise may be related to pollutant emissions generated by construction activities.
[0046] 7. Pedestrians: Noise from human activity on sidewalks can increase exposure to dust and pollution.
[0047] Subsequently, a pre-trained deep learning model was employed and fine-tuned for an urban sidewalk soundscape classification task. Ultimately, the model was used to process audio clips collected from the study area, predicting the probability of the existence of seven predefined soundscape types.
[0048] The fine-tuned model used in this application is the Audio Spectrograph Transformer (AST) model. The AST model achieves state-of-the-art results in audio classification by converting audio into spectrograph images and applying a visual transformer to process the audio data. In model training, 70% of the data is used for training, and the remaining 30% is used as the validation set (the test set accounts for 0.3%). The learning rate is set to 5.00E-05. To evaluate model performance, multiple evaluation metrics are used, including classification cross-entropy loss, accuracy, recall, precision, and F1 score. The formulas for calculating these metrics are shown in formulas (1)-(5): The classification cross-entropy loss function is used to measure the true class probability y. j With the probability of predicting soundscape The differences between them. Where n represents the total number of categories. TP j FP represents the number of samples correctly classified as soundscape category j, and T is the total number of samples. j FN represents the number of samples incorrectly predicted as sound scene category j.j It represents the number of samples belonging to soundscape category j but incorrectly classified into other categories, while w j The sound scene category weights are determined based on the number of real samples in each category. Generally, the lower the loss value, the higher the accuracy, precision, recall, and F1 score, indicating better model performance. As training progresses, the validation loss continuously decreases and eventually stabilizes at a low level. Meanwhile, the accuracy and other evaluation metrics curves show an inverse trend compared to the loss curve. Initially, these metrics improve significantly, indicating that the model is effectively learning to make more accurate predictions. As training continues, the curves gradually stabilize, reflecting that the model has entered the convergence phase. The final model's accuracy and other evaluation metrics approach 90%, demonstrating excellent performance in sound scene prediction. Therefore, the best-performing model from the training period was selected to predict seven sound scene types within the study area and calculate their occurrence probabilities.
[0049] In some embodiments of this application, S102 can be implemented by S201-S202, as follows: S201. Based on the panoramic street video, extract video frames at a preset frame sampling rate to obtain multiple images; and extract visual features from the multiple images to obtain visual features.
[0050] In some embodiments of this application, the visual features include semantic segmentation region features and target detection object features. Semantic segmentation region features are multiple segmentation regions; target detection object features are multiple objects in each image.
[0051] In some embodiments of this application, image segmentation and target detection are performed simultaneously on multiple images to obtain multiple segmentation regions corresponding to each of the multiple images and multiple objects corresponding to each of the multiple images; visual features are determined based on the multiple segmentation regions and multiple objects.
[0052] For example, during mobile air pollution monitoring operations, environmental data was collected using video recording technology. To reduce image redundancy and computational burden, video frames were extracted at a fixed rate of one frame every 5 seconds (i.e., a preset frame sampling rate), with each frame having a resolution of 1280×640 pixels. Subsequently, panoramic segmentation technology was used to detect traffic-related objects and classify image pixels. The panoramic segmentation model used in this application is a Panoptic Feature Pyramid Network implemented based on Detectron2, an AI vision library developed by Facebook. This model achieved an average bounding box accuracy of 42.4% and an average mask accuracy of 38.5% on the COCO dataset. Pre-trained models were used to perform inference analysis on self-collected video frames, ultimately determining the number of traffic-related objects in the images and the proportion of pixel area occupied by specific urban environmental visual elements. The object types detected in the study (i.e., multiple objects) included pedestrians, cars, bicycles, buses, and trucks. Segmented pixel region variables (i.e., multiple segmented regions) included sky, buildings, trees, grass, roads, sidewalks, cars, buses, and trucks.
[0053] S202. Based on the panoramic street video, audio segments are extracted in preset time segments to obtain multiple audio segments; and sound scene features are extracted from the multiple audio segments using a pre-determined sound scene classification model to obtain sound scene features.
[0054] In some embodiments of this application, audio segments are extracted from street panoramic video in preset time segments to obtain multiple audio segments; multiple audio segments are converted into multiple spectrogram images through a sound scene classification model; based on the multiple spectrogram images, sound scene classification is performed through the sound scene classification model to obtain multiple sound classification values for each audio segment; and sound scene features are determined based on the multiple sound classification values for each audio segment.
[0055] For example, multiple audio segments were extracted by dividing the self-captured video into 10-second segments (i.e., preset time intervals), and these segments were saved in WAV format with a 44.1 kHz sampling rate. To analyze the audio information, a Python toolkit called scikit-maad was used, which allows for quantitative analysis of the audio recordings.
[0056] While existing tools can calculate general soundscape indices, there are currently no specific indices or tools for analyzing urban traffic or sidewalk soundscapes. To accurately capture the unique soundscape characteristics of the sampling routes, a soundscape classification model was used to categorize the soundscape types on urban sidewalks, resulting in multiple sounds: ambient sounds, animal sounds, car horns, car engine sounds, vehicle traffic sounds, construction noise, and pedestrian sounds. These multiple sounds were then identified as soundscape features.
[0057] In some embodiments of this application, S104 can be implemented by S301-S302, as follows: S301. Match the target visual features, meteorological data, and target soundscape features to obtain a matched feature data set; wherein, the feature data set consists of the target visual features, meteorological data, and target soundscape features within the same time period; S302. Based on the feature data set, the short-term particulate matter concentration of urban sidewalks is predicted using an air particulate matter concentration prediction model to obtain the air particulate matter concentration prediction results.
[0058] For example, the audiovisual data includes visual, audio, and meteorological data, namely, panoramic street video and meteorological data, the specific content of which is shown in Table 1. By calculating the average value over the same 10-second interval, the images (i.e., target visual features) and meteorological data are matched with each audio segment (i.e. target soundscape features) to obtain feature data sets.
[0059] Table 1 Table 1 shows the sound pressure level features, acoustic index features, and temporal acoustic features of multiple audio segments extracted using Python scripts.
[0060] In some embodiments of this application, after S103, the following is also included: Obtain geospatial data corresponding to the panoramic street video; wherein, the geospatial data includes static feature data characterizing the urban built environment; The target's visual features, meteorological data, and target's acoustic features are matched to obtain a set of matched feature data. Based on feature data sets and geospatial data, a pre-determined mixed prediction model for air particulate matter concentration is used to predict short-term particulate matter concentration on urban sidewalks, resulting in air particulate matter concentration prediction results. The mixed prediction model for air particulate matter concentration is a mixed prediction model that integrates audiovisual variables and geospatial variables.
[0061] For example, the geospatial data is shown in Table 2. The geospatial data includes building forms (i.e., buildings), points of interest, and green spaces (i.e., land cover). For the geospatial data, the time dimension is correlated with built environment and land use data through latitude and longitude information. To further analyze the impact of geospatial variables at a finer scale, this application employs a buffer feature enhancement method, setting five different radius levels: 100 meters, 200 meters, 300 meters, 400 meters, and 500 meters.
[0062] Table 2 This application develops a hybrid model that integrates audiovisual and geospatial variables. Audiovisual variables primarily reflect the dynamic characteristics of the city, while the inclusion of geospatial variables allows the model to more accurately capture static features affecting air pollution levels. Most high-impact geospatial variables are positively correlated with air pollution levels. For example, points of interest (such as traffic lights) are typically located in areas with high pedestrian and vehicle traffic. Studies have shown that frequent vehicle starts and stops near traffic lights significantly increase exhaust emissions. Previous studies have also confirmed high pollution exposure risks around the commercial and financial districts of CBD areas. Furthermore, the hybrid model integrates 34 audiovisual variables and 45 geospatial variables. Notably, among the identified high-impact variables, the number of geospatial variables is far less than that of audiovisual variables. This finding highlights the unique role of audiovisual variables in predicting PM concentrations—they can capture dynamic and localized environmental characteristics that geospatial variables cannot fully represent.
[0063] Understandably, this involves a hybrid model that integrates audiovisual and geospatial variables. Audiovisual variables primarily reflect the dynamic characteristics of the city, while the inclusion of geospatial variables allows the model to more accurately capture the static characteristics affecting air pollution levels, thereby improving the accuracy of particulate matter concentration prediction.
[0064] Based on the machine vision and hearing-based method for predicting air particulate matter concentration described in the above embodiments, this application also provides a machine vision and hearing-based system for predicting air particulate matter concentration, such as... Figure 3 As shown, Figure 3 This is a schematic diagram of the structure of an air particulate matter concentration prediction system based on machine vision and hearing, provided in an embodiment of this application. The air particulate matter concentration prediction system 3 based on machine vision and hearing includes: an acquisition module 301, an extraction module 302, and a prediction module 303, wherein... The acquisition module 301 is used to acquire panoramic street video through a panoramic camera and to collect meteorological data through a sensor. The extraction module 302 is used to extract features from the street panoramic video to obtain audiovisual features; wherein, the audiovisual features include: visual features and sound features; based on the sound features and the visual features, target sound features and target visual features are selected respectively; wherein, the target sound features represent sound features that directly affect the air pollution level; the target visual features represent visual features that directly affect the air pollution level; The prediction module 303 is used to predict the short-term particulate matter concentration of urban sidewalks based on the meteorological data, the target soundscape features, and the target visual features, using a pre-determined air particulate matter concentration prediction model, and to obtain the air particulate matter concentration prediction result; wherein, the air particulate matter concentration prediction model is trained based on historical street panoramic videos of sidewalks, historical meteorological data, and historical air particulate matter concentrations.
[0065] In some embodiments of this application, the extraction module 302 is further configured to extract video frames based on the street panoramic video at a preset frame sampling rate to obtain multiple images; and to extract visual features from the multiple images to obtain the visual features; to extract audio segments based on the street panoramic video at a preset time segmentation to obtain multiple audio segments; and to extract sound scene features from the multiple audio segments using a pre-determined sound scene classification model to obtain the sound scene features.
[0066] In some embodiments of this application, the acquisition module 301 is further configured to extract sound scene features from the plurality of audio segments using a pre-determined sound scene classification model, and before obtaining the sound scene features, acquire a sound dataset and determine a research objective; wherein, the sound dataset includes multiple audio samples corresponding to multiple different sounds; the research objective is the air pollution situation in a preset area; based on the research objective, the sound dataset is classified to obtain multiple sound scene types; wherein, the multiple sound scene types include ambient sound, animal sounds, car horn sounds, car engine sounds, vehicle traffic sounds, construction noise, and pedestrian sounds; based on the multiple sound scene types and the sound dataset, an audio spectrogram converter model is trained to determine the sound scene classification model.
[0067] In some embodiments of this application, the extraction module 302 is further configured to convert the plurality of audio segments into a plurality of spectrogram images using the soundscape classification model; classify the soundscape based on the plurality of spectrogram images using the soundscape classification model to obtain a plurality of sound classification values for each audio segment; and determine the soundscape features based on the plurality of sound classification values for each audio segment.
[0068] In some embodiments of this application, the extraction module 302 is further configured to simultaneously perform image segmentation and target detection on the multiple images to obtain multiple segmentation regions corresponding to each of the multiple images and multiple objects corresponding to each of the multiple images; and to determine the visual features based on the multiple segmentation regions and the multiple objects.
[0069] In some embodiments of this application, the prediction module 303 is further configured to match the target visual features, the meteorological data, and the target soundscape features to obtain a matched feature data set; wherein, the feature data set consists of the target visual features, meteorological data, and target soundscape features within the same time period; based on the feature data set, the short-term particulate matter concentration of urban sidewalks is predicted using the air particulate matter concentration prediction model to obtain the air particulate matter concentration prediction result.
[0070] In some embodiments of this application, the acquisition module 301 is further configured to perform filtering based on the soundscape features and the visual features respectively, select target soundscape features and target visual features, and then acquire geospatial data corresponding to the street panoramic video; wherein, the building form, points of interest and green space; and match the target visual features, the meteorological data and the target soundscape features to obtain a matched feature data group; The prediction module 303 is further configured to predict the short-term particulate matter concentration of urban sidewalks based on the feature data group and the geospatial data using a pre-determined mixed prediction model for air particulate matter concentration, and obtain the air particulate matter concentration prediction result; wherein, the mixed prediction model for air particulate matter concentration represents a mixed prediction model that integrates audiovisual variables and geospatial variables.
[0071] Based on the machine vision and hearing-based method for predicting air particulate matter concentration described in the above embodiments, this application also provides a machine vision and hearing-based device for predicting air particulate matter concentration, such as... Figure 4 As shown, Figure 4 This is a schematic diagram of a machine vision and hearing-based air particulate matter concentration prediction device provided in an embodiment of this application. The machine vision and hearing-based air particulate matter concentration prediction device 4 includes a processor 401 and a memory 402. The memory 402 is used to store computer programs; the processor 401 is used to call and run the computer programs from the memory to execute the machine vision and hearing-based air particulate matter concentration prediction method as described in the above embodiment.
[0072] In the embodiments of this application, the processor 401 described above can be at least one of the following: Application-Specific Integrated Circuit (ASIC), Digital Signal Processor (DSP), Digital Signal Processing Device (DSPD), Programmable Logic Device (PLD), Field-Programmable Gate Array (FPGA), Central Processing Unit (CPU), Controller, Microcontroller, and Microprocessor. It is understood that for different devices, the electronic device used to implement the above processor function can also be other types, and the embodiments of this application do not specifically limit it.
[0073] This application provides a computer-readable storage medium storing a computer program for implementing, when executed by a processor, the air particulate matter concentration prediction method based on machine vision and hearing as described in any of the above embodiments.
[0074] For example, the program instructions corresponding to the air particulate matter concentration prediction method based on machine vision and hearing in this embodiment can be stored on storage media such as optical discs, hard disks, and USB flash drives. When the program instructions corresponding to the air particulate matter concentration prediction method based on machine vision and hearing in the storage media are read or executed by an electronic device, the air particulate matter concentration prediction method based on machine vision and hearing as described in any of the above embodiments can be realized.
[0075] Furthermore, in the embodiments of this application, the functional modules can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional module.
[0076] If the integrated unit is implemented as a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the method of this embodiment. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0077] It should be understood that the phrases "one embodiment," "an embodiment," or "some embodiments" mentioned throughout the specification mean that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment," "in one embodiment," or "in some embodiments" appearing throughout the specification do not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above-described processes do not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the embodiments in this application are merely for descriptive purposes and do not represent the superiority or inferiority of the embodiments. The descriptions of the various embodiments above tend to emphasize the differences between the various embodiments; their similarities or commonalities can be referred to mutually, and for the sake of brevity, these will not be repeated here.
[0078] The modules described above as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules. They may be located in one place or distributed across multiple network units. Some or all of the modules may be selected to achieve the purpose of this embodiment according to actual needs.
[0079] In addition, each functional module in the various embodiments of this application can be integrated into one processing unit, or each module can be a separate unit, or two or more modules can be integrated into one unit; the integrated modules can be implemented in hardware or in the form of hardware plus software functional units.
[0080] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memory (ROM), magnetic disks, or optical disks.
[0081] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0082] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0083] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method or device embodiments.
[0084] The above description is merely an embodiment of this application, but the protection scope of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the protection scope of this application. Therefore, the protection scope of this application should be determined by the protection scope of the claims.
Claims
1. A method for predicting air particulate matter concentration based on machine vision and hearing, characterized in that, The method includes: The system uses panoramic cameras to capture panoramic street videos and sensors to collect meteorological data. Feature extraction is performed on the panoramic street video to obtain audiovisual features; wherein, the audiovisual features include: visual features and sound features; Based on the soundscape features and the visual features, target soundscape features and target visual features are selected; wherein, the target soundscape features represent soundscape features that directly affect the air pollution level; and the target visual features represent visual features that directly affect the air pollution level. Based on the meteorological data, the target acoustic features, and the target visual features, a pre-determined air particulate matter concentration prediction model is used to predict the short-term particulate matter concentration of urban sidewalks, and the air particulate matter concentration prediction result is obtained; wherein, the air particulate matter concentration prediction model is trained based on historical street panoramic videos of sidewalks, historical meteorological data, and historical air particulate matter concentrations.
2. The method according to claim 1, characterized in that, The panoramic street video is subjected to feature extraction to obtain audiovisual features, including: Based on the street panoramic video, video frames are extracted at a preset frame sampling rate to obtain multiple images; and visual features are extracted from the multiple images to obtain the visual features. Based on the panoramic street video, audio segments are extracted in preset time segments to obtain multiple audio segments; and sound scene features are extracted from the multiple audio segments using a pre-determined sound scene classification model to obtain the sound scene features.
3. The method according to claim 2, characterized in that, Before extracting soundscape features from the multiple audio segments using a pre-determined soundscape classification model to obtain the soundscape features, the method further includes: Acquire a sound dataset and determine the research objective; wherein, the sound dataset includes multiple audio samples corresponding to multiple different sounds; the research objective is the air pollution situation in a preset area; Based on the research objectives, the sound dataset is classified to obtain multiple sound scene types; wherein, the multiple sound scene types include ambient sound, animal sound, car horn sound, car engine sound, vehicle traffic sound, construction noise, and pedestrian sound. Based on the multiple soundscape types and the sound dataset, the audio spectrogram converter model is trained to determine the soundscape classification model.
4. The method according to claim 2, characterized in that, The process involves extracting soundscape features from the multiple audio segments using a pre-defined soundscape classification model to obtain the soundscape features, including: The soundscape classification model is used to convert the multiple audio segments into multiple spectrogram images. Based on the multiple spectrogram images, sound scene classification is performed using the sound scene classification model to obtain multiple sound classification values for each audio segment; The soundscape features are determined based on multiple sound classification values for each audio segment.
5. The method according to claim 2, characterized in that, The step of extracting visual features from the multiple images to obtain the visual features includes: Image segmentation and target detection are performed simultaneously on the multiple images to obtain multiple segmentation regions and multiple objects corresponding to each of the multiple images; The visual features are determined based on the multiple segmented regions and the multiple objects.
6. The method according to claim 1, characterized in that, The method of predicting short-term particulate matter concentration on urban sidewalks based on the meteorological data, the target acoustic features, and the target visual features, using a pre-determined air particulate matter concentration prediction model, and obtaining the air particulate matter concentration prediction result includes: The target visual features, the meteorological data, and the target soundscape features are matched to obtain a matched feature data set; wherein, the feature data set consists of target visual features, meteorological data, and target soundscape features within the same time period; Based on the aforementioned feature data set, the air particulate matter concentration prediction model is used to predict the short-term particulate matter concentration on urban sidewalks, resulting in the predicted air particulate matter concentration.
7. The method according to claim 1, characterized in that, After filtering based on the soundscape features and the visual features respectively, and selecting target soundscape features and target visual features, the method further includes: Obtain geospatial data corresponding to the panoramic street video; wherein, the geospatial data includes static feature data characterizing the urban built environment; The target visual features, the meteorological data, and the target soundscape features are matched to obtain a set of matched feature data. Based on the feature data set and the geospatial data, a short-term particulate matter concentration prediction for urban sidewalks is performed using a pre-determined mixed prediction model for air particulate matter concentration, resulting in the air particulate matter concentration prediction result; wherein, the mixed prediction model for air particulate matter concentration represents a mixed prediction model that integrates audiovisual variables and geospatial variables.
8. A system for predicting air particulate matter concentration based on machine vision and hearing, characterized in that, The air particulate matter concentration prediction system based on machine vision and hearing includes: an acquisition module, an extraction module, and a prediction module, wherein, The acquisition module is used to acquire panoramic street video through a panoramic camera and to collect meteorological data through a sensor. The extraction module is used to extract features from the street panoramic video to obtain audiovisual features; wherein, the audiovisual features include: visual features and sound features; based on the sound features and the visual features, target sound features and target visual features are selected respectively; wherein, the target sound features represent sound features that directly affect the air pollution level; the target visual features represent visual features that directly affect the air pollution level; The prediction module is used to predict the short-term particulate matter concentration of urban sidewalks based on the meteorological data, the target soundscape features, and the target visual features, using a pre-determined air particulate matter concentration prediction model, and to obtain the air particulate matter concentration prediction result; wherein, the air particulate matter concentration prediction model is trained based on historical street panoramic videos of the sidewalks, historical meteorological data, and historical air particulate matter concentrations.
9. A device for predicting air particulate matter concentration based on machine vision and hearing, characterized in that, include: Processor and memory, of which, The memory is used to store computer programs; The processor is configured to call and run the computer program from the memory to perform the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The device stores executable instructions for causing a processor to execute the method according to any one of claims 1 to 7.