Methods and apparatus for analyzing urban road humidity, electronic equipment, and storage media

By extracting frames, preprocessing, identifying and segmenting wet areas from urban road video stream data, and calculating wetness using scale weighting coefficients, the problem of insufficient quantitative measurement in existing wetness analysis is solved, achieving high-precision quantitative assessment and cost reduction.

CN122493356APending Publication Date: 2026-07-31RUNKAI ARTIFICIAL INTELLIGENCE TECHNOLOGY (HEBEI) CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
RUNKAI ARTIFICIAL INTELLIGENCE TECHNOLOGY (HEBEI) CO LTD
Filing Date
2026-04-24
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies lack quantitative measurement standards for urban road humidity analysis, which cannot meet the needs of multi-segment and cross-time period road condition trend analysis. The detection results have a high misjudgment rate and cannot support the decision-making needs of refined urban road management.

Method used

By acquiring road video stream data, performing frame extraction and preprocessing, wet area scale identification and segmentation, segmentation based on scale weight coefficients, and finally calculating and normalizing the wetness, the process is upgraded from qualitative judgment to quantitative assessment.

Benefits of technology

It significantly improves detection accuracy and scenario robustness, achieving an upgrade from qualitative judgment to quantitative evaluation, solving the problem of insufficient detection accuracy in existing technologies, and reducing system deployment and maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122493356A_ABST
    Figure CN122493356A_ABST
Patent Text Reader

Abstract

This application provides a method, apparatus, electronic device, and storage medium for analyzing urban road humidity, belonging to the field of intelligent analysis technology. The method includes: acquiring road video stream data; performing frame extraction processing on the road video stream data to obtain road image frame data; preprocessing the road image frame data to obtain target road image frame data; performing wet area scale recognition processing on the target road image frame data to obtain scale weight coefficients; performing segmentation processing based on the scale weight coefficients and the target road image frame data to obtain a road wet area segmentation mask; and performing humidity calculation and normalization processing based on the road wet area segmentation mask and the target road image frame data to obtain a road humidity value. The urban road humidity analysis method, apparatus, electronic device, and storage medium provided in this application can improve the accuracy of quantifying urban road humidity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of intelligent analysis technology, and more specifically, relates to a method and device for analyzing the humidity of urban roads, electronic equipment, and storage medium. Background Technology

[0002] With the deepening of smart city construction, real-time and accurate perception of road conditions has become a key requirement for intelligent traffic management. The wetness of urban roads directly affects vehicle braking distance and driving stability. Slippery road surfaces in rain, snow, or condensation weather are prone to traffic accidents. Therefore, efficient analysis of urban road wetness is an important method to improve traffic safety early warning and urban emergency response capabilities. Currently, the mainstream technology in the industry uses computer vision to analyze road wetness based on urban road monitoring videos. After processing frame segments of the monitoring video, wet areas on the road surface are identified using traditional convolutional models or manual feature extraction methods to determine the road wetness status. However, in practical applications, existing technologies lack quantitative measurement standards for detection results. Existing methods mostly only output qualitative judgments of "wet / dry," without a unified quantifiable indicator of wetness. This makes it difficult to support multi-segment, cross-time period road condition trend analysis and fails to meet the decision-making needs of refined urban road management. Summary of the Invention

[0003] The purpose of this application is to provide a method, apparatus, electronic device, and storage medium for analyzing urban road humidity, which can improve the accuracy of quantifying urban road humidity. To achieve the above objective, the technical solution provided by this application is as follows: Firstly, a method for analyzing the humidity of urban roads is provided, including: Acquire road video stream data, and perform frame extraction processing on the road video stream data to obtain road image frame data; The road image frame data is preprocessed to obtain the target road image frame data; The target road image frame data is subjected to wet area scale recognition processing to obtain scale weight coefficients; and based on the scale weight coefficients and the target road image frame data, segmentation processing is performed to obtain a road wet area segmentation mask; Based on the road wet area segmentation mask and the target road image frame data, the wetness is calculated and normalized to obtain the road wetness value.

[0004] Secondly, a device for analyzing the humidity of urban roads is provided, comprising: The image acquisition module is used to acquire road video stream data and perform frame extraction processing on the road video stream data to obtain road image frame data. The image preprocessing module is used to preprocess the road image frame data to obtain target road image frame data; The segmentation processing module is used to perform wet area scale recognition processing on the target road image frame data to obtain scale weight coefficients; and to perform segmentation processing based on the scale weight coefficients and the target road image frame data to obtain a road wet area segmentation mask. The humidity value calculation module is used to perform humidity calculation and normalization processing based on the road humidity area segmentation mask and the target road image frame data to obtain the road humidity value.

[0005] Thirdly, embodiments of this application also provide an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the urban road humidity analysis method provided by any possible implementation of the first aspect.

[0006] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the urban road humidity analysis method provided by any possible implementation of the first aspect.

[0007] The beneficial effects of the technical solution provided in this application are as follows: The urban road wetness analysis method, apparatus, electronic device, and storage medium provided in this application, compared with related technologies, directly utilize existing urban road video stream data for analysis, eliminating the need for additional dedicated sensing equipment and reducing system deployment and maintenance costs. This application's embodiment, through standardized preprocessing of road image frame data, eliminates data differences caused by different camera devices and lighting environments, significantly improving recognition stability in complex scenes and overcoming the problems of traditional methods being susceptible to interference and having high false positive and false negative rates. This application's embodiment, by introducing wet area scale recognition processing and combining it with scale weight coefficients for segmentation, can accurately adapt to multi-scale wetness targets such as small raindrops, localized slipperiness, and large areas of water accumulation on the road surface, achieving pixel-level accurate segmentation of wet areas and significantly improving detection accuracy and scene robustness. This application's embodiment, through wetness calculation and normalization processing, transforms the wetness state into a unified standard quantitative value, achieving an upgrade from qualitative judgment to quantitative evaluation, solving the problems of existing technologies lacking unified measurement standards and being unable to support cross-road segment and cross-time period comparative analysis, effectively addressing the problem of insufficient detection accuracy in existing road wetness state detection technologies. Attached Figure Description

[0008] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below.

[0009] Figure 1 A flowchart illustrating the urban road wetness analysis method provided in this application embodiment; Figure 2 This is a schematic diagram of the C3K2 module provided in an embodiment of this application; Figure 3 This is a schematic diagram of PANet feature fusion provided in an embodiment of this application; Figure 4 A diagram illustrating the overall architecture of urban road humidity analysis provided in this application embodiment; Figure 5 A data processing flowchart provided for embodiments of this application; Figure 6 This is a structural block diagram of the urban road humidity analysis device provided in the embodiments of this application; Figure 7 A schematic block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0010] The embodiments of this application are described below with reference to the accompanying drawings. It should be understood that the embodiments described below with reference to the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions of the embodiments of this application.

[0011] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the terms “comprising” and “including” as used in embodiments of this application mean that the corresponding feature can be implemented as the presented feature, information, data, step, operation, element, and / or component, but do not exclude implementation as other features, information, data, step, operation, element, component, and / or combinations thereof supported by the art. It should be understood that when we say that an element is “connected” or “coupled” to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. Furthermore, “connected” or “coupled” as used herein can include wireless connection or wireless coupling. The term “and / or” as used herein indicates at least one of the items defined by the term; for example, “A and / or B” can be implemented as “A,” or as “B,” or as “A and B.” When describing multiple (two or more) items, if the relationship between the multiple items is not explicitly defined, the multiple items can refer to one, several or all of the multiple items. For example, the description of "parameter A includes A1, A2, A3" can be implemented as parameter A includes A1 or A2 or A3, or it can be implemented as parameter A includes at least two of the three items A1, A2 and A3.

[0012] It is understood that in the embodiments of this application, data such as user information are involved. When the embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with relevant laws, regulations and standards.

[0013] To make the objectives, technical solutions, and advantages of this application clearer, the following description will be provided in conjunction with the accompanying drawings and specific embodiments.

[0014] This application provides a method for analyzing the humidity of urban roads, which can be executed by an electronic device, such as... Figure 1 As shown, the method may include: S101: Acquire road video stream data, perform frame extraction processing on the road video stream data, and obtain road image frame data.

[0015] In this embodiment, the road video stream data is continuous dynamic data of road images continuously collected by urban road monitoring equipment, such as real-time dynamic image data of road scenes transmitted by monitoring cameras at urban main roads and intersections; frame extraction processing is an image processing operation that extracts a single static image from continuous video stream data, such as the operation of capturing a single frame from continuous road video at fixed time intervals; road image frame data is static image data of road scenes obtained after frame extraction processing of road video stream data, such as a single static image of road monitoring extracted from the road video stream according to preset rules.

[0016] Considering that road video stream data is continuous and dynamic image data, it cannot be directly adapted to image analysis processing logic in the field of computer vision. It needs to be converted into static image data before subsequent road wetness analysis can be carried out. Furthermore, the video stream data generated by urban road monitoring is large in volume, and direct processing would increase computational resource consumption. Frame extraction processing can filter out effective static images, reducing the data volume. Frame extraction processing can completely preserve the visual features of the road scene, ensuring that the obtained road image frame data accurately reflects the actual road conditions.

[0017] In this embodiment, high-definition cameras deployed at various traffic nodes along urban roads continuously capture dynamic images of road scenes around the clock according to preset acquisition parameters such as resolution and frame rate, forming road video stream data. The cameras establish a stable connection with a backend computing terminal via a wired network, transmitting the acquired road video stream data to the backend computing terminal in real time. Next, a frame extraction processing program runs on the backend computing terminal to decode the received road video stream data, parsing out a continuous image sequence. A fixed frame extraction time interval is set according to the real-time requirements of road moisture analysis. Subsequently, the frame extraction processing program selects a single static image from the decoded image sequence at the set time intervals, completing the frame extraction process. During extraction, the original pixel and color visual features of the image are preserved, and no additional image modification operations are performed. Finally, the extracted single static images are structurally labeled according to the deployment location of the camera equipment and the image acquisition time, and the labeled static images are stored in the image database of the backend computing terminal, forming road image frame data that can be directly accessed.

[0018] This embodiment acquires road video stream data, leveraging existing urban road monitoring systems for data collection without requiring dedicated road data acquisition equipment, effectively reducing hardware costs. It also enables 24 / 7, contactless data collection of road scenes, ensuring continuous data acquisition. Frame extraction processing yields road image frame data, converting the dynamic video stream into static images suitable for subsequent image analysis, resolving the issue of dynamic video streams being unsuitable for direct road wetness analysis. Furthermore, frame extraction processing simplifies the amount of data to be processed, reducing the computational load on the backend computing terminal and improving the overall efficiency of the road wetness analysis process.

[0019] S102: Preprocess the road image frame data to obtain the target road image frame data.

[0020] In this embodiment, the road image frame data is preprocessed to obtain target road image frame data, including: The road image frame data is subjected to image scaling and resolution standardization to obtain road image frame data with uniform resolution. Color normalization and channel conversion are performed on road image frame data with uniform resolution to obtain road image frame data with uniform color distribution; Image enhancement processing is performed on road image frame data with uniform color distribution to obtain target road image frame data; image enhancement processing includes one or more combinations of random cropping, rotation, brightness and contrast adjustment, mirror transformation, Gaussian blur and noise injection.

[0021] In this embodiment, preprocessing is a set of image operations that standardize and optimize road image frame data, such as unifying the specifications and color calibration of road images of different sizes; image scaling and resolution standardization are operations that adjust the image size to unify the image resolution, such as adjusting road images of different resolutions to a fixed pixel specification; color normalization and channel conversion are operations that calibrate the image color distribution and convert the color channel format, such as converting the image color channels to a model-adaptive format and unifying the color mean; image enhancement is operations that perform diverse transformations on the image to improve feature representation, such as rotating the road image to enrich the image samples. Road image frame data with unified resolution is the image data after resolution standardization; road image frame data with unified color distribution is the image data after color normalization and channel conversion; and the target road image frame data is the final image data after all preprocessing is completed.

[0022] Considering that road image frame data comes from different camera devices with varying resolutions and color parameters, directly inputting it into the model would lead to biased analysis results. Therefore, standardization is necessary to unify the data. Subsequent semantic segmentation models have fixed requirements for the color channels and color distribution of the input images; color normalization and channel conversion are needed to adapt the images to the model's input standards. Real-world road scenes vary in lighting and angles, which can easily lead to unclear image features. Image enhancement processing is required to enrich the image samples and improve the model's adaptability to different scenes. Step-by-step preprocessing optimizes the image data layer by layer, ensuring the relevance of each step and ultimately obtaining target road image frame data that meets the analysis requirements.

[0023] For example, in this embodiment, stored road image frame data is retrieved, and an image scaling and resolution normalization processing program is started on the computing terminal. Based on the input requirements of the subsequent semantic segmentation model, fixed image resolution parameters are set, and road image frame data of different resolutions are scaled or interpolated proportionally to remove the influence of resolution inconsistencies, generating road image frame data with uniform resolution and temporarily storing it. Next, a color normalization and channel conversion processing program is called to perform color mean and variance calibration on the road image frame data with uniform resolution, unifying the image color distribution. Simultaneously, the image color channels are converted to a channel format suitable for the model, generating road image frame data with uniform color distribution. Subsequently, an image enhancement processing program is started, selecting multiple combinations from random cropping, rotation, and other processing methods to process the road image frame data with uniform color distribution. For example, the image is first randomly cropped and then rotated to retain the core features of the road area. Finally, the image data after all image enhancement processing is checked for integrity. After confirming that no features are lost, it is defined as target road image frame data and stored in a designated database to prepare data for subsequent semantic segmentation processing.

[0024] This embodiment achieves uniformity in image resolution and color distribution by performing step-by-step preprocessing on road image frame data, allowing image data from different sources to adapt to the input standards of subsequent models and eliminating analysis errors caused by device differences. Image enhancement processing enriches the feature representation of images through diverse image transformations, improving the data's coverage of different road scenes.

[0025] S103: Perform wet area scale recognition processing on the target road image frame data to obtain scale weight coefficients; and perform segmentation processing based on the scale weight coefficients and the target road image frame data to obtain the road wet area segmentation mask.

[0026] In this embodiment, the wet area scale recognition processing is an image processing operation that identifies and determines the size of wet areas in the target road image frame data, such as identifying the scale features of small raindrops or large areas of water accumulation on the road surface in the image; the scale weight coefficient is a numerical coefficient assigned based on the scale features of the wet area, for example, matching a high proportion of scale weight coefficients to large areas of water accumulation. The segmentation processing is an operation that divides the target road image frame data into pixel-level regions, such as distinguishing between wet and non-wet areas in the image; the road wet area segmentation mask is image data that marks the pixel positions of wet areas after segmentation processing, for example, image data that marks road water accumulation areas with specific pixel identifiers.

[0027] Considering that there are multi-scale wet areas in the target road image frame data, a single processing method cannot accurately capture the features of wet areas at different scales. It is necessary to first complete scale identification and then match the corresponding processing strategy. The scale weight coefficient can quantify the processing priority of wet areas at different scales. Based on it, the segmentation process can be more targeted.

[0028] In this embodiment, the target road image frame data is input to the scale recognition processing module. This module performs multi-scale feature extraction on the image, identifies the pixel range and size characteristics of each wet region in the image, and distinguishes between small, medium, and large-scale wet regions. Next, based on the identified wet region scale features, corresponding scale weight coefficients are generated according to a preset allocation rule. Different scale wet regions are matched with different weight coefficient values, and the coefficient allocation fits the segmentation requirements of each scale region. Subsequently, the target road image frame data and the scale weight coefficients are simultaneously input to the segmentation processing module. The module adjusts the segmentation parameters of each region according to the scale weight coefficients, strengthening pixel-level feature extraction for wet regions with high weight coefficients. Finally, the segmentation processing module completes pixel-level region division of the target road image frame data, accurately marks the pixel positions of all wet regions, generates a corresponding road wet region segmentation mask, and stores this mask in a designated data storage area to provide data support for subsequent wetness calculations.

[0029] The existing C3K2 module only uses fixed multi-scale convolutional kernels (3×3 / 5×5 / 7×7), and cannot dynamically adjust the receptive field according to the actual scale of the wet areas in the image (such as small raindrops, large areas of water accumulation, or localized slipperiness). That is, the C3K2 module uses 3×3, 5×5, and 7×7 convolutional kernels to extract features simultaneously, but the weights of the three convolutions are fixed and averaged, and do not automatically adjust according to whether the image contains "small raindrops" or "large areas of water accumulation." In contrast, the Dynamic Receptive Field Adaptive Module (DRFAM) first determines the size of the wet area in the image, then automatically calculates the α, β, and γ weights, and finally allows C3K2 to dynamically change the proportions of the three convolutional kernels, instead of using a fixed average fusion.

[0030] Figure 2 This is a schematic diagram of the C3K2 module provided in an embodiment of this application. The C3K2 module uses CBS units as input and output. First, it performs channel dimensionality reduction on the input features using a CBS unit with a stride of 1 and a 1×1 kernel. Then, it performs channel separation using a Split unit, dividing the feature channels into two independent subsets. One subset is directly transmitted via a direct connection path, while the other subset is sequentially input into two C3K Bottleneck units. Multi-scale feature extraction and reuse are achieved through residual backpropagation. The two feature paths are finally concatenated in the channel dimension using a Concat unit, and then output after feature calibration by a 1×1 CBS unit. This module significantly reduces computation through a channel separation strategy, while simultaneously enhancing the ability to extract features from multi-scale wet regions by relying on a multi-scale residual structure, providing high-quality feature input for subsequent dynamic weighting and feature fusion.

[0031] For example, the DRFAM module determines the target scale in real time through a scale-aware branch and dynamically allocates the proportion of convolutional kernel combinations. A new scale evaluation sub-network is added: it takes the preprocessed image as input and outputs three scale weight coefficients (α, β, γ, α+β+γ=1), corresponding to the weight proportions of small (3×3), medium (5×5), and large (7×7) convolutional kernels, respectively. Dynamic convolutional kernel combination: the multi-path convolutional outputs of the C3K2 module are no longer equally fused, but rather weighted and summed according to the weight coefficients (α, β, γ). For example, if a large area of ​​water is detected (large-scale target), γ approaches 0.7, and α and β are 0.15, prioritizing the preservation of global features from large convolutional kernels; if small raindrops are detected (small-scale target), α approaches 0.8, enhancing the detail capture of small convolutional kernels.

[0032] In this embodiment, the specific processing procedure of the DRFAM module is as follows: (1) Input: Input the preprocessed target road image frame data into the DRFAM module.

[0033] (2) Scale Evaluation Sub-network Processing: The scale evaluation sub-network within the DRFAM module performs global feature analysis on the target road image frame data to identify the actual size type of wet areas in the image, including small-scale raindrops, medium-scale localized slipperiness, and large-scale road surface water accumulation. Based on the identification results, three scale weight coefficients α, β, and γ are output, satisfying: α + β + γ = 1, where: α corresponds to the weight of a 3×3 small-scale convolution kernel; β corresponds to the weight of a 5×5 medium-scale convolution kernel; and γ corresponds to the weight of a 7×7 large-scale convolution kernel.

[0034] (3) Input the weight coefficients into the C3K2 module: The C3K2 module performs channel separation on the target road image frame data to form multiple feature channels, which are then fed into different convolution kernel paths of 3×3, 5×5, and 7×7 to obtain multiple convolution output features.

[0035] (4) Dynamic weighted fusion: The C3K2 module no longer performs fixed average fusion of multi-path convolution output features, but instead performs weighted summation based on α, β, and γ output by the DRFAM module: Dynamic multi-scale road image feature data = α × small convolution feature + β × medium convolution feature + γ × large convolution feature.

[0036] (5) Output dynamically adjusted features: After weighted fusion, dynamic multi-scale road image feature data is obtained. This feature can adaptively highlight the feature information that best matches the current wet area scale, thereby improving the segmentation accuracy.

[0037] In this embodiment, the wet area scale recognition processing can accurately determine wet areas of different scales in the image, providing targeted parameter basis for segmentation processing and avoiding the recognition bias of a single segmentation strategy for multi-scale targets. Segmentation processing based on scale weight coefficients ensures that the segmentation process conforms to the characteristics of wet areas at different scales, effectively improving the accuracy and completeness of road wet area segmentation. The generated road wet area segmentation mask can accurately mark the pixel positions of wet areas, providing accurate and reliable basic data for subsequent quantitative calculation of road wetness, and ensuring the effectiveness of the overall road wetness analysis results.

[0038] In this embodiment, segmentation processing is performed based on scale weight coefficients and target road image frame data to obtain a road wet area segmentation mask, including: Multi-path convolution feature extraction processing is performed on the target road image frame data to obtain multi-path convolution output feature data; Based on the scale weight coefficient, the multi-path convolution output feature data is weighted and summed to obtain dynamic multi-scale road image feature data. Multi-scale feature stitching and fusion processing is performed on dynamic multi-scale road image feature data to obtain fused road image feature data. Semantic segmentation processing is performed on the fused road image feature data to obtain the road wet area segmentation mask.

[0039] In this embodiment, multi-path convolution feature extraction involves performing convolution operations on the image using multiple parallel paths to extract features, such as using convolution kernels of different sizes to extract image features separately. The multi-path convolution output feature data is the multi-path feature information output after multi-path convolution processing. Weighted summation processing is the operation of merging and calculating the multi-path features according to their weight values. Dynamic multi-scale road image feature data is the multi-scale fused feature information obtained after weight adjustment. Multi-scale feature stitching and fusion processing is the operation of integrating multi-level features. The fused road image feature data is the comprehensive feature information after fusion. Semantic segmentation processing is the operation of dividing image regions according to pixel categories. The road wet region segmentation mask is image data marking the locations of wet pixels.

[0040] Considering that a single convolutional path is insufficient to fully capture the features of multi-scale wet regions, multi-path convolution can improve feature richness; different scale wet regions need to be matched with different feature weights, and weighted summation can strengthen effective features and suppress invalid information; multi-level features need to be fully integrated to improve expressive power, and multi-scale splicing fusion can integrate global and local features; fused features are more conducive to category determination, and semantic segmentation can achieve accurate pixel-level division.

[0041] In this embodiment, channel separation is first performed on the target road image frame data, dividing the feature channels into multiple subsets. Convolution calculations are then performed on different paths using convolution kernels of different sizes to complete multi-path convolution feature extraction, outputting multi-path convolution output feature data. Next, weights are assigned to the multiple output features according to scale weight coefficients, and weighted summation is performed to generate dynamic multi-scale road image feature data. Subsequently, bidirectional feature transfer from top to bottom and bottom to top is performed on the dynamic multi-scale road image feature data, and tensor concatenation of features at different levels is performed to complete multi-scale feature concatenation and fusion processing, obtaining fused road image feature data. Finally, pixel-level category determination is performed on the fused road image feature data to distinguish between wet and non-wet regions, completing semantic segmentation processing, generating a road wet region segmentation mask, and storing the data.

[0042] In this embodiment, multi-path convolutional feature extraction can comprehensively capture multi-scale wet region features, and weighted summation makes the features more closely match the actual wet target scale. Multi-scale feature concatenation and fusion enhances feature representation capabilities and improves feature stability in complex scenes. Semantic segmentation relies on high-quality fused features to achieve accurate pixel-level division, effectively reducing false positives and false negatives.

[0043] For example, YOLOv11-seg improves upon the YOLOv8 architecture by adding the following key modules: Backbone network: It adopts an improved C3K2 structure to achieve deeper feature extraction.

[0044] Traditional C3 modules primarily rely on multiple 3×3 standard convolutions and partial channel residual connections to extract local features. However, in practical applications, using only fixed-size convolutional kernels results in a limited receptive field, especially in scenarios with complex backgrounds, blurred boundaries, and multi-scale targets, easily leading to insufficient feature extraction or information loss. To address this issue, the C3K2 module introduces a combination of multi-scale convolutional kernels, with common configurations including 3×3, 5×5, and 7×7 variable convolutional kernels. This embodiment allows the model to simultaneously capture image structural information from different receptive fields within the same layer, thereby more comprehensively perceiving target contours and contextual background.

[0045] Furthermore, C3K2 incorporates a channel separation strategy, dividing the input feature channels into multiple independent channel subsets, each assigned to a different feature extraction sub-path. Different scale convolutional kernels are used for feature extraction in each sub-path: a 3×3 kernel for the first sub-path, a 5×5 kernel for the second, and a 7×7 kernel for the third. After feature extraction is completed, the multiple output features are fused. This embodiment effectively reduces the model's computational load by combining channel separation with multi-scale convolution, while simultaneously improving the model's ability to extract parallel features from multi-scale wet regions.

[0046] The Neck architecture introduces a Path Aggregation Network (PANet) to enhance multi-scale feature fusion capabilities. PANet receives dynamic multi-scale road image feature data, performs bidirectional feature fusion on it through top-down high-level semantic transfer and bottom-up low-level detail feedback, and finally outputs fused road image feature data for use in subsequent semantic segmentation branches.

[0047] PANet is a network architecture used to enhance the information flow of feature pyramid structures, commonly employed in object detection tasks. Traditional feature pyramids only transmit high-level semantic information from top to bottom, while PANet introduces a bottom-up path enhancement module, enabling lower-level features to be effectively fed back to higher levels, improving the fusion effect of multi-scale features. Furthermore, PANet uses tensor concatenation instead of addition in feature fusion, preserving feature information from each layer more completely. It also introduces a global information path to improve the perception of small targets and complex backgrounds. The structure is further optimized, eliminating redundant modules and enhancing feature stacking effects, providing the detection head with richer semantic and positional information.

[0048] Figure 3 This is a schematic diagram of PANet feature fusion provided in an embodiment of this application. Tensor concatenation is a feature fusion operation that concatenates multiple feature tensors along the feature channel dimension. The input consists of multiple sets of independent feature tensors with consistent spatial dimensions. Each set of feature tensors corresponds to a different feature source, such as features extracted by convolutions at different scales or features output by networks at different levels. Through tensor concatenation, the multiple features are merged along the channel dimension, and the fused feature tensor is output. This tensor fully retains the original information of all input features and has the expressive power of multiple features, providing high-quality fused feature support for subsequent semantic segmentation processing.

[0049] Head Output: Building upon bounding box detection, a semantic segmentation branch is added to generate a segmentation mask for each frame of the image, marking wet regions. This model balances detection speed and segmentation accuracy, enabling real-time processing of image streams from cameras. The original YOLOv11-segHead output includes bounding boxes, class confidence scores, and a segmentation mask, applicable to general targets such as people, vehicles, or objects. This embodiment removes the bounding boxes, outputting only pixel-level segmentation masks for wet road regions, specifically for wet region labeling. The Head structure enhances and separates the semantic segmentation branch from the original bounding box detection branch, removing redundant bounding box detection and multi-class classification outputs, retaining only pixel-level segmentation for wet road regions. For each frame of the target road image data, a wet region segmentation mask is directly output, accurately distinguishing wet and non-wet regions in the image with pixel-level annotation, providing an accurate data foundation for subsequent wetness calculations.

[0050] In this embodiment, multi-path convolution feature extraction processing is performed on the target road image frame data to obtain multi-path convolution output feature data, including: Channel separation processing is performed on the target road image frame data to obtain multiple sets of road image channel feature data; Convolutional kernel feature extraction was performed on multiple sets of road image channel feature data to obtain multiple sets of convolutional feature data; Multiple sets of convolutional feature data are concatenated to obtain multi-path convolutional output feature data.

[0051] In this embodiment, channel separation processing is a process of dividing the image feature channels into multiple subsets, for example, uniformly dividing the feature channels of the input image into three independent data sets. The multiple sets of road image channel feature data are the multi-channel feature information obtained after channel separation. Convolution kernel feature extraction processing is an operation of performing local feature extraction on the image channel features using convolution kernels, for example, using convolution kernels of different sizes to extract texture and contour information respectively. The multiple sets of convolution feature data are the feature results output after multi-path convolution processing. Feature concatenation processing is an operation of merging multiple independent convolution feature sets. The multi-path convolution output feature data is the comprehensive feature data obtained after concatenation.

[0052] Considering that the target road image frame data contains rich channel information, direct convolution can easily lead to feature redundancy and computational waste. Channel separation can be used to achieve feature splitting. Different channels correspond to different image attributes, and performing convolution kernel feature extraction separately can improve the targeting and completeness of feature extraction. Multiple independent features need to be integrated into a unified expression form. Feature concatenation can retain all effective information. Wet region segmentation needs to take into account both detailed and global information. This step-by-step processing can reduce the amount of computation while enhancing the multi-scale feature capture capability and improving the stability and accuracy of subsequent segmentation.

[0053] For example, this embodiment first reads the target road image frame data, and performs channel separation processing on the data according to a preset channel division rule, splitting the complete feature channel into multiple independent sets of data, resulting in multiple sets of road image channel feature data. Then, each set of road image channel feature data is fed into a corresponding convolutional processing unit, and convolutional kernel feature extraction processing is performed independently using convolutional kernels of different sizes to extract local texture edges and global structural information, resulting in multiple sets of convolutional feature data. Next, the multiple sets of convolutional feature data are sequentially merged according to channel dimensions, and feature stitching processing is performed to maintain the integrity of feature position and dimensional information. Finally, the stitched data is regularized and optimized to output multi-path convolutional output feature data.

[0054] In this embodiment, channel separation effectively reduces feature redundancy and computational resource consumption. Grouped convolutional kernel feature extraction improves the comprehensiveness and detail of feature capture. Feature concatenation fully preserves multi-path feature information, enhancing feature representation capabilities.

[0055] In this embodiment, multi-scale feature stitching and fusion processing is performed on dynamic multi-scale road image feature data to obtain fused road image feature data, including: Top-down high-level semantic feature transfer processing is performed on dynamic multi-scale road image feature data to obtain high-level road image feature data. Bottom-up low-level detail feature feedback processing is performed on dynamic multi-scale road image feature data to obtain low-level road image feature data; Tensor stitching is performed on the feature data of high-level road images and low-level road images to obtain fused road image feature data.

[0056] In this embodiment, multi-scale feature stitching and fusion processing integrates image features from different levels, such as fusing global contour and local texture features. Top-down high-level semantic feature propagation processing propagates global semantic information from deep networks downwards, such as propagating semantic information about the overall wet area of ​​the road surface. High-level road image feature data represents feature information characteristic of global semantics. Bottom-up low-level detail feature feedback processing propagates detail information from shallow networks upwards, such as propagating detail information about road surface edges and raindrop textures. Low-level road image feature data represents feature information characteristic of local details. Tensor stitching processing merges feature tensors by dimension. The fused road image feature data represents the comprehensive feature information after stitching.

[0057] Top-down high-level semantic feature transfer processing processes the high-level features output by deep networks, passing the global semantic information learned by the deep layers of the network down to the shallow features layer by layer. This enables the shallow features to have an overall understanding of the category of wet areas on the road surface, clearly identifying which areas in the image belong to wet areas and improving the semantic accuracy of the features.

[0058] Bottom-up low-level detail feature feedback processing processes the low-level features output by the shallow network, passing the local edges, textures, and positional details learned by the shallow layers of the network upwards to the high-level features, so that the high-level features have accurate position and boundary information, improving the edge accuracy of wet region segmentation.

[0059] Tensor stitching is a process that merges the high-level road image feature data and the low-level road image feature data, which have undergone the bidirectional transmission mentioned above, in the channel dimension, fully preserving semantic and detailed information to form a fused road image feature data with stronger expressive power.

[0060] Considering that dynamic multi-scale road image feature data contains multi-level information, single-direction transmission cannot fully utilize semantic and detailed features; high-level features have global semantics but lack details, while low-level features have details but lack global cognition, so bidirectional transmission is required to achieve complementarity; tensor stitching can completely preserve the feature information of both paths and avoid feature loss; wet area segmentation requires the simultaneous identification of large-scale water accumulation and small slippery areas, and bidirectional transmission and tensor stitching can improve feature integrity and enhance the stability and accuracy of segmentation in complex scenes.

[0061] In this embodiment, dynamic multi-scale road image feature data is first input into the feature fusion unit, initiating a top-down high-level semantic feature transfer process. This process propagates the global semantic information extracted by the deep network layer by layer downwards, combining it with lower-level features to obtain high-level road image feature data. Subsequently, a bottom-up low-level detail feature feedback process is performed on the same dynamic multi-scale road image feature data. This process propagates the edge texture and other detail information extracted by the shallow network layer by layer upwards, enhancing the detail expression of the upper-level features to obtain low-level road image feature data. Next, the generated high-level and low-level road image feature data are aligned along a specified dimension, and tensor concatenation is performed to merge the two feature streams in an orderly manner. Finally, the concatenated feature data is regularized and optimized to remove redundant information, resulting in fused road image feature data, providing high-quality feature input for subsequent semantic segmentation.

[0062] In this embodiment, bidirectional feature transfer, both top-down and bottom-up, achieves full complementarity between global semantics and local details. Tensor concatenation fully preserves multi-level feature information, enhancing feature representation capabilities. The fused road image feature data possesses both global cognition and detail perception capabilities, significantly improving the accuracy and robustness of wet region segmentation.

[0063] Specifically, this embodiment uses the improved YOLOv11-seg model, whose structural improvements are as follows: Backbone employs the C3K2 module, which introduces 3×3, 5×5, and 7×7 multi-scale convolutional kernels. Combined with a channel separation strategy, it extracts features from multiple paths and then fuses them, thereby improving the multi-scale feature capture capability and reducing computational cost.

[0064] Neck employs PANet: based on top-down propagation of high-level semantics, it adds a bottom-up low-level detail feedback path and uses tensor concatenation to complete feature fusion, thus fully preserving the feature information of each layer.

[0065] The Head branch adds a semantic segmentation branch: based on object detection, it outputs a pixel-level segmentation mask to accurately mark wet regions.

[0066] The scale identification and processing of humid areas is implemented through the DRFAM module: The built-in scale evaluation subnetwork takes target road image frame data as input and outputs three scale weight coefficients α, β, and γ, which satisfy α+β+γ=1. The coefficients correspond to the weights of small, medium, and large scale convolution kernels, respectively. Large areas of water accumulation increase the weight of large convolution kernels, while small raindrops increase the weight of small convolution kernels.

[0067] Semantic segmentation processing performs pixel-level classification on the fused road image feature data to obtain a segmentation mask for the wet area of ​​the road.

[0068] In this embodiment, domain-invariant feature extraction processing is performed on the fused road image feature data to obtain domain-invariant feature vector data; The domain-invariant feature vector data is processed by an adversarial loss function to obtain the model optimization loss value; during the training phase of the improved YOLOv11-seg model, the parameters of the improved YOLOv11-seg model are iteratively optimized based on the model optimization loss value. In this embodiment, a wet domain adversarial discriminator is constructed. Its core logic is as follows: using the adversarial learning idea, a "domain discriminator" is trained to distinguish whether the image comes from the "training domain" (known scene) or the "target domain" (unknown deployment scene). That is, the target road image frame data is divided into training domain road image frame data and target domain road image frame data. At the same time, the YOLOv11-seg model is trained to generate "domain invariant features" (that is, the features of the wet area remain consistent in different scenes).

[0069] Generator: Improved YOLOv11-seg model (including DRFAM and BiA-PANet), outputting wet region segmentation mask and domain-invariant feature vector; Discriminator: Lightweight CNN, input feature vector, output "domain label" (training domain / target domain); Adversarial loss function:

[0070] Where G is the generator and D is the discriminator. To counteract the loss function value, To train the domain road image frame data, For target domain road image frame data, The mathematical expectation of the training domain data distribution. The mathematical expectation of the data distribution in the target domain; For generator G pairs Extracted domain-invariant features; For generator G pairs Extracted domain-invariant features; For discriminator D pair The discrimination probability; For discriminator D pair The discrimination probability.

[0071] The generator drives the discriminator to correctly identify road image frame data in the training domain. The more stable the improved YOLOv11-seg model is in extracting domain-invariant features from road image frame data in the training domain, the easier it is for the discriminator to identify it as part of the training domain, and the lower the loss value of this term. This ensures the model's basic segmentation ability on the training data. The generator drives the discriminator to be unable to distinguish target domain road image frame data, mistaking domain-invariant features of the target domain for "training domain-invariant features." The lower this loss value is when the model successfully confuses the discriminator. This is key to achieving cross-scene generalization, allowing the model to output a feature distribution consistent with the training scene even in unfamiliar target scenes. At the minimum, the generator G has successfully learned a cross-scene general feature representation of wet regions, making the feature distributions of the training domain and the target domain tend to be consistent. At this point, the model's wet region segmentation accuracy in the target domain will be significantly improved, solving the problem of model performance degradation under different urban road scenes, lighting, and weather conditions.

[0072] The generator employs an improved YOLOv11-seg model, integrating a DRFAM module, a bidirectional feature fusion structure BiA-PANet, and a dedicated segmentation branch, enabling multi-scale wet region feature extraction and pixel-level segmentation. The discriminator uses a lightweight convolutional neural network structure, used only during model training. It receives the domain-invariant feature vectors output by the generator and outputs domain labels to distinguish between the training and target domains.

[0073] In this embodiment, an adversarial loss function is used during model training. By constraining the feature extraction process of the generator through the adversarial loss function, the domain-invariant features output by the generator cannot be distinguished from the source domain by the discriminator. This allows the generator to learn humid region features with domain invariance, thereby improving the model's generalization ability and detection stability under different cities, climates and acquisition devices.

[0074] During the model training phase, an adversarial learning mechanism is employed. Target road image frames are input into the generator to obtain a wet region segmentation mask and domain-invariant feature vectors. These domain-invariant feature vectors are then input into the discriminator, which outputs a domain label classification result. Based on the classification result, an adversarial loss is calculated, and this loss is used to simultaneously optimize both the generator and discriminator. This allows the generator to gradually learn domain-invariant wet features unaffected by the scene, improving cross-scene adaptability.

[0075] In the practical application stage of road wetness analysis, only the generator is retained for inference, while the discriminator is completely removed and does not participate in any computation. The generator receives preprocessed target road image frame data and sequentially performs dynamic scale recognition, multi-path convolutional feature extraction, multi-scale feature fusion, and semantic segmentation, directly outputting a wet region segmentation mask for subsequent wetness calculation and visualization. This application stage does not involve adversarial learning, does not use a discriminator, and does not calculate any loss function.

[0076] In this embodiment, the discriminator is specifically optimized for "humidity domain features" (rather than general domain features). The adversarial training process does not increase the computational load during the inference phase and does not affect real-time performance. This embodiment employs adversarial learning, setting the generator to an improved YOLOv11-seg model that outputs domain-invariant features; setting the discriminator to a lightweight CNN that distinguishes between the training domain and the target domain; and optimizing the model through an adversarial loss function to ensure that humidity features remain consistent across city and climate scenarios. Adversarial training is performed only during the training phase and does not increase computational load during the inference phase.

[0077] For example, the improved YOLOv11-seg model is trained on 500 hours of real urban road video, covering all weather conditions, different seasons, multiple climates, various terrains, and urban scenes in both northern and southern China. The training process integrates visual features from publicly available datasets and adopts a multi-scale anchoring mechanism to adapt to the image structure of different camera models and angles, ensuring the model's robustness and versatility in complex scenarios.

[0078] S104: Based on the road wet area segmentation mask and the target road image frame data, perform wetness calculation and normalization to obtain the road wetness value.

[0079] In this embodiment, the road wet area segmentation mask includes multiple segmentation masks corresponding to road wet areas; Based on the road wet region segmentation mask and target road image frame data, wetness is calculated and normalized to obtain road wetness values, including: Pixel counting is performed on each wet area of ​​the road to obtain the number of pixels in multiple wet areas. Pixel counting is performed on the target road image frame data to obtain the total number of pixels in the image, and the total number of pixels in the image is used as the original humidity calculation data. The pixel count of each humid region and the original humidity calculation data are processed by ratio calculation to obtain multiple original humidity ratios; The weighted average of multiple raw humidity ratios is processed to obtain weighted average humidity data; The weighted average humidity data is normalized to obtain the road humidity value.

[0080] In this embodiment, the humidity calculation and normalization process involves combining the mask and image data to calculate and standardize the humidity level. Pixel counting is the operation of counting the number of effective pixels in the image. The number of pixels in the wet region is the total number of pixels marked as wet regions in the mask. The total number of pixels in the image is the total number of pixels contained in a single frame. The original humidity calculation data is the total number of pixels in the image used for basic calculations. The ratio operation is the operation of dividing two sets of values ​​to obtain a relative proportion. The original humidity ratio is the basic ratio value of wet pixels to total pixels. The weighted average process is the process of averaging the values ​​according to weights. The weighted average humidity data is the intermediate value after weighted calculation. The normalization process is the process of mapping the data to a unified interval. The road humidity value is the final output quantitative index.

[0081] Considering that the degree of humidity needs to be objectively quantified, pixel counting and ratio calculation can intuitively reflect the proportion of wet areas; different regions of the image have different levels of importance, and weighted averaging can highlight key observation areas; the numerical ranges of different images and scenes are inconsistent, and normalization can make data comparable across road sections and time periods.

[0082] In this embodiment, firstly, the road wet area segmentation mask and the target road image frame data are read. Pixel counting is performed on the mask, traversing all pixels and counting the number of pixels marked as wet areas to obtain the number of wet area pixels. Next, pixel counting is performed on the target road image frame data, counting the product of the horizontal and vertical pixels to obtain the total number of image pixels, which is used as the original wetness calculation data. Then, the wet area pixel count and the original wetness calculation data are compared to obtain the original wetness ratio. Next, a weighted average is performed on the original wetness ratio according to the preset regional weights, strengthening the weight of the core image region to obtain the weighted average wetness data. Finally, historical sliding window extreme value data is retrieved, and the weighted average wetness data is normalized to map the result to a standard interval to obtain the road wetness value and complete the storage.

[0083] Pixel counting is performed on the segmentation mask corresponding to each wet area of ​​the road to obtain the number of pixels in the wet area. ; Perform pixel counting on the target road image frame data to obtain the total number of pixels in the image. , as the original humidity calculation data; The ratio of the number of pixels in each humid region to the original humidity calculation data is calculated to obtain the original humidity ratio. Its value is between 0 and 1; A weighted average of multiple raw humidity ratios is performed, and a pixel weight function is introduced. By strengthening the weight of key areas such as the image center, weighted average humidity data is obtained. ; in, This is a weighted average humidity data; This is the pixel weighting function, which can be set to Gaussian weighting.

[0084] The weighted average humidity data is normalized, including: historical sliding window normalization: based on historical maximum values. and minimum value Perform interval mapping; standard score normalization: standardize based on the sliding window mean and standard deviation; finally output standardized road wetness values. .Will It is mapped to intuitive numerical or color levels in the user interface.

[0085] ; The model evaluation uses the Intersection over Union (IOU) metric, which calculates the IOU between the predicted mask (i.e., the road wet area segmentation mask) and the ground truth mask (i.e., the manually labeled wet areas) to measure segmentation accuracy. The ground truth mask is standard mask data obtained by manually annotating wet areas in the road image at the pixel level, representing the actual location and extent of wet areas in the image; the predicted mask is the road wet area segmentation mask output by the model in this embodiment, representing the location and extent of wet areas identified by the model.

[0086] In this embodiment, a threshold (e.g., 0.5) is set. When the IOU between the predicted mask and the true mask is ≥ 0.5, the prediction is considered correct. IOU is an important indicator for evaluating the accuracy of the algorithm, reflecting the degree of overlap between the prediction and the true value.

[0087] As can be seen from the above, this application embodiment directly utilizes existing video stream data of urban roads for analysis, eliminating the need for additional dedicated sensing equipment and reducing system deployment and maintenance costs. This application embodiment, through standardized preprocessing of road image frame data, can eliminate data differences caused by different camera devices and lighting environments, significantly improving recognition stability in complex scenes and overcoming the problems of traditional methods being susceptible to interference and having high false positive and false negative rates. This application embodiment, by introducing wet area scale recognition processing and combining it with scale weight coefficients for segmentation processing, can accurately adapt to multi-scale wet targets such as small raindrops on the road surface, localized slipperiness, and large areas of water accumulation, achieving pixel-level accurate segmentation of wet areas and significantly improving detection accuracy and scene robustness. This application embodiment, through wetness calculation and normalization processing, transforms the wetness state into a unified standard quantitative value, achieving an upgrade from qualitative judgment to quantitative evaluation, solving the problems of existing technologies lacking unified measurement standards and being unable to support cross-road segment and cross-time period comparative analysis, effectively addressing the problem of insufficient detection accuracy in existing road wetness state detection technologies.

[0088] In one embodiment of this application, the urban road wetness analysis method further includes: The road moisture content values ​​are subjected to grade threshold matching processing to obtain road moisture content grade data; Visualize the road wet area segmentation mask by rendering a layer to obtain wet area mask layer data. Visualization basemap adaptation processing is performed on the target road image frame data to obtain road visualization basemap data; Based on the humid area mask layer data, the road visualization base map data is overlaid to obtain masked overlaid road image data. Based on road humidity level data, state association rendering processing is performed on camera map point data to obtain road humidity map visualization data; the camera map point data is map coordinate information data that corresponds one-to-one with the actual deployment location of urban road camera equipment. The road image data with mask overlay, road humidity level data and road humidity map visualization data are integrated and processed by a visualization interface to obtain comprehensive road humidity visualization interface data. Multi-terminal adaptation rendering is performed on the comprehensive road moisture visualization interface data to obtain multi-terminal visualization display data; The system performs multi-terminal visualization processing to push and display road moisture values.

[0089] In this embodiment, the grade threshold matching process involves comparing the road wetness value with a preset grade range to determine the grade, such as determining whether the value is dry or slippery after comparison. The visualization layer rendering process involves converting a segmented mask into a displayable layer, such as marking wet areas with a colored layer. The visualization base map adaptation process involves adjusting the image to display the base map. The layer overlay process involves merging the mask layer with the base map. The state-related rendering process involves setting the display style for points based on their grade. The visualization interface integration process involves combining multiple display contents into a unified interface. The multi-terminal adaptation rendering process involves adjusting the interface to adapt to different terminal display specifications. The push display process involves sending data to the terminal and displaying it.

[0090] Considering that road moisture levels are not intuitive enough, matching with graded thresholds can simplify the understanding process for managers; single data points cannot quickly pinpoint problems, but layer overlay can visually represent the location of moist areas. Given that urban management requires a global perspective, map point plotting and status rendering can provide a comprehensive overview of road conditions; and considering its diverse application scenarios, multi-terminal adaptation can meet the viewing needs of different scenarios. In this embodiment, the road moisture content value is first read, and a level threshold matching process is performed according to a preset level interval to obtain road moisture content level data. Next, a visualization layer rendering process is performed on the road moisture area segmentation mask to generate moisture area mask layer data with clear markings. Visualization base map adaptation processing is performed on the target road image frame data to generate road visualization base map data that conforms to display specifications. The mask layer and the road base map are overlaid to obtain mask-overlaid road image data. Urban camera deployment information data is retrieved, and map point plotting processing is performed on the electronic map to obtain camera map point data. Based on the road moisture content level data, status association rendering processing is performed on the points to obtain road moisture content map visualization data. The mask overlaid image, level data, and map data are integrated into a visualization interface to obtain comprehensive road moisture content visualization interface data. Multi-terminal adaptation rendering processing is performed on the comprehensive interface to generate multi-terminal visualization display data adapted to different display devices. Finally, the display data is pushed to the corresponding terminal for display, realizing a complete visualization display of the road moisture content value.

[0091] This embodiment transforms abstract numerical values ​​into intuitive graphics and level indicators, lowering the barrier to understanding and use. Layer overlay and map plotting enable simultaneous display of local details and global situation, improving management efficiency. Multi-terminal adaptation and push display meet the real-time viewing needs of different scenarios, improving response speed.

[0092] The visualization also supports timeline scrolling, historical record playback, humidity heatmap, trend prediction, event ranking, and alarm prompts.

[0093] For example, the detection and calculation results of this embodiment are presented in real time through a front-end interface, including the following functions: displaying the current image frame and its wet area mask overlay; displaying the wetness value and grade; marking the location of each camera and its current wetness status on the map; and allowing users to slide through the timeline to select historical playback records. This graphical interface is deployed on the management platform terminal and supports interactive viewing of the wetness status of any intersection at any time period.

[0094] In this embodiment, all analysis results and original images are automatically archived by the system. Each record includes: timestamp; camera number and location information; humidity value; and corresponding segmented image. Retrieval based on time range, location conditions, etc., is supported. This data is used for later review and model training data accumulation.

[0095] This embodiment archives and stores road wetness values, road wetness area segmentation masks, target road image frame data, timestamps, camera numbers, and location information in a linked manner. The system supports retrieval by time range and location criteria for later review. Simultaneously, the improved YOLOv11-seg model is continuously iterated and optimized based on the archived data, forming a data closed loop and continuously improving detection accuracy and generalization ability.

[0096] like Figure 4 As shown, the embodiments of this application can be divided into three layers: Input layer: city road cameras, edge computing devices; Intermediate processing layer: image preprocessing module, image enhancement module, wet area detection module, and humidity calculation module; Output layer: Visual interface, historical record backtracking interface, data archiving and query system.

[0097] like Figure 5 As shown, this embodiment first performs video acquisition. High-definition cameras can be deployed at the front-end monitoring points to continuously collect video stream data from major intersections and main roads in the city. Super-resolution image enhancement technology is used to optimize image clarity, making subsequent algorithm analysis more accurate. Effective image recognition quality is maintained even in low-light scenarios such as high rainfall frequency or nighttime, laying a good foundation for humidity determination. Next, the images in the video are processed by frame segmentation to obtain image frame data. Then, a series of preprocessing operations, such as scaling, normalization, and enhancement, are performed on the image frame data to obtain the target frame image data. This embodiment uses an improved YOLO series target detection algorithm to segment the monitoring screen in real time, identifying areas that may be slippery and generating a wet area mask. To enhance the ability to perceive humidity changes, a humidity calculation module based on image region proportion is proposed. This module can accurately extract the pixel proportion of wet areas in the image and convert it into a quantifiable "regional humidity index" through a normalization function. This module can not only dynamically track humid areas, but also maintain a high recognition rate under complex weather conditions. Its output results serve as an important input for subsequent event classification and alarm mechanisms.

[0098] This embodiment has been successfully deployed in the urban road management platform, providing a highly visualized human-computer interaction interface. It supports the display of multi-dimensional data views on the command center's large screen, the web management terminal, and mobile terminals, including: real-time humidity maps and heat map annotations; historical curves and trend predictions of humidity levels for each road segment; real-time video windows and highlighted key humidity areas; humidity event rankings, alarm prompts, and detailed backtracking; and multi-segment comparative analysis and weather warning linkage display. This embodiment will respond quickly, using spatiotemporal data indexing technology to retrieve humidity indicators and video event analysis results for corresponding monitoring points and time periods, and present them to users in the form of charts, video playback, and indicator curves.

[0099] Each core model in this embodiment is trained on 500 hours of video data from real urban road cameras, covering different seasons, all weather conditions, and varying terrain and climate. The data includes road sections with different precipitation characteristics in cities in both the north and south, ensuring the generalization and adaptability of the algorithm models. Simultaneously, the training process incorporates basic visual feature models from multiple publicly available datasets and introduces a multi-scale anchoring mechanism to adapt to image resolutions and scene structures from different angles and camera models, improving the overall robustness of humid area detection and behavior recognition.

[0100] Based on the same principle as the urban road humidity analysis method provided in the embodiments of this application, the embodiments of this application also provide an urban road humidity analysis device, such as... Figure 6 As shown, the urban road humidity analysis device 20 may specifically include: an image acquisition module 21, an image preprocessing module 22, a segmentation processing module 23, and a humidity value calculation module 24. The image acquisition module 21 is used to acquire road video stream data and perform frame extraction processing on the road video stream data to obtain road image frame data. Image preprocessing module 22 is used to preprocess road image frame data to obtain target road image frame data; The segmentation processing module 23 is used to perform wet area scale recognition processing on the target road image frame data to obtain scale weight coefficients; and to perform segmentation processing based on the scale weight coefficients and the target road image frame data to obtain a road wet area segmentation mask. The humidity value calculation module 24 is used to calculate and normalize the humidity based on the road wet area segmentation mask and the target road image frame data to obtain the road humidity value.

[0101] In one embodiment of this application, the image preprocessing module 22 is specifically used for: The road image frame data is subjected to image scaling and resolution standardization to obtain road image frame data with uniform resolution. Color normalization and channel conversion are performed on road image frame data with uniform resolution to obtain road image frame data with uniform color distribution; Image enhancement processing is performed on road image frame data with uniform color distribution to obtain target road image frame data; image enhancement processing includes one or more combinations of random cropping, rotation, brightness and contrast adjustment, mirror transformation, Gaussian blur and noise injection.

[0102] In one embodiment of this application, the segmentation processing module 23 is specifically used for: Multi-path convolution feature extraction processing is performed on the target road image frame data to obtain multi-path convolution output feature data; Based on the scale weight coefficient, the multi-path convolution output feature data is weighted and summed to obtain dynamic multi-scale road image feature data. Multi-scale feature stitching and fusion processing is performed on dynamic multi-scale road image feature data to obtain fused road image feature data. Semantic segmentation processing is performed on the fused road image feature data to obtain the road wet area segmentation mask.

[0103] In one embodiment of this application, the segmentation processing module 23 is specifically used for: Channel separation processing is performed on the target road image frame data to obtain multiple sets of road image channel feature data; Convolutional kernel feature extraction was performed on multiple sets of road image channel feature data to obtain multiple sets of convolutional feature data; Multiple sets of convolutional feature data are concatenated to obtain multi-path convolutional output feature data.

[0104] In one embodiment of this application, the segmentation processing module 23 is specifically used to: perform top-down high-level semantic feature transfer processing on dynamic multi-scale road image feature data to obtain high-level road image feature data. Bottom-up low-level detail feature feedback processing is performed on dynamic multi-scale road image feature data to obtain low-level road image feature data; Tensor stitching is performed on the feature data of high-level road images and low-level road images to obtain fused road image feature data.

[0105] In one embodiment of this application, the road wet area segmentation mask includes segmentation masks corresponding to multiple road wet areas; the wetness value calculation module 24 is specifically used for: Pixel counting is performed on the road wet area segmentation mask corresponding to each road wet area to obtain the number of pixels in multiple wet areas; Pixel counting is performed on the target road image frame data to obtain the total number of pixels in the image, and the total number of pixels in the image is used as the original humidity calculation data. The pixel count of each humid region and the original humidity calculation data are processed by ratio calculation to obtain multiple original humidity ratios; The weighted average of multiple raw humidity ratios is processed to obtain weighted average humidity data; The weighted average humidity data is normalized to obtain the road humidity value.

[0106] In one embodiment of this application, the urban road humidity analysis device 20 further includes: The road moisture content values ​​are subjected to grade threshold matching processing to obtain road moisture content grade data; Visualize the road wet area segmentation mask by rendering a layer to obtain wet area mask layer data. Visualization basemap adaptation processing is performed on the target road image frame data to obtain road visualization basemap data; Based on the humid area mask layer data, the road visualization base map data is overlaid to obtain masked overlaid road image data. Based on road humidity level data, state association rendering processing is performed on camera map point data to obtain road humidity map visualization data; the camera map point data is map coordinate information data that corresponds one-to-one with the actual deployment location of urban road camera equipment. The road image data with mask overlay, road humidity level data and road humidity map visualization data are integrated and processed by a visualization interface to obtain comprehensive road humidity visualization interface data. Multi-terminal adaptation rendering is performed on the comprehensive road moisture visualization interface data to obtain multi-terminal visualization display data; The system performs multi-terminal visualization processing to push and display road moisture values.

[0107] The apparatus in this application embodiment can execute the method provided in this application embodiment, and the implementation principle is similar. The actions performed by each module in the apparatus of each embodiment of this application correspond to the steps in the method of each embodiment of this application. For detailed functional descriptions of each module of the apparatus, please refer to the descriptions in the corresponding methods shown above, which will not be repeated here.

[0108] Figure 7 A schematic diagram of the structure of an electronic device to which this application embodiment applies is shown, such as... Figure 7 As shown, the electronic device can be used to implement the methods provided in any embodiment of this application.

[0109] like Figure 7 As shown, the electronic device 300 may primarily include at least one processor 301. Figure 7 The diagram shows components such as a memory 302, a communication module 303, and an input / output interface 304. Optionally, these components can be connected and communicate with each other via a bus 305. It should be noted that... Figure 7 The structure of the electronic device 300 shown is merely illustrative and does not constitute a limitation on the electronic devices to which the methods provided in the embodiments of this application are applicable.

[0110] The memory 302 can be used to store operating systems and applications, etc. The applications can include computer programs that implement the methods shown in the embodiments of this application when invoked by the processor 301, and can also include programs for implementing other functions or services. The memory 302 can be ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, RAM (Random Access Memory) or other types of dynamic storage devices that can store information and computer programs, or it can be EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer, but is not limited thereto.

[0111] Processor 301 is connected to memory 302 via bus 305 and implements corresponding functions by calling the application programs stored in memory 302. Processor 301 can be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 301 can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0112] Electronic device 300 can connect to a network via communication module 303 (which may include, but is not limited to, components such as a network interface) to communicate with other devices (such as user terminals or servers) through the network and achieve data interaction, such as sending data to or receiving data from other devices. Communication module 303 may include wired network interfaces and / or wireless network interfaces, meaning the communication module may include at least one of wired or wireless communication modules.

[0113] The electronic device 300 can connect to necessary input / output devices, such as a keyboard and display device, via the input / output interface 304. The electronic device 300 itself may have a display device, and other display devices can also be connected externally via the interface 304. Optionally, a storage device, such as a hard drive, can also be connected via the interface 304 to store data from the electronic device 300, retrieve data from the storage device, or store data from the storage device in the memory 302. It is understood that the input / output interface 304 can be a wired interface or a wireless interface. Depending on the actual application scenario, the device connected to the input / output interface 304 can be a component of the electronic device 300 or an external device connected to the electronic device 300 when needed.

[0114] The bus 305 used to connect the components may include a path for transmitting information between the components. The bus 305 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Depending on its function, the bus 305 may be divided into an address bus, a data bus, a control bus, etc.

[0115] Optionally, for the solution provided in the embodiments of this application, the memory 302 can be used to store a computer program that executes the solution of this application, and the processor 301 runs the computer program. When the processor 301 runs the computer program, it implements the operation of the method or apparatus provided in the embodiments of this application.

[0116] Based on the same principle as the method provided in the embodiments of this application, the embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, can implement the corresponding content of the aforementioned method embodiments.

[0117] It should be noted that the terms "first," "second," "third," "fourth," "1," "2," etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in a sequence other than that shown in the figures or text.

[0118] In the embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0119] It should be understood that although arrows indicate various operation steps in the flowcharts of this application's embodiments, the order in which these steps are implemented is not limited to the order indicated by the arrows. Unless explicitly stated herein, in some implementation scenarios of this application's embodiments, the implementation steps in each flowchart can be executed in other orders as required. Furthermore, some or all steps in each flowchart, based on the actual implementation scenario, may include multiple sub-steps or multiple stages. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage can also be executed at different times. In scenarios where execution times differ, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and this application's embodiments do not limit this.

[0120] The above description is only an optional implementation method for some implementation scenarios of this application. It should be noted that for those skilled in the art, other similar implementation methods based on the technical concept of this application without departing from the technical concept of this application also fall within the protection scope of the embodiments of this application.

Claims

1. A method for analyzing the humidity of urban roads, characterized in that, include: Acquire road video stream data, and perform frame extraction processing on the road video stream data to obtain road image frame data; The road image frame data is preprocessed to obtain the target road image frame data; The target road image frame data is subjected to wet area scale recognition processing to obtain scale weight coefficients; and based on the scale weight coefficients and the target road image frame data, segmentation processing is performed to obtain a road wet area segmentation mask; Based on the road wet area segmentation mask and the target road image frame data, the wetness is calculated and normalized to obtain the road wetness value.

2. The method for analyzing urban road humidity as described in claim 1, characterized in that, The preprocessing of the road image frame data to obtain target road image frame data includes: The road image frame data is subjected to image scaling and resolution standardization to obtain road image frame data with uniform resolution. The road image frame data with uniform resolution is subjected to color normalization and channel conversion to obtain road image frame data with uniform color distribution. Image enhancement processing is performed on the road image frame data with uniform color distribution to obtain the target road image frame data; the image enhancement processing includes one or more combinations of random cropping, rotation, brightness and contrast adjustment, mirror transformation, Gaussian blur and noise injection.

3. The method for analyzing urban road humidity as described in claim 1, characterized in that, The segmentation process based on the scale weight coefficients and the target road image frame data to obtain a road wet area segmentation mask includes: The target road image frame data is subjected to multi-path convolution feature extraction processing to obtain multi-path convolution output feature data; Based on the scale weight coefficients, the multi-path convolution output feature data is weighted and summed to obtain dynamic multi-scale road image feature data. The dynamic multi-scale road image feature data is subjected to multi-scale feature stitching and fusion processing to obtain fused road image feature data; The fused road image feature data is subjected to semantic segmentation processing to obtain the road wet area segmentation mask.

4. The method for analyzing urban road humidity as described in claim 3, characterized in that, The step of performing multi-path convolution feature extraction processing on the target road image frame data to obtain multi-path convolution output feature data includes: The target road image frame data is subjected to channel separation processing to obtain multiple sets of road image channel feature data; Convolutional kernel feature extraction processing is performed on the multiple sets of road image channel feature data to obtain multiple sets of convolutional feature data; The multiple sets of convolutional feature data are concatenated to obtain the multi-path convolutional output feature data.

5. The method for analyzing urban road humidity as described in claim 3, characterized in that, The process of performing multi-scale feature stitching and fusion on the dynamic multi-scale road image feature data to obtain fused road image feature data includes: The dynamic multi-scale road image feature data is subjected to top-down high-level semantic feature transfer processing to obtain high-level road image feature data. The dynamic multi-scale road image feature data is subjected to bottom-up low-level detail feature feedback processing to obtain low-level road image feature data; The high-level road image feature data and the low-level road image feature data are subjected to tensor concatenation processing to obtain the fused road image feature data.

6. The method for analyzing urban road humidity as described in claim 1, characterized in that, The road wet area segmentation mask includes multiple segmentation masks corresponding to road wet areas; The process of calculating and normalizing the road wet region segmentation mask and the target road image frame data to obtain the road wetness value includes: Pixel counting is performed on the segmentation mask corresponding to each wet area of ​​the road to obtain the number of pixels in multiple wet areas; The target road image frame data is processed by pixel counting to obtain the total number of image pixels, and the total number of image pixels is used as the original humidity calculation data. The ratio of the number of pixels in each humid region to the original humidity calculation data is calculated to obtain multiple original humidity ratios. The weighted average of multiple raw humidity ratios is processed to obtain weighted average humidity data; The weighted average humidity data is normalized to obtain the road humidity value.

7. The method for analyzing urban road humidity as described in claim 1, characterized in that, Also includes: The road moisture content values ​​are subjected to grade threshold matching processing to obtain road moisture content grade data; The road wet area segmentation mask is subjected to visualization layer rendering processing to obtain wet area mask layer data; The target road image frame data is subjected to visualization base map adaptation processing to obtain road visualization base map data; Based on the humid area mask layer data, the road visualization base map data is overlaid to obtain masked overlaid road image data. Based on the road moisture level data, state association rendering processing is performed on the camera map point data to obtain road moisture map visualization data. The camera map location data is map coordinate information data that corresponds one-to-one with the actual deployment location of the urban road camera equipment; The road image data with mask overlay, the road humidity level data, and the road humidity map visualization data are integrated and processed by a visualization interface to obtain comprehensive road humidity visualization interface data. The road moisture content comprehensive visualization interface data is subjected to multi-terminal adaptation rendering processing to obtain multi-terminal visualization display data. The multi-terminal visualization data is pushed and displayed on multiple terminals to complete the visualization of road moisture values.

8. A device for analyzing the humidity of urban roads, characterized in that, include: The image acquisition module is used to acquire road video stream data and perform frame extraction processing on the road video stream data to obtain road image frame data. The image preprocessing module is used to preprocess the road image frame data to obtain target road image frame data; The segmentation processing module is used to perform wet area scale recognition processing on the target road image frame data to obtain scale weight coefficients; and to perform segmentation processing based on the scale weight coefficients and the target road image frame data to obtain a road wet area segmentation mask. The humidity value calculation module is used to perform humidity calculation and normalization processing based on the road humidity area segmentation mask and the target road image frame data to obtain the road humidity value.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the urban road humidity analysis method according to any one of claims 1 to 7 when running the computer program.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the urban road humidity analysis method according to any one of claims 1 to 7.