Road waterlogging early warning method and device based on segmentation distillation enhancement, and electronic equipment
By adopting the method of segmented distillation enhancement technology and multimodal information fusion in the road water accumulation early warning system, the problems of low accuracy and single results in the existing technology are solved, and more accurate and rich water accumulation detection results are achieved.
Patent Information
- Application Number
- CN202411998454.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-06
AI Technical Summary
In the prior art, road water accumulation is detected through target detection, and there are problems such as low accuracy of the detection results and single content of the detection results.
The road water accumulation warning method based on segmentation distillation is adopted, and road images are collected and prompt text are obtained, and inputted to the water accumulation recognition model for identification. The model consists of a visual encoder, a multimodal fusion network and a large language model. It improves the accuracy of feature extraction and region perception through segmentation distillation enhancement technology, and generates richer recognition results through multimodal information fusion.
It significantly improves the accuracy and richness of the water accumulation test results, solves the problems of low accuracy and single content of the test results, and improves the intelligence and accuracy of road water accumulation warnings.
Smart Images

Figure CN119942111A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing or other related technical fields, and in particular to a road waterlogging warning method and device, and electronic equipment based on segmentation and distillation enhancement. Background Art
[0002] With the acceleration of urbanization and the frequent occurrence of extreme weather, road waterlogging has become a common hidden danger that affects traffic safety, urban operation efficiency and the quality of life of citizens. Road waterlogging not only hinders the normal passage of vehicles and pedestrians, but may also cause traffic accidents, resulting in property losses and casualties. Therefore, timely and accurate road waterlogging warnings are crucial for urban traffic management and disaster response.
[0003] Among the related technologies, road waterlogging detection mainly relies on target detection technology in the field of computer vision. Target detection technology analyzes image data collected by surveillance cameras or drones to identify waterlogging features in images, such as the area, location, and shape of waterlogging. However, due to the complexity and diversity of road environments, existing single target detection technologies face many challenges in practical applications. Detection technology relies on image features, and when the boundaries of the waterlogged area are blurred or the surrounding environment is similar, false alarms are prone to occur, resulting in low accuracy of detection results and single detection results.
[0004] To address the above-mentioned problems, no effective solution has been proposed yet. Summary of the invention
[0005] The embodiments of the present invention provide a road water accumulation warning method, device, and electronic device based on segmentation and distillation enhancement, so as to at least solve the technical problems in the related art of detecting road water accumulation by target detection, such as low detection result accuracy and single detection result content.
[0006] According to one aspect of an embodiment of the present invention, a road waterlogging warning method based on segmentation distillation enhancement is provided, comprising: collecting a road image of a target road and obtaining a prompt text, wherein the prompt text comprises a plurality of heuristic prompt words associated with road waterlogging; inputting the road image and the prompt text into a waterlogging recognition model, and outputting a waterlogging recognition result for the target road, wherein the waterlogging recognition model is a model pre-trained based on the segmentation distillation enhancement technology and is used to recognize waterlogging conditions on the road; extracting features from the road image through a visual encoder of the waterlogging recognition model to obtain waterlogging features, and outputting the waterlogging recognition results for the target road. The waterlogging feature is input into the multimodal fusion network of the waterlogging recognition model, and the multimodal information of the waterlogging feature and the prompt text is fused through the multimodal fusion network to obtain a fused feature, and the fused feature is input into the large language model of the waterlogging recognition model, and the fused feature is analyzed through the large language model to obtain a waterlogging recognition result, and the waterlogging recognition result is output as output data of the waterlogging recognition model, and the waterlogging recognition result includes at least one of the following: waterlogging state, waterlogging scene, and waterlogging impact; and early warning information and processing strategies are generated for the target road according to the waterlogging recognition result.
[0007] Optionally, the step of extracting features from the road image through the visual encoder of the waterlogging recognition model to obtain the waterlogging features includes: performing region segmentation on the road image through the visual encoder to obtain K waterlogging areas in the road image, wherein K is a positive integer; performing feature extraction on each of the waterlogging areas to obtain waterlogging features of each of the waterlogging areas, wherein the waterlogging features include at least one of the following: waterlogging area features, waterlogging depth features, waterlogging scene features, and waterlogging associated features associated with waterlogging areas.
[0008] Optionally, the step of training the waterlogging recognition model includes: obtaining a set of historical road images within a historical time period, and generating instruction following data for the historical road images in the historical road image set; constructing training samples based on the historical road images and the instruction following data to obtain a training sample set; dividing the training sample set to obtain a training set and a test set; constructing an initial waterlogging recognition model, wherein the initial waterlogging recognition model includes at least a visual encoder, a multimodal fusion network, a segmentation distillation enhancement network and a large language model; using the training set to iteratively train the initial waterlogging recognition model to obtain the trained waterlogging recognition model; using the test set to test the trained waterlogging recognition model to obtain the trained waterlogging recognition model.
[0009] Optionally, the steps of obtaining a set of historical road images within a historical time period and generating instruction follow-up data for the historical road images in the set of historical road images include: acquiring historical road images through an image acquisition device to obtain a set of historical road images, wherein the set of historical road images covers road images of multiple road types, multiple scene types and multiple waterlogged states; constructing heuristic prompt words to obtain prompt text; extracting waterlogging features from the historical road images to obtain a set of historical waterlogging features, and configuring seed answers for each heuristic prompt word in the prompt text based on the set of historical waterlogging features; using a language model to perform information expansion based on the historical road images, the heuristic prompt words and the seed answers to generate multiple instruction follow-up data corresponding to the historical road images, wherein the language model is a pre-trained model for expanding text data according to text context, and each piece of the instruction follow-up data includes a heuristic prompt word and a seed answer corresponding to the heuristic prompt word.
[0010] Optionally, the step of iteratively training the initial water accumulation recognition model using the training set includes: step one, inputting the sample training in the training set into the initial water accumulation recognition model, guiding the visual encoder to extract regional features through the segmentation distillation enhancement network, and obtaining water accumulation training features; step two, fusing the water accumulation training features and prompt texts through a multimodal fusion network to obtain fused training features; step three, analyzing the fused training features through the large language model to obtain water accumulation recognition training results; step four, calculating the loss function based on the water accumulation recognition training results and the instruction following data; step five, performing back propagation according to the calculated loss function to update the model parameters of the initial water accumulation recognition model; repeating the above steps one to five to iteratively train the initial water accumulation recognition model until the iteration termination condition is reached and the iteration is ended.
[0011] Optionally, the step of guiding the visual encoder to perform regional feature extraction through the segmentation distillation enhancement network includes: for each iterative training, inputting the training samples in the training set into the segmentation distillation enhancement network in the initial water accumulation recognition model; performing regional segmentation on the historical road image in the training sample through the segmentation distillation enhancement network to obtain regional segmentation results, and generating target information based on the regional segmentation results, wherein the target information at least includes: the probability that each pixel point in the historical road image belongs to the corresponding segmented region; inputting the training sample into the visual encoder in the initial water accumulation recognition model, and using the target information as a supervisory signal to guide the visual encoder to perform regional segmentation and feature extraction on the training sample to obtain water accumulation training features, wherein, in each iterative training, the distillation loss function of the visual encoder is calculated using the target information and the water accumulation training features, and the visual encoder is guided to learn regional feature extraction by minimizing the distillation loss function.
[0012] Optionally, the step of generating warning information and processing strategies for the target road based on the water accumulation identification results includes: determining the water accumulation level of the target road based on the water accumulation identification results; when the water accumulation level is higher than a preset level threshold, determining that the target road is a warning road, and generating warning information for the warning road; for the warning road, matching a processing strategy for the warning road from a strategy database based on the water accumulation state of the warning road, the water accumulation scene and the water accumulation impact.
[0013] According to another aspect of an embodiment of the present invention, a road waterlogging warning device based on segmentation distillation enhancement is also provided, comprising: an acquisition unit, used to acquire a road image of a target road and obtain a prompt text, wherein the prompt text contains a plurality of heuristic prompt words associated with road waterlogging; an output unit, used to input the road image and the prompt text into a waterlogging recognition model, and output a waterlogging recognition result for the target road, wherein the waterlogging recognition model is a model pre-trained based on the segmentation distillation enhancement technology, and is used to recognize waterlogging conditions on the road, and features of the road image are extracted through a visual encoder of the waterlogging recognition model to obtain waterlogging features. Characteristic, and input the waterlogging feature into the multimodal fusion network of the waterlogging recognition model, perform multimodal information fusion on the waterlogging feature and the prompt text through the multimodal fusion network to obtain fused features, input the fused features into the large language model of the waterlogging recognition model, analyze the fused features through the large language model to obtain waterlogging recognition results, and output the waterlogging recognition results as output data of the waterlogging recognition model, the waterlogging recognition results include at least one of the following: waterlogging state, waterlogging scene, and waterlogging impact; a generating unit is used to generate warning information and processing strategies for the target road according to the waterlogging recognition results.
[0014] Optionally, the output unit includes: a first segmentation module, used to perform region segmentation on the road image through the visual encoder to obtain K waterlogged areas in the road image, wherein K is a positive integer; a first extraction module, used to perform feature extraction on each of the waterlogged areas to obtain waterlogging features of each of the waterlogged areas, wherein the waterlogging features include at least one of the following: waterlogging area features, waterlogging depth features, waterlogging scene features, and waterlogging associated features associated with the waterlogged areas.
[0015] Optionally, the road waterlogging warning device based on segmentation, distillation and enhancement also includes: a first acquisition unit, used to acquire a set of historical road images within a historical time period, and generate instruction following data for the historical road images in the historical road image set; a first construction unit, used to construct training samples according to the historical road images and the instruction following data to obtain a training sample set; a first division unit, used to divide the training sample set to obtain a training set and a test set; a second construction unit, used to construct an initial waterlogging recognition model, wherein the initial waterlogging recognition model includes at least a visual encoder, a multimodal fusion network, a segmentation, distillation and enhancement network, and a large language model; a first training unit, used to iteratively train the initial waterlogging recognition model using the training set to obtain the trained waterlogging recognition model; a first testing unit, used to test the trained waterlogging recognition model using the test set to obtain the trained waterlogging recognition model.
[0016] Optionally, the first acquisition unit includes: a first acquisition module, used to acquire historical road images through an image acquisition device to obtain a historical road image set, wherein the historical road image set covers road images of multiple road types, multiple scene types and multiple waterlogged states; a first construction module, used to construct heuristic prompt words to obtain prompt text; a second extraction module, used to extract waterlogging features from the historical road images to obtain a historical waterlogging feature set, and configure seed answers for each heuristic prompt word in the prompt text based on the historical waterlogging feature set; a first generation module, used to use a language model to perform information expansion based on the historical road images, the heuristic prompt words and the seed answers, and generate multiple instruction follow-up data corresponding to the historical road images, wherein the language model is a pre-trained model for expanding text data according to text context, and each of the instruction follow-up data includes a heuristic prompt word and a seed answer corresponding to the heuristic prompt word.
[0017] Optionally, the first training unit includes: a first guidance module, used to execute step one, input the sample training in the training set into the initial water accumulation recognition model, guide the visual encoder to perform regional feature extraction through the segmentation distillation enhancement network, and obtain water accumulation training features; a first fusion module, used to execute step two, fuse the water accumulation training features and the prompt text through the multimodal fusion network to obtain fused training features; a first analysis module, used to execute step three, analyze the fused training features through the large language model to obtain water accumulation recognition training results; a first calculation module, used to execute step four, calculate the loss function based on the water accumulation recognition training results and the instruction following data; a first update module, used to execute step five, perform back propagation according to the calculated loss function, and update the model parameters of the initial water accumulation recognition model; a first repetition module, used to repeat the above steps one to five, iteratively train the initial water accumulation recognition model until the iteration termination condition is reached and the iteration is ended.
[0018] Optionally, the first guidance module includes: a first input submodule, used to input the training samples in the training set into the segmentation distillation enhancement network in the initial waterlogging recognition model for each iterative training; a first generation submodule, used to perform regional segmentation on the historical road image in the training sample through the segmentation distillation enhancement network to obtain regional segmentation results, and generate target information based on the regional segmentation results, wherein the target information at least includes: the probability that each pixel point in the historical road image belongs to the corresponding segmented region; a first guidance submodule, used to input the training sample into the visual encoder in the initial waterlogging recognition model, and use the target information as a supervisory signal to guide the visual encoder to perform regional segmentation and feature extraction on the training sample to obtain waterlogging training features, wherein in each iterative training, the distillation loss function of the visual encoder is calculated using the target information and the waterlogging training features, and the visual encoder is guided to learn regional feature extraction by minimizing the distillation loss function.
[0019] Optionally, the generation unit includes: a first determination module, used to determine the water accumulation level of the target road based on the water accumulation identification result; a second generation module, used to determine that the target road is a warning road when the water accumulation level is higher than a preset level threshold, and generate warning information for the warning road; a first matching module, used to match a processing strategy for the warning road from a strategy database based on the water accumulation state of the warning road, the water accumulation scene and the water accumulation impact.
[0020] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is also provided, wherein the computer-readable storage medium includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute any one of the above-mentioned road water accumulation warning methods based on segmentation and distillation enhancement.
[0021] According to another aspect of an embodiment of the present invention, an electronic device is also provided, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement any one of the above-mentioned road water accumulation warning methods based on segmentation and distillation enhancement.
[0022] In the present application, the following steps are performed: collecting a road image of a target road and obtaining a prompt text, wherein the prompt text includes a plurality of heuristic prompt words associated with waterlogging on the road; then inputting the road image and the prompt text into a waterlogging recognition model, and outputting a waterlogging recognition result for the target road, wherein the waterlogging recognition model is a model pre-trained based on the segmentation distillation enhancement technology and is used to recognize waterlogging conditions on the road; extracting features from the road image through a visual encoder of the waterlogging recognition model to obtain waterlogging features; and inputting the waterlogging features into a multimodal fusion network of the waterlogging recognition model; performing multimodal information fusion on the waterlogging features and the prompt text through the multimodal fusion network to obtain fused features; and inputting the fused features into a large language model of the waterlogging recognition model; analyzing the fused features through the large language model to obtain waterlogging recognition results; and outputting the waterlogging recognition results as output data of the waterlogging recognition model, wherein the waterlogging recognition results include at least one of the following: waterlogging status, waterlogging scene, and waterlogging impact; and finally generating warning information and processing strategies for the target road according to the waterlogging recognition results.
[0023] In this application, road images and prompt texts are used as input data, and the waterlogging status of roads is identified and evaluated through a waterlogging recognition model. The waterlogging recognition model uses distillation segmentation enhancement technology to guide the model's visual encoder to focus more accurately on the waterlogged area in the image, significantly improving the accuracy of feature extraction and regional perception capabilities. At the same time, based on a multimodal fusion network, multimodal input data is fused to more comprehensively analyze the waterlogging situation of the target road, and more abundant recognition results are output to achieve the purpose of comprehensive and accurate waterlogging detection, and improve the accuracy of waterlogging detection results. This solves the technical problem of low accuracy and single content of detection results in the related art of detecting road waterlogging through target detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0025] Figure 1 is a flow chart of an optional road waterlogging early warning method based on segmentation and distillation enhancement according to an embodiment of the present invention;
[0026] Figure 2 is a schematic diagram of an optional road waterlogging warning process based on segmentation and distillation enhancement according to an embodiment of the present invention;
[0027] Figure 3 is a schematic diagram of an optional training sample generation process according to an embodiment of the present invention;
[0028] Figure 4 is a schematic diagram of an optional road waterlogging warning device based on segmentation and distillation enhancement according to an embodiment of the present invention;
[0029] Figure 5 It is a hardware structure block diagram of an electronic device (or mobile device) that executes a road waterlogging warning method based on segmentation and distillation enhancement according to an embodiment of the present invention. DETAILED DESCRIPTION
[0030] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0031] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0032] It should be noted that the road waterlogging warning method and device based on segmentation and distillation enhancement in the present application can be used in the field of image processing. When the waterlogging situation on the road is identified and warned based on segmentation and distillation enhancement, it can also be used in any field except the field of image processing. When the waterlogging situation on the road is identified and warned based on segmentation and distillation enhancement, the present application does not limit the application field of the road waterlogging warning method and device based on segmentation and distillation enhancement.
[0033] It should be noted that the relevant information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with the relevant laws, regulations and standards of the relevant regions, take necessary confidentiality measures, do not violate public order and good customs, and provide corresponding operation entrances for users to choose to authorize or refuse. For example, an interface is set up between this system and relevant users or organizations. Before obtaining relevant information, it is necessary to send an acquisition request to the aforementioned user or organization through the interface, and obtain relevant information after receiving the consent information fed back by the aforementioned user or organization.
[0034] It should be noted that in this application, when collecting and analyzing customer information, corresponding operation entrances are provided for users to choose to agree or reject the automated decision-making results; if the user chooses to reject, the expert decision-making process will be entered.
[0035] The following embodiments of the present invention can be applied to various road waterlogging warning systems / applications / devices. The present invention utilizes the designed segmentation distillation enhancement module, and based on the segmentation distillation enhancement mechanism, it delegates the segmentation task in the process of training the waterlogging recognition model, guides the visual encoder of the waterlogging recognition model to focus on the key areas in the image, thereby improving its feature extraction accuracy and regional perception ability, and improving the accuracy of waterlogging detection results.
[0036] The present invention uses multimodal information to fuse road images and text prompts, combined with scene understanding and cognitive reasoning, so that the waterlogging warning system can analyze the waterlogging situation more comprehensively, not only focusing on visual features, but also understanding the semantic information of the road scene through text prompts. This method improves the system's ability to perceive the hazards of waterlogging, especially in complex scenes, and can make warning judgments that meet actual needs.
[0037] By designing heuristic prompt words, the present invention can help the waterlogging recognition model to make accurate inferences in waterlogging recognition, impact assessment, warning classification and personalized processing. Different from the prompt words of general large models, the prompt word design for waterlogging detection tasks needs to focus on the characteristics of waterlogging, scene impact and actual needs, and gradually guide the model to provide intelligent decisions and warnings that conform to actual scenes in a chain of thought.
[0038] The present invention is described in detail below in conjunction with various embodiments.
[0039] Embodiment 1
[0040] According to an embodiment of the present invention, an embodiment of a road water accumulation warning method based on segmentation distillation enhancement is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0041] Figure 1 is a flow chart of an optional road waterlogging early warning method based on segmentation and distillation enhancement according to an embodiment of the present invention. Figure 1 As shown, the method comprises the following steps:
[0042] Step S101, collecting a road image of a target road and obtaining a prompt text.
[0043] It should be noted that the implementing body of the embodiment of the present invention is a road waterlogging warning system, which interacts with the urban management system and can receive road images collected by the road image acquisition equipment in the urban management system in real time, thereby jointly forming a road waterlogging detection network, realizing accurate identification and timely warning of urban road waterlogging, and improving the city's emergency management capabilities.
[0044] It should be noted that the waterlogging warning system receives real-time image data of the target road transmitted by the image acquisition device to obtain the road image, which may include the image at the current moment and the road image at the historical moment before the current moment. The image acquisition device may be a surveillance camera or drone deployed along the road, which takes images of the target road at fixed time intervals or under specific conditions (such as rain sensor triggering, user report, etc.). The images taken should be able to reflect information such as weather conditions, overall road conditions, and surrounding environment conditions.
[0045] Furthermore, in the process of real-time detection of waterlogging, it is also necessary to receive pre-built prompt text, which contains heuristic prompt words related to road waterlogging. Multiple heuristic prompt words form a thinking chain, which is progressive and used to gradually guide the model to output recognition results that are strongly related to road waterlogging. For example, "Please analyze the depth and area of waterlogging in the image and evaluate its impact on traffic."
[0046] Multimodal recognition of road waterlogging conditions is achieved through road images and prompt texts, enabling the road waterlogging warning system to more comprehensively analyze the waterlogging conditions of the target roads. It not only focuses on the visual features of the road, but also understands the semantic information of the road scene through text prompts, and outputs more accurate recognition results.
[0047] Step S102: input the road image and the prompt text into a waterlogging recognition model, and output a waterlogging recognition result for the target road.
[0048] It should be noted that after obtaining the input data, the road image and prompt text are encoded and input into the pre-built water accumulation recognition model. The water accumulation recognition model is composed of a visual encoder enhanced by segmentation and distillation, a multimodal fusion network and a large language model. Through the information recognition of the above multiple networks, the water accumulation recognition result of the target road is finally output.
[0049] Furthermore, the above-mentioned waterlogging recognition model is a model for identifying waterlogging on roads, which is obtained by pre-training based on the segmentation distillation enhancement technology. The segmentation distillation enhancement technology mainly performs segmentation distillation enhancement processing on the visual encoder of the waterlogging recognition model, guiding the training process of the visual encoder from the original picture text alignment to the regional text alignment. Existing visual encoders rely on large multimodal models, which are prone to "hallucination" phenomena in applications, causing the model to generate inaccurate or off-topic descriptions and making it difficult to focus on key areas in the image. This problem poses a challenge to the accuracy and reliability of the model. The embodiment of the present invention uses segmentation distillation enhancement to enable the visual encoder of the waterlogging recognition model to focus on key areas in the road image, thereby improving its feature extraction accuracy and regional perception capabilities, and significantly reducing the "hallucination" phenomenon of the visual encoder without increasing the amount of inference calculations.
[0050] When the waterlogging recognition model identifies waterlogging, the visual encoder of the waterlogging recognition model extracts features of the road image to obtain waterlogging features, and the waterlogging features are input into the multimodal fusion network of the waterlogging recognition model. The multimodal fusion network performs multimodal information fusion on the waterlogging features and prompt texts to obtain fused features, and the fused features are input into the large language model of the waterlogging recognition model. The fused features are analyzed by the large language model to obtain waterlogging recognition results, and the waterlogging recognition results are output as output data of the waterlogging recognition model. The waterlogging recognition results include at least one of the following: waterlogging status, waterlogging scene, and waterlogging impact, so as to obtain more comprehensive and rich recognition results and enhance the system's waterlogging warning capability and emergency handling capability.
[0051] Optionally, the step of extracting features of the road image through the visual encoder of the waterlogging recognition model to obtain waterlogging features includes: performing region segmentation on the road image through the visual encoder to obtain K waterlogging areas in the road image, where K is a positive integer; performing feature extraction on each waterlogging area to obtain waterlogging features of each waterlogging area, where the waterlogging features include at least one of the following: waterlogging area features, waterlogging depth features, waterlogging scene features, and waterlogging associated features associated with the waterlogging areas.
[0052] It should be noted that the waterlogging recognition model enhanced by segmentation and distillation can achieve regional segmentation and focus on feature extraction of waterlogging areas. Specifically, the collected image data of the target road is input into the visual encoder, and the visual encoder uses the guidance of the segmentation and distillation enhancement module to perform accurate regional segmentation on the image. This process aims to divide the image into multiple potential waterlogging areas, focusing on the key areas in the road image, that is, the parts where waterlogging may exist, and finally identifying K waterlogging areas in the road image, where K is a positive integer dynamically determined according to the number of waterlogging areas in the image. This step ensures that the model can accurately identify all waterlogging areas and provides a clear target for subsequent feature extraction.
[0053] Furthermore, after identifying K waterlogged areas, feature extraction is performed on each area to obtain the waterlogging features of each waterlogged area. The waterlogging feature extraction is obtained by further processing through the visual encoder, which aims to analyze the detailed features of each waterlogged area, including but not limited to: waterlogging area features, calculating the area size of the waterlogged area, which is crucial for evaluating the impact of waterlogging on road traffic; waterlogging depth features, the depth of waterlogging is directly related to the degree of harm caused by waterlogging, and the extraction of depth information helps the system to evaluate the risk of waterlogging; waterlogging scene features, analyzing the environmental conditions around the waterlogged area, such as traffic flow, obstacle distribution, road type, etc. This information helps to understand the specific scene of waterlogging and provide a basis for the classification of warning levels; waterlogging related features associated with the waterlogged area, identifying other elements related to the waterlogged area, such as pedestrians, vehicles, drainage facilities, etc. These features help the system evaluate the impact of waterlogging on pedestrian safety, traffic order, and the status of the road drainage system.
[0054] The intelligence and accuracy of road waterlogging warnings are significantly improved by using visual encoders to perform precise region segmentation and extract key features.
[0055] Optionally, the step of training the waterlogging recognition model includes: obtaining a set of historical road images within a historical time period, and generating instruction following data for the historical road images in the historical road image set; constructing training samples based on the historical road images and the instruction following data to obtain a training sample set; dividing the training sample set to obtain a training set and a test set; constructing an initial waterlogging recognition model, wherein the initial waterlogging recognition model includes at least a visual encoder, a multimodal fusion network, a segmentation distillation enhancement network, and a large language model; using the training set to iteratively train the initial waterlogging recognition model to obtain a trained waterlogging recognition model; using the test set to test the trained waterlogging recognition model to obtain a trained waterlogging recognition model.
[0056] It should be noted that the specific steps of training the waterlogging recognition model include obtaining historical data, constructing training samples, model initialization, iterative training and model testing to ensure the accuracy and reliability of the model in identifying road waterlogging. Specifically, first of all, it is necessary to collect a large number of historical road images. These images should be taken under different weather conditions, light intensities, and time periods, covering various road types and scenes. Ideally, the images should include multiple samples of the presence and absence of waterlogging, as well as instances of different degrees of waterlogging, to ensure that the model can fully understand the characteristics and impacts of road waterlogging. Secondly, it is necessary to configure labels for each historical road image. For each historical road image collected, the corresponding instruction following data needs to be generated. This step specifically includes multiple processes such as preliminary annotation, content expansion, and manual screening.
[0057] Furthermore, historical road images are paired with the generated instruction-following data to construct training samples. Each training sample consists of a road image and the corresponding instruction-following data. These samples will be used to train the model to learn how to extract waterlogging features from the image and generate corresponding recognition results. In order to ensure the robustness of the training and the generalization ability of the model, the training sample set needs to be fully diversified to cover different types of waterlogging scenarios. The constructed training sample set is then divided into two parts: a training set and a test set. The training set is used to update the model parameters and improve the model performance; the test set is used to verify the performance of the model on unseen data, ensure the generalization ability of the model, and avoid overfitting.
[0058] Construct an initial waterlogging recognition model, which includes at least the following components: a visual encoder, which is used to extract waterlogging features from road images; a multimodal fusion network, which is used to fuse waterlogging features with text instructions to achieve comprehensive analysis of multimodal information. A segmentation and distillation enhancement network, which is used to enhance the model's ability to focus on key areas (waterlogging areas) during training and reduce the "hallucination" phenomenon, which only exists during training. A large language model, which is used to generate recognition results for road flooding based on fused features.
[0059] Then, using the samples in the training set, the initial waterlogging recognition model is iteratively trained through optimization algorithms such as back propagation and gradient descent, and the model parameters are updated so that the model can identify waterlogging features from road images and generate corresponding recognition results according to the instruction-following data. During the training process, the segmentation distillation enhancement network guides the visual encoder to more accurately identify and segment waterlogging areas through distillation loss, thereby improving the accuracy of feature extraction.
[0060] After the model training is completed, the trained waterlogging recognition model is tested using the test set to evaluate its performance and accuracy. The test set contains road images and instruction following data that the model has never seen, which is used to verify the effect of the model in practical applications, including the accuracy of waterlogging recognition, the rationality of warning information, and the response speed of the model. Through testing, a trained waterlogging recognition model can be obtained, which has the ability to process new data and can be used in actual road waterlogging warning systems.
[0061] The training of the above-mentioned waterlogging recognition model, through the segmentation and distillation enhancement network, enables the model to more accurately identify and segment waterlogging areas, reduce the "hallucination" phenomenon, and improve the accuracy of feature extraction. The multimodal fusion network enables the model to combine visual features and text instructions to understand the complexity of waterlogging scenes, including the depth, area, location of waterlogging, and its impact on traffic. The large language model generates auxiliary warnings and strategy matching recognition results based on fusion features, providing intelligent decision support for road management, reducing the occurrence of false alarms and missed alarms, and improving the efficiency of responding to waterlogging incidents.
[0062] Optionally, the steps of obtaining a set of historical road images within a historical time period and generating instruction following data for the historical road images in the set of historical road images include: acquiring historical road images through an image acquisition device to obtain a set of historical road images, wherein the set of historical road images covers road images of multiple road types, multiple scene types, and multiple waterlogged states; constructing heuristic prompt words to obtain prompt text; extracting waterlogging features from the historical road images to obtain a set of historical waterlogging features, and configuring seed answers for each heuristic prompt word in the prompt text based on the set of historical waterlogging features; using a language model to perform information expansion based on the historical road images, the heuristic prompt words, and the seed answers to generate multiple instruction following data corresponding to the historical road images, wherein the language model is a pre-trained model for expanding text data according to a text context, and each instruction following data includes a heuristic prompt word and a seed answer corresponding to the heuristic prompt word.
[0063] It should be noted that the instruction following data is the label configured for the historical road images, which is equivalent to the standard answer of the waterlogging recognition model when identifying historical road images, and can guide the model to output the required recognition results. Specifically, the first step is to collect a large number of historical road images through image acquisition devices (such as surveillance cameras, drones, etc.) installed on different road types. These images should cover multiple road types (such as urban roads, rural roads, highways), multiple scene types (such as congestion, unmanned areas, rainy seasons, sunny days, etc.), and road images with different waterlogging conditions (from no waterlogging to severe waterlogging). Building a comprehensive and diverse image collection is the key to training models to recognize waterlogging in different scenarios.
[0064] In order to guide the model to learn how to identify and understand the characteristics of waterlogging from images, a series of heuristic prompt words can be designed. These prompt words can help the model focus on key information, such as "detecting waterlogging areas", "assessing waterlogging depth", "analyzing the impact of waterlogging on traffic", etc. These prompt words are combined into prompt texts, and each historical road image will be equipped with one or more prompt texts to guide the model to perform specific multimodal reasoning tasks.
[0065] For the collected historical road images, a seed answer can be configured for each heuristic prompt word through manual annotation. For example, for the evaluation of the waterlogged area, the seed answer may be "the waterlogged area is 20 square meters." These seed answers will serve as a reference for model training to help the model learn how to generate correct descriptions based on the prompt words.
[0066] Using the pre-trained language model, multiple instruction follow-up data are generated based on historical road images, prompt texts and their corresponding seed answers. Each instruction follow-up data contains a heuristic prompt word and its corresponding detailed description, such as "Detection of waterlogging area: Waterlogging is located on the left side of the road, with an area of 30 square meters and a depth of about 10 centimeters, affecting the passage of vehicles." Through the expansion capabilities of the language model, rich and diverse descriptions can be generated to improve the training effect of the model.
[0067] After the instruction following data is generated, it can be manually reviewed and screened to ensure data quality. Through manual screening, descriptions that do not meet expectations or have confusing semantics are removed, leaving high-quality data for model training. If the data quality does not meet the requirements after a round of generation and screening, the above steps can be repeated until a set of instruction following data that meets the quality standards is generated.
[0068] Optionally, the step of iteratively training the initial water accumulation recognition model using the training set includes: step one, inputting the sample training in the training set into the initial water accumulation recognition model, guiding the visual encoder to extract regional features through the segmentation distillation enhancement network, and obtaining water accumulation training features; step two, fusing the water accumulation training features and the prompt text through the multimodal fusion network to obtain fused training features; step three, analyzing the fused training features through the large language model to obtain the water accumulation recognition training results; step four, calculating the loss function based on the water accumulation recognition training results and the instruction following data; step five, performing back propagation according to the calculated loss function to update the model parameters of the initial water accumulation recognition model; repeating the above steps one to five, iteratively training the initial water accumulation recognition model until the iteration termination condition is reached, and the iteration ends.
[0069] It should be noted that when training the waterlogging recognition model based on the segmentation distillation enhancement technology, the samples in the training set (including historical road images and corresponding instruction follow-up data) are first input into the visual encoder in the initial waterlogging recognition model, and the input image is initially processed to extract the waterlogging features in the image, and the sample data is simultaneously input into the image segmentation model (for example, Segment Anything Model l, referred to as SAM). The segmentation distillation enhancement network is used to transfer the knowledge of the image segmentation sub-model to the visual encoder through information distillation during the training process, and the visual encoder is guided by the distillation loss to focus more on identifying and analyzing the features of the waterlogging area, reducing attention to irrelevant areas, thereby obtaining more accurate waterlogging training features.
[0070] Furthermore, the obtained waterlogging training features are fused with the prompt text in the instruction following data through a multimodal fusion network. Aligning the image features with the text information at the semantic level enables the model to understand both the visual features in the image and the semantic information in the prompt text. For example, the model can not only identify the waterlogged area in the image, but also understand the meaning of "the depth of waterlogging affects traffic", thereby generating richer recognition results.
[0071] The fused features are passed to the large language model, which performs in-depth analysis and reasoning based on the fused features to generate training results for waterlogging recognition. This result includes a detailed description of the waterlogging area, such as the depth, area, and location of the waterlogging, as well as an assessment of the impact of the waterlogging on road traffic. Through the analysis of the large language model, the model can learn how to generate recognition results that meet user preferences.
[0072] Based on the water accumulation recognition training results and the ideal results in the command following data, the model's loss function is calculated. This loss function is used to measure the difference between the recognition results generated by the model and the information provided in the command following data. By calculating the loss function, the model's performance in identifying and analyzing water accumulation characteristics can be quantified, providing a basis for subsequent parameter updates.
[0073] According to the calculated loss function, the model parameters of the initial waterlogging recognition model are updated through the back propagation algorithm. The back propagation algorithm adjusts the weights and biases of the model according to the gradient of the loss function to minimize the loss function. This process is repeated until the performance of the model on the training set reaches the predetermined iteration termination condition, such as the loss function is lower than a certain threshold, the maximum number of iterations is reached, or the model performance is stable, and the iterative training ends.
[0074] Through a sophisticated training process, combined with segmentation, distillation enhancement and multimodal information processing, the model is guided to output recognition results that are more in line with user needs, which significantly improves the performance of the water accumulation recognition model and the accuracy of the recognition results.
[0075] Optionally, the step of guiding the visual encoder to perform regional feature extraction through a segmentation distillation enhancement network includes: for each iterative training, inputting the training samples in the training set into the segmentation distillation enhancement network in the initial water accumulation recognition model; performing regional segmentation on the historical road images in the training samples through the segmentation distillation enhancement network to obtain regional segmentation results, and generating target information based on the regional segmentation results, wherein the target information at least includes: the probability that each pixel point in the historical road image belongs to the corresponding segmented region; inputting the training samples into the visual encoder in the initial water accumulation recognition model, and using the target information as a supervisory signal to guide the visual encoder to perform regional segmentation and feature extraction on the training samples to obtain water accumulation training features, wherein in each iterative training, the distillation loss function of the visual encoder is calculated using the target information and the water accumulation training features, and the visual encoder is guided to learn regional feature extraction by minimizing the distillation loss function.
[0076] It should be noted that when the segmentation distillation enhancement network is used to guide the visual encoder to extract regional features, it specifically includes: at the beginning of each iterative training, the training samples in the training set (i.e., samples containing historical road images and instruction following data) are input into the initial waterlogging recognition model, and in particular, the segmentation distillation enhancement network input into the model performs regional segmentation on the input historical road image to generate regional segmentation results. The regional segmentation result here is the result of automatically identifying the waterlogging area through the deep understanding of the image by the image segmentation sub-model in the segmentation distillation enhancement network. Based on the regional segmentation result, target information is generated, and the target information at least includes the probability that each pixel in the historical road image belongs to the corresponding segmented area. Using the target information as a supervisory signal can guide the visual encoder to more accurately identify which areas are waterlogged and which are background or non-waterlogged objects, and guide the visual encoder to learn the ability of regional segmentation.
[0077] Furthermore, the training samples are simultaneously input into the visual encoder for feature extraction. In this process, the target information serves as a supervisory signal to guide the visual encoder to focus more on the waterlogged area and extract features related to the waterlogged area. By calculating the difference between the target information (i.e., the probability that each pixel belongs to the waterlogged segmented area) and the waterlogged training features generated by the visual encoder, the distillation loss function of the visual encoder can be obtained. During multiple iterative training processes, by minimizing the distillation loss function, the visual encoder can be guided to learn how to de-segment and focus on the waterlogged area, thereby extracting regional features more accurately.
[0078] Step S103, generating warning information and processing strategies for the target road according to the waterlogging identification result.
[0079] It should be noted that according to the water accumulation status in the water accumulation identification results, that is, whether there is water accumulation, the water accumulation scene, such as urban roads or rural roads, and the impact of water accumulation, that is, the impact of water accumulation on pedestrians, vehicles, or public facilities, it can be determined whether the water accumulation on the target road is dangerous, and automatically generate warning information and treatment strategies for dangerous road water accumulation. Based on the precise identification and analysis of the model, the system can provide more specific and accurate warning information, as well as targeted treatment suggestions, to ensure that road management departments can respond quickly and take effective measures in a timely manner.
[0080] Optionally, the step of generating warning information and processing strategies for the target road based on the water accumulation identification results includes: determining the water accumulation level of the target road based on the water accumulation identification results; when the water accumulation level is higher than a preset level threshold, determining the target road as a warning road, and generating warning information for the warning road; for the warning road, matching the processing strategy for the warning road from the strategy database based on the water accumulation status, water accumulation scene and water accumulation impact of the warning road.
[0081] Specifically, when generating warning information and processing strategies, the water accumulation level of the target road is first determined based on the water accumulation identification results. The water accumulation level is divided according to information such as the depth, area and degree of impact on road traffic, ranging from mild to severe, such as "mild water accumulation", "moderate water accumulation", "severe water accumulation", etc. If the water accumulation level is higher than the preset level threshold, that is, the water accumulation level reaches a level that requires attention, the system will determine that the target road is a warning road. Warning roads refer to roads where water accumulation may have a significant impact on road traffic safety, traffic efficiency or the surrounding environment. For warning roads, the system automatically generates warning information, which includes the specific location of the water accumulation, the water accumulation level, the possible scope of impact, etc., and is presented in the form of text, images or voice to ensure that road management departments can quickly understand and respond.
[0082] While generating warning information, the system also needs to formulate a processing strategy. The processing strategy can be pre-configured based on historical waterlogging conditions. The processing strategy is matched for the warning road from the strategy database through the waterlogging status, waterlogging scene, and waterlogging impact of the warning road. The formulation of the processing strategy should be combined with the waterlogging identification results and the specific scene information of the road. For example, on a busy urban trunk road, even if the area of waterlogging is not large, it may affect the safety of vehicle traffic. The system may recommend that personnel be dispatched immediately for drainage. On remote rural roads, where the area of waterlogging is large but the impact on traffic is small, the system may recommend observing changes in waterlogging and deciding whether it needs to be handled depending on the situation.
[0083] Through the above steps, a road image of the target road is collected, and a prompt text is obtained, wherein the prompt text includes a plurality of heuristic prompt words associated with waterlogging on the road, and then the road image and the prompt text are input into a waterlogging recognition model, and a waterlogging recognition result of the target road is output, wherein the waterlogging recognition model is a model pre-trained based on the segmentation distillation enhancement technology, and is used to recognize waterlogging conditions on the road. The road image is feature extracted by the visual encoder of the waterlogging recognition model to obtain waterlogging features, and the waterlogging features are input into a multimodal fusion network of the waterlogging recognition model. The multimodal fusion network performs multimodal information fusion on the waterlogging features and the prompt text to obtain fused features, and the fused features are input into a large language model of the waterlogging recognition model. The fused features are analyzed by the large language model to obtain waterlogging recognition results, and the waterlogging recognition results are output as output data of the waterlogging recognition model, and the waterlogging recognition results include at least one of the following: waterlogging status, waterlogging scene, and waterlogging impact. Finally, according to the waterlogging recognition results, warning information and processing strategies are generated for the target road.
[0084] In this embodiment, the road image and prompt text are used as input data, and the waterlogging status of the road is identified and evaluated through the waterlogging recognition model. The waterlogging recognition model uses distillation segmentation enhancement technology to guide the model's visual encoder to focus more accurately on the waterlogging area in the image, significantly improving the accuracy of feature extraction and regional perception capabilities. At the same time, based on the multimodal fusion network, the multimodal input data is fused, which can more comprehensively analyze the waterlogging situation of the target road and output richer recognition results to achieve the purpose of comprehensive and accurate waterlogging detection, and improve the accuracy of waterlogging detection results. This solves the technical problems of low accuracy and single content of detection results in the related technology of detecting road waterlogging through target detection.
[0085] Another optional specific implementation is described in detail below.
[0086] Figure 2 is a schematic diagram of an optional road waterlogging warning process based on segmentation and distillation enhancement according to an embodiment of the present invention, such as Figure 2As shown, the entire waterlogging warning process specifically includes: in the training stage, the road image (i.e., historical road image) is recognized simultaneously through SAM and the visual encoder, and then the knowledge of SAM is transferred to the visual encoder through the segmentation distillation enhancement module, so that the visual encoder can learn the regional segmentation ability, and the visual encoder can focus on the key areas in the image when performing feature extraction. In the real-time reasoning stage, the real-time road image is input into the visual encoder, the recognition encoder outputs the waterlogging features, the waterlogging features and the prompt text are fused through the multimodal fusion network, and the enhanced visual encoding (waterlogging features) is mapped to the text space to realize the fusion of multimodal information and obtain the fusion features. Finally, the waterlogging recognition results are output through the language model. Specifically,
[0087] Data generation: Its main function is to provide the model with a rich and diverse input data set required for training, including the collection of road scene data and the generation of corresponding instruction follow-up data for each picture.
[0088] The heuristic prompt word design provides directional guidance to the model, enabling it to generate waterlogging recognition results that meet the requirements according to the actual scenario during the reasoning process.
[0089] The visual encoder is responsible for processing image input and extracting water accumulation and scene information, such as the area, depth, and location of water accumulation, as well as vehicles and pedestrians in the scene.
[0090] The Segmentation Distillation Enhanced Module (SAM-adapter) transfers the knowledge of SAM to the visual encoder through information distillation during the training process. The Segmentation Distillation Enhanced Module guides the training process of the waterlogging recognition model from the original image text alignment to regional text alignment through the proxy segmentation task, so that the visual encoder of the waterlogging recognition model focuses on the key areas in the image, thereby improving its feature extraction accuracy and regional perception ability. Without increasing the amount of inference calculation, the "hallucination" phenomenon of the visual encoder is significantly reduced.
[0091] The multimodal fusion network module maps the enhanced visual encoding (water accumulation features) to the text space to achieve the fusion of multimodal information, so that the visual information can be reasoned and understood together with the text information. The core task of this module is to ensure that the features extracted from the image can be aligned with the text features in the same semantic space.
[0092] Large language model, the pre-trained large language model is integrated in the waterlogging recognition model. Through the input of multimodal information (images and text prompts), it can infer the state of waterlogging, the impact of waterlogging and other information, and output the waterlogging recognition results.
[0093] When iteratively training the waterlogging recognition model, you first need to build training samples, which is the data basis for model training. Figure 3 is a schematic diagram of an optional training sample generation process according to an embodiment of the present invention, such as Figure 3 As shown in Figure 1, the training sample generation process specifically includes:
[0094] Step 1: image acquisition;
[0095] Historical road images under different times, weather, and lighting conditions are collected through surveillance cameras, drones, and other equipment. The dataset covers various road types (urban roads, rural roads, highways, etc.) and complex scenes (congestion, uninhabited areas, etc.).
[0096] Step 2: Annotate the image and generate seed answers;
[0097] First, it is necessary to design heuristic prompt words for the waterlogging recognition model based on the waterlogging scene on the road, guide the model to reason about specific tasks, provide context for model analysis, guide the model to focus on the specific problem of waterlogging, help the model understand the details in complex images, and combine multimodal information for reasonable reasoning. For example, "Please analyze the area and depth of waterlogging on the road and evaluate its impact on traffic." Through this prompt, the model will focus on the image features related to waterlogging and provide reasonable results based on the prompt words.
[0098] Heuristic prompt words can also guide multimodal information fusion. Through clear language expression, it guides the waterlogging recognition model to extract key information that matches the text context from the image and enhance the semantic consistency between modalities. In complex scenes, it helps the model focus on visual elements related to the task and ignore irrelevant information to avoid reasoning bias. For example, "Does the road waterlogging in the image pose a threat to pedestrian safety? Please evaluate the waterlogging level of the target road." The model will find the depth of the waterlogging from the visual information based on the prompt and interact with the pedestrians in the scene to generate more targeted results.
[0099] Heuristic prompt words can also guide the model to reason step by step through the "thinking chain", understand the implicit meaning in the visual information, and help the model generate output results that are more in line with the actual scene during the reasoning process. For example, "The current scene is a busy road in the city center. What problems will waterlogging cause?" The model understands the context of the prompt words, not only analyzes the waterlogging situation, but also infers its impact on traffic and outputs information about the impact of waterlogging.
[0100] Furthermore, based on the heuristic prompt words, a seed answer is provided for each historical road image to help it extract key information from the image. For example, if there is water in the image, the question "Is there water in the image?" is marked as "there is water".
[0101] Step 3: Expand through a large language model;
[0102] Based on the collected pictures, heuristic prompt words and seed answers, the language model is used to automatically expand the seed answers and generate detailed instruction following data.
[0103] Step 4: Manual filtration;
[0104] Manually screen the instruction-following data generated by the language model to remove content that does not meet expectations or is semantically confusing, and use the screening results as seed answers to generate richer data until the requirements are met.
[0105] Step 5: End. Generate training samples based on historical road images and the instruction following data obtained after annotation.
[0106] In an embodiment of the present invention, the training process keeps the weights of the visual encoder and the large language model unchanged, and only updates the segmentation distillation enhancement module and the visual encoding mapping module. The segmentation distillation enhancement module is only used during training and not during deployment and inference, thereby improving the accuracy of feature extraction and regional perception capabilities without increasing the amount of computation.
[0107] The segmentation distillation enhancement module designed in the embodiment of the present invention delegates the segmentation task during the training of the water accumulation recognition model based on the segmentation distillation enhancement mechanism, and guides the visual encoder of the water accumulation recognition model to focus on key areas in the image, thereby improving its feature extraction accuracy and regional perception ability, and improving the accuracy of water accumulation detection results.
[0108] The embodiment of the present invention uses multimodal information to fuse road images and text prompts, and combines scene understanding and cognitive reasoning to enable the waterlogging warning system to analyze the waterlogging situation more comprehensively, not only focusing on visual features, but also understanding the semantic information of the road scene through text prompts. This method improves the system's ability to perceive the hazards of waterlogging, especially in complex scenes, and can make warning judgments that meet actual needs.
[0109] The embodiment of the present invention can help the waterlogging recognition model to make accurate inferences in waterlogging recognition, impact assessment, warning classification and personalized processing by designing heuristic prompt words. Different from the prompt words of general large models, the prompt word design for waterlogging detection tasks needs to focus on the characteristics of waterlogging, scene impact and actual needs, and gradually guide the model to provide intelligent decisions and warnings that conform to actual scenes in a chain of thought mode.
[0110] The following is a detailed description in conjunction with another embodiment.
[0111] Embodiment 2
[0112] A road water accumulation warning device based on segmentation distillation enhancement provided in this embodiment includes multiple implementation units, each implementation unit corresponds to each implementation step in the above-mentioned embodiment one. Its specific implementation method and beneficial effects can refer to the above-mentioned method embodiment and will not be repeated here.
[0113] Figure 4 is a schematic diagram of an optional road waterlogging warning device based on segmentation and distillation enhancement according to an embodiment of the present invention, such as Figure 4 As shown, the road waterlogging warning device based on segmentation and distillation enhancement may include: a collection unit 41, an output unit 42, and a generation unit 43, wherein:
[0114] The acquisition unit 41 is used to acquire a road image of a target road and obtain a prompt text, wherein the prompt text includes a plurality of heuristic prompt words associated with road waterlogging;
[0115] The output unit 42 is used to input the road image and the prompt text into the waterlogging recognition model, and output the waterlogging recognition result of the target road, wherein the waterlogging recognition model is a model pre-trained based on the segmentation distillation enhancement technology, and is used to identify the waterlogging situation of the road. The visual encoder of the waterlogging recognition model extracts features of the road image to obtain waterlogging features, and the waterlogging features are input into the multimodal fusion network of the waterlogging recognition model. The multimodal fusion network performs multimodal information fusion on the waterlogging features and the prompt text to obtain fusion features, and the fusion features are input into the large language model of the waterlogging recognition model. The fusion features are analyzed by the large language model to obtain waterlogging recognition results, and the waterlogging recognition results are output as output data of the waterlogging recognition model, and the waterlogging recognition results include at least one of the following: waterlogging state, waterlogging scene, and waterlogging impact;
[0116] The generating unit 43 is used to generate warning information and processing strategies for the target road according to the water accumulation identification result.
[0117] The above-mentioned road waterlogging warning device based on segmentation and distillation enhancement collects the road image of the target road through the collection unit 41 and obtains the prompt text, wherein the prompt text includes multiple heuristic prompt words associated with road waterlogging; the road image and the prompt text are input into the waterlogging recognition model through the output unit 42, and the waterlogging recognition result of the target road is output, wherein the waterlogging recognition model is a model pre-trained based on the segmentation and distillation enhancement technology, and is used to identify the waterlogging situation of the road; the road image is extracted by the visual encoder of the waterlogging recognition model to obtain the waterlogging feature, and the waterlogging feature is input into the multimodal fusion network of the waterlogging recognition model; the waterlogging feature and the prompt text are multimodally fused through the multimodal fusion network to obtain the fusion feature, and the fusion feature is input into the large language model of the waterlogging recognition model; the fusion feature is analyzed through the large language model to obtain the waterlogging recognition result, and the waterlogging recognition result is output as the output data of the waterlogging recognition model, and the waterlogging recognition result includes at least one of the following: waterlogging state, waterlogging scene, and waterlogging impact; the warning information and processing strategy are generated for the target road according to the waterlogging recognition result through the generation unit 43.
[0118] In this embodiment, the road image and prompt text are used as input data, and the waterlogging status of the road is identified and evaluated through the waterlogging recognition model. The waterlogging recognition model uses distillation segmentation enhancement technology to guide the model's visual encoder to focus more accurately on the waterlogging area in the image, significantly improving the accuracy of feature extraction and regional perception capabilities. At the same time, based on the multimodal fusion network, the multimodal input data is fused, which can more comprehensively analyze the waterlogging situation of the target road and output richer recognition results to achieve the purpose of comprehensive and accurate waterlogging detection, and improve the accuracy of waterlogging detection results. This solves the technical problems of low accuracy and single content of detection results in the related technology of detecting road waterlogging through target detection.
[0119] Optionally, the output unit 42 includes: a first segmentation module, used to perform region segmentation on the road image through a visual encoder to obtain K waterlogged areas in the road image, where K is a positive integer; a first extraction module, used to perform feature extraction on each waterlogged area to obtain waterlogging features of each waterlogged area, where the waterlogging features include at least one of the following: waterlogging area features, waterlogging depth features, waterlogging scene features, and waterlogging associated features associated with the waterlogged areas.
[0120] Optionally, the road waterlogging warning device based on segmentation distillation enhancement also includes: a first acquisition unit, used to acquire a set of historical road images within a historical time period, and generate instruction following data for the historical road images in the historical road image set; a first construction unit, used to construct training samples based on the historical road images and the instruction following data to obtain a training sample set; a first division unit, used to divide the training sample set to obtain a training set and a test set; a second construction unit, used to construct an initial waterlogging recognition model, wherein the initial waterlogging recognition model includes at least a visual encoder, a multimodal fusion network, a segmentation distillation enhancement network and a large language model; a first training unit, used to iteratively train the initial waterlogging recognition model using the training set to obtain a trained waterlogging recognition model; a first testing unit, used to test the trained waterlogging recognition model using the test set to obtain a trained waterlogging recognition model.
[0121] Optionally, the first acquisition unit includes: a first acquisition module, which is used to acquire historical road images through an image acquisition device to obtain a historical road image set, wherein the historical road image set covers road images of multiple road types, multiple scene types and multiple waterlogged states; a first construction module, which is used to construct heuristic prompt words to obtain prompt text; a second extraction module, which is used to extract waterlogging features from historical road images to obtain a historical waterlogging feature set, and configure seed answers for each heuristic prompt word in the prompt text based on the historical waterlogging feature set; a first generation module, which is used to use a language model to perform information expansion based on historical road images, heuristic prompt words and seed answers, and generate multiple instruction follow-up data corresponding to the historical road images, wherein the language model is a pre-trained model for expanding text data according to text context, and each instruction follow-up data includes a heuristic prompt word and a seed answer corresponding to the heuristic prompt word.
[0122] Optionally, the first training unit includes: a first guidance module, used to execute step one, input the sample training in the training set into the initial water accumulation recognition model, guide the visual encoder to perform regional feature extraction through the segmentation distillation enhancement network, and obtain the water accumulation training feature; a first fusion module, used to execute step two, fuse the water accumulation training feature and the prompt text through the multimodal fusion network to obtain the fused training feature; a first analysis module, used to execute step three, analyze the fused training feature through the large language model, and obtain the water accumulation recognition training result; a first calculation module, used to execute step four, calculate the loss function based on the water accumulation recognition training result and the instruction following data; a first update module, used to execute step five, perform back propagation according to the calculated loss function, and update the model parameters of the initial water accumulation recognition model; a first repetition module, used to repeat the above steps one to five, iteratively train the initial water accumulation recognition model until the iteration termination condition is reached and the iteration ends.
[0123] Optionally, the first guidance module includes: a first input submodule, used to input the training samples in the training set into the segmentation distillation enhancement network in the initial water accumulation recognition model for each iterative training; a first generation submodule, used to perform regional segmentation on the historical road images in the training samples through the segmentation distillation enhancement network to obtain regional segmentation results, and generate target information based on the regional segmentation results, wherein the target information at least includes: the probability that each pixel point in the historical road image belongs to the corresponding segmented region; a first guidance submodule, used to input the training samples into the visual encoder in the initial water accumulation recognition model, and use the target information as a supervision signal to guide the visual encoder to perform regional segmentation and feature extraction on the training samples to obtain water accumulation training features, wherein in each iterative training, the distillation loss function of the visual encoder is calculated using the target information and the water accumulation training features, and the visual encoder is guided to learn regional feature extraction by minimizing the distillation loss function.
[0124] Optionally, the generation unit 43 includes: a first determination module, used to determine the water accumulation level of the target road based on the water accumulation identification result; a second generation module, used to determine that the target road is a warning road when the water accumulation level is higher than a preset level threshold, and generate warning information for the warning road; a first matching module, used to match a processing strategy for the warning road from a strategy database based on the water accumulation state, water accumulation scene and water accumulation impact of the warning road.
[0125] The above-mentioned road water accumulation warning device based on segmentation distillation enhancement can also include a processor and a memory. The above-mentioned collection unit 41, output unit 42, generation unit 43, etc. are all stored in the memory as program units, and the processor executes the above-mentioned program units stored in the memory to realize the corresponding functions.
[0126] The processor includes a kernel, which retrieves the corresponding program unit from the memory. One or more kernels can be set, and the road waterlogging situation can be identified and warned by adjusting kernel parameters.
[0127] The above-mentioned memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (fl ash RAM), and the memory includes at least one storage chip.
[0128] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute any one of the above-mentioned road waterlogging warning methods based on segmentation and distillation enhancement.
[0129] According to another aspect of an embodiment of the present invention, an electronic device is also provided, including one or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by one or more processors, the one or more processors implement any one of the above-mentioned road water accumulation warning methods based on segmentation distillation enhancement.
[0130] According to another aspect of an embodiment of the present invention, a computer program product is also provided, which includes a computer program, wherein when the computer program is executed by a processor, any one of the above-mentioned road waterlogging warning methods based on segmentation and distillation enhancement is implemented.
[0131] The present application also provides a computer program product, which, when executed on a data processing device, is suitable for executing a program that is initialized with the following method steps: collecting a road image of a target road and obtaining a prompt text, wherein the prompt text includes a plurality of heuristic prompt words associated with waterlogging on the road; inputting the road image and the prompt text into a waterlogging recognition model, and outputting a waterlogging recognition result for the target road, wherein the waterlogging recognition model is a model pre-trained based on the segmentation distillation enhancement technology and is used to recognize waterlogging conditions on the road; extracting features of the road image through a visual encoder of the waterlogging recognition model to obtain waterlogging features, and inputting the waterlogging features into a multimodal fusion network of the waterlogging recognition model; performing multimodal information fusion on the waterlogging features and the prompt text through the multimodal fusion network to obtain fused features; inputting the fused features into a large language model of the waterlogging recognition model; analyzing the fused features through the large language model to obtain waterlogging recognition results; and outputting the waterlogging recognition results as output data of the waterlogging recognition model, wherein the waterlogging recognition results include at least one of the following: waterlogging status, waterlogging scene, and waterlogging impact; generating warning information and processing strategies for the target road according to the waterlogging recognition results.
[0132] Optionally, the step of extracting features of the road image through the visual encoder of the waterlogging recognition model to obtain waterlogging features includes: performing region segmentation on the road image through the visual encoder to obtain K waterlogging areas in the road image, where K is a positive integer; performing feature extraction on each waterlogging area to obtain waterlogging features of each waterlogging area, where the waterlogging features include at least one of the following: waterlogging area features, waterlogging depth features, waterlogging scene features, and waterlogging associated features associated with the waterlogging areas.
[0133] Optionally, the step of training the waterlogging recognition model includes: obtaining a set of historical road images within a historical time period, and generating instruction following data for the historical road images in the historical road image set; constructing training samples based on the historical road images and the instruction following data to obtain a training sample set; dividing the training sample set to obtain a training set and a test set; constructing an initial waterlogging recognition model, wherein the initial waterlogging recognition model includes at least a visual encoder, a multimodal fusion network, a segmentation distillation enhancement network, and a large language model; using the training set to iteratively train the initial waterlogging recognition model to obtain a trained waterlogging recognition model; using the test set to test the trained waterlogging recognition model to obtain a trained waterlogging recognition model.
[0134] Optionally, the steps of obtaining a set of historical road images within a historical time period and generating instruction following data for the historical road images in the set of historical road images include: acquiring historical road images through an image acquisition device to obtain a set of historical road images, wherein the set of historical road images covers road images of multiple road types, multiple scene types, and multiple waterlogged states; constructing heuristic prompt words to obtain prompt text; extracting waterlogging features from the historical road images to obtain a set of historical waterlogging features, and configuring seed answers for each heuristic prompt word in the prompt text based on the set of historical waterlogging features; using a language model to perform information expansion based on the historical road images, the heuristic prompt words, and the seed answers to generate multiple instruction following data corresponding to the historical road images, wherein the language model is a pre-trained model for expanding text data according to a text context, and each instruction following data includes a heuristic prompt word and a seed answer corresponding to the heuristic prompt word.
[0135] Optionally, the step of iteratively training the initial water accumulation recognition model using the training set includes: step one, inputting the sample training in the training set into the initial water accumulation recognition model, guiding the visual encoder to extract regional features through the segmentation distillation enhancement network, and obtaining water accumulation training features; step two, fusing the water accumulation training features and the prompt text through the multimodal fusion network to obtain fused training features; step three, analyzing the fused training features through the large language model to obtain the water accumulation recognition training results; step four, calculating the loss function based on the water accumulation recognition training results and the instruction following data; step five, performing back propagation according to the calculated loss function to update the model parameters of the initial water accumulation recognition model; repeating the above steps one to five, iteratively training the initial water accumulation recognition model until the iteration termination condition is reached, and the iteration ends.
[0136] Optionally, the step of guiding the visual encoder to perform regional feature extraction through a segmentation distillation enhancement network includes: for each iterative training, inputting the training samples in the training set into the segmentation distillation enhancement network in the initial water accumulation recognition model; performing regional segmentation on the historical road images in the training samples through the segmentation distillation enhancement network to obtain regional segmentation results, and generating target information based on the regional segmentation results, wherein the target information at least includes: the probability that each pixel point in the historical road image belongs to the corresponding segmented region; inputting the training samples into the visual encoder in the initial water accumulation recognition model, and using the target information as a supervisory signal to guide the visual encoder to perform regional segmentation and feature extraction on the training samples to obtain water accumulation training features, wherein in each iterative training, the distillation loss function of the visual encoder is calculated using the target information and the water accumulation training features, and the visual encoder is guided to learn regional feature extraction by minimizing the distillation loss function.
[0137] Optionally, the step of generating warning information and processing strategies for the target road based on the water accumulation identification results includes: determining the water accumulation level of the target road based on the water accumulation identification results; when the water accumulation level is higher than a preset level threshold, determining the target road as a warning road, and generating warning information for the warning road; for the warning road, matching the processing strategy for the warning road from the strategy database based on the water accumulation status, water accumulation scene and water accumulation impact of the warning road.
[0138] Figure 5 is a hardware structure block diagram of an electronic device (or mobile device) for executing a road waterlogging warning method based on segmentation and distillation enhancement according to an embodiment of the present invention. Figure 5 As shown, the electronic device may include one or more processors ( Figure 5 502a, 502b, ..., 502n are used to illustrate that the processor may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 504 for storing data. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a keyboard, a power supply and / or a camera. A person skilled in the art can understand that Figure 5 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 5 More or fewer components as shown, or with Figure 5 Different configurations shown.
[0139] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0140] In the above embodiments of the present invention, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0141] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units can be a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0142] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0143] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0144] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk and other media that can store program codes.
[0145] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A road flooding warning method based on segmentation and distillation enhancement, characterized in that: include: Collecting a road image of a target road and acquiring a prompt text, wherein the prompt text includes a plurality of heuristic prompt words associated with road waterlogging; The road image and the prompt text are input into a waterlogging recognition model, and a waterlogging recognition result for a target road is output, wherein the waterlogging recognition model is a model pre-trained based on the segmentation distillation enhancement technology and used to recognize waterlogging conditions on roads; the road image is subjected to feature extraction by a visual encoder of the waterlogging recognition model to obtain waterlogging features; the waterlogging features are input into a multimodal fusion network of the waterlogging recognition model; the waterlogging features and the prompt text are subjected to multimodal information fusion by the multimodal fusion network to obtain fusion features; the fusion features are input into a large language model of the waterlogging recognition model; the fusion features are analyzed by the large language model to obtain waterlogging recognition results; and the waterlogging recognition results are output as output data of the waterlogging recognition model, and the waterlogging recognition results include at least one of the following: waterlogging status, waterlogging scene, and waterlogging impact; According to the waterlogging identification result, early warning information and processing strategies are generated for the target road.
2. The method according to claim 1, characterized in that The step of extracting features from the road image by using the visual encoder of the waterlogging recognition model to obtain the waterlogging features comprises: Performing region segmentation on the road image by using the visual encoder to obtain K waterlogged areas in the road image, where K is a positive integer; Feature extraction is performed on each of the waterlogged areas to obtain waterlogged features of each of the waterlogged areas, wherein the waterlogged features include at least one of the following: waterlogged area features, waterlogged depth features, waterlogged scene features, and waterlogged associated features associated with the waterlogged areas.
3. The method according to claim 1, characterized in that The steps of training the water accumulation recognition model include: Acquire a historical road image set within a historical time period, and generate instruction follow-up data for the historical road images in the historical road image set; Constructing training samples according to the historical road images and the instruction following data to obtain a training sample set; Dividing the training sample set to obtain a training set and a test set; Constructing an initial water accumulation recognition model, wherein the initial water accumulation recognition model at least includes a visual encoder, a multimodal fusion network, a segmentation distillation enhancement network, and a large language model; Iteratively training the initial waterlogging recognition model using the training set to obtain the trained waterlogging recognition model; The trained waterlogging recognition model is tested using the test set to obtain the trained waterlogging recognition model.
4. The method according to claim 3, characterized in that The steps of acquiring a historical road image set within a historical time period and generating instruction follow-up data for the historical road images in the historical road image set include: Collecting historical road images through an image acquisition device to obtain a historical road image set, wherein the historical road image set covers road images of multiple road types, multiple scene types, and multiple waterlogged conditions; Construct heuristic prompt words and obtain prompt text; Extracting waterlogging features from the historical road image to obtain a historical waterlogging feature set, and configuring a seed answer for each heuristic prompt word in the prompt text based on the historical waterlogging feature set; A language model is used to perform information expansion based on the historical road image, the heuristic prompt word and the seed answer to generate multiple instruction follow-up data corresponding to the historical road image, wherein the language model is a pre-trained model for expanding text data according to text context, and each instruction follow-up data includes a heuristic prompt word and a seed answer corresponding to the heuristic prompt word.
5. The method according to claim 3, characterized in that: The step of iteratively training the initial waterlogging recognition model using the training set comprises: Step 1: Input the sample training in the training set into the initial waterlogging recognition model, guide the visual encoder to extract regional features through the segmentation distillation enhancement network, and obtain waterlogging training features; Step 2: The water accumulation training features and the prompt text are fused through a multimodal fusion network to obtain fused training features; Step 3, analyzing the fusion training features through the large language model to obtain water accumulation recognition training results; Step 4, calculating a loss function based on the water accumulation recognition training result and the instruction following data; Step 5, performing back propagation according to the calculated loss function to update the model parameters of the initial waterlogging recognition model; Repeat the above steps 1 to 5 to iteratively train the initial waterlogging recognition model until the iteration termination condition is reached and the iteration is terminated.
6. The method according to claim 5, characterized in that The step of guiding the visual encoder to extract regional features through the segmentation distillation enhancement network includes: For each iterative training, the training samples in the training set are input into the segmentation and distillation enhancement network in the initial water accumulation recognition model; Performing regional segmentation on the historical road image in the training sample through the segmentation distillation enhancement network to obtain a regional segmentation result, and generating target information based on the regional segmentation result, wherein the target information at least includes: a probability that each pixel point in the historical road image belongs to a corresponding segmented area; The training samples are input into the visual encoder in the initial water accumulation recognition model, and the target information is used as a supervisory signal to guide the visual encoder to perform region segmentation and feature extraction on the training samples to obtain water accumulation training features, wherein in each iterative training, a distillation loss function of the visual encoder is calculated using the target information and the water accumulation training features, and the visual encoder is guided to learn regional feature extraction by minimizing the distillation loss function.
7. The method according to claim 1, characterized in that The step of generating warning information and processing strategies for the target road based on the water accumulation identification result includes: Determining a waterlogging level of the target road based on the waterlogging identification result; When the water accumulation level is higher than a preset level threshold, determining the target road as a warning road, and generating warning information for the warning road; For the warning road, a processing strategy is matched for the warning road from a strategy database based on the waterlogging state, the waterlogging scene and the waterlogging impact of the warning road.
8. A road waterlogging warning device based on segmentation and distillation enhancement, characterized in that: include: A collection unit, used for collecting a road image of a target road and obtaining a prompt text, wherein the prompt text includes a plurality of heuristic prompt words associated with road waterlogging; an output unit, for inputting the road image and the prompt text into a waterlogging recognition model, and outputting a waterlogging recognition result for a target road, wherein the waterlogging recognition model is a model pre-trained based on the segmentation distillation enhancement technology, and is used to recognize waterlogging conditions on a road; a visual encoder of the waterlogging recognition model is used to extract features from the road image to obtain waterlogging features; the waterlogging features are input into a multimodal fusion network of the waterlogging recognition model; the multimodal fusion network is used to perform multimodal information fusion on the waterlogging features and the prompt text to obtain fusion features; the fusion features are input into a large language model of the waterlogging recognition model; the fusion features are analyzed by the large language model to obtain waterlogging recognition results; and the waterlogging recognition results are output as output data of the waterlogging recognition model, and the waterlogging recognition results include at least one of the following: waterlogging status, waterlogging scene, and waterlogging impact; A generating unit is used to generate warning information and a processing strategy for the target road according to the water accumulation identification result.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the road waterlogging warning method based on segmentation distillation enhancement as described in any one of claims 1 to 7.
10. An electronic device, characterized in that: It includes one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the road water accumulation warning method based on segmentation distillation enhancement as described in any one of claims 1 to 7.
Citation Information
Cited By
Express delivery network inlet water treatment method, electronic equipment and program product
CN121708540A