Large waterlogging detection model training method, training device, detection method and equipment

By building a large flood detection model, using labeled and unlabeled data sets to generate pseudo-labels and optimize model parameters, the problems of high labeling costs and limited model performance in the existing technology are solved, and high-precision and low-cost flood detection are achieved.

CN120356034AActive Publication Date: 2025-07-22CHONGQING UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510445471.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-04-09
Filing Date
2025-04-10
Publication Date
2025-07-22
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

The existing urban flood detection technology relies on a fully supervised semantic segmentation model, which requires a large amount of labeling data and is expensive to label. The semi-supervised algorithm has insufficient consistency regular constraints, the pseudo-label confidence threshold is too high, resulting in low data utilization, and the visual basic model lacks domain knowledge. The traditional adapter fine-tuning method limits the model performance.

Method used

A large model of flood detection is constructed, a model is trained using labeled and unlabeled data sets, a pseudo-label is generated through preset rules, and a loss function is calculated to optimize the model parameters, explicitly constrain the consistency of the prediction results of strong perturbation branches, and a pseudo-label that meets the threshold is selected for model training.

Benefits of technology

The data labeling volume is reduced, the accuracy and generalization of the flood detection model are improved, the negative impact of noise pseudo-labels is reduced, and the robustness and generalization ability of the model are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356034A_ABST
    Figure CN120356034A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, and relates to a waterlogging detection large model training method, a training device, a detection method and equipment. The waterlogging detection large model training method comprises the steps that a training data set is acquired, and the training data set comprises an annotated data set and an unannotated data set; constructing a network structure of the large waterlogging detection model; training a network of a large waterlogging detection model by using the labeled data set and the unlabeled data set, obtaining prediction of a labeled waterlogging image by using the large waterlogging detection model during each training, and processing the unlabeled waterlogging image according to a preset rule to obtain a pseudo label of the unlabeled waterlogging image; and calculating a loss function of the large waterlogging detection model by using the prediction of the labeled waterlogging image and the pseudo label of the unlabeled waterlogging image, and optimizing network parameters of the large waterlogging detection model by using the loss function to obtain a final large waterlogging detection model. According to the invention, high-precision and low-cost pixel-level waterlogging detection can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and in particular, to a training method for a large model for waterlogging detection, a training device, a detection method, and a device. Background Art

[0002] Urban waterlogging detection is an important means to cope with modern urban flood disasters and an important part of smart city management. Current urban waterlogging detection technologies are mainly divided into two categories: those based on physical sensors (such as water level gauges) and those based on visual images. Image-based waterlogging detection methods can be mainly divided into methods based on object detection and methods based on semantic segmentation. The method based on object detection gives a detection box (Bounding Box) for the waterlogging area, while the method based on semantic segmentation can give a pixel-level segmentation mask for the waterlogging area, which belongs to a binary classification semantic segmentation task.

[0003] Current image-based urban waterlogging detection methods mainly rely on fully supervised semantic segmentation models, which require a large amount of labeled data and have high labeling costs. Existing semi-supervised algorithms (such as UniMatch V2) have problems such as insufficient consistency regularization constraints and too high pseudo-label confidence thresholds, resulting in low data utilization. In addition, although vision foundation models (such as SAM2, Segment Anything in Images and Videos) have general segmentation capabilities, they lack domain knowledge for waterlogging scenarios. Traditional adapter fine-tuning methods (such as SAM2-Adapter) only use fully supervised learning, which limits the model performance. Summary of the Invention

[0004] This application aims to at least solve the technical problems existing in the prior art, and provides a training method for a large model for waterlogging detection, a training device, a detection method, and a device.

[0005] In a first aspect, a training method for a large model for waterlogging detection provided by the present invention includes:

[0006] Obtain a training data set, which includes a labeled data set and an unlabeled data set; the labeled data set includes labeled waterlogging images and mask labels corresponding to the labeled waterlogging images, and the unlabeled data set includes unlabeled waterlogging images;

[0007] Construct the network structure of the large model for waterlogging detection;

[0008] Train the network of the large-scale waterlogging detection model using the labeled dataset and the unlabeled dataset. During each training, use the large-scale waterlogging detection model to obtain the predictions of the labeled waterlogging images, and process the unlabeled waterlogging images according to the preset rules to obtain the pseudo-labels of the unlabeled waterlogging images; calculate the loss function of the large-scale waterlogging detection model using the predictions of the labeled waterlogging images and the pseudo-labels of the unlabeled waterlogging images, and use the loss function to optimize the network parameters of the large-scale waterlogging detection model to obtain the final large-scale waterlogging detection model.

[0009] In a second aspect, the present invention provides a large-scale waterlogging detection model training device, including:

[0010] An acquisition module, configured to acquire a training dataset, where the training dataset includes a labeled dataset and an unlabeled dataset; the labeled dataset includes labeled waterlogging images and mask labels corresponding to the labeled waterlogging images, and the unlabeled dataset includes unlabeled waterlogging images;

[0011] A construction module, configured to construct the network structure of the large-scale waterlogging detection model;

[0012] A training module, configured to train the network of the large-scale waterlogging detection model using the labeled dataset and the unlabeled dataset. During each training, use the large-scale waterlogging detection model to obtain the predictions of the labeled waterlogging images, and process the unlabeled waterlogging images according to the preset rules to obtain the pseudo-labels of the unlabeled waterlogging images; calculate the loss function of the large-scale waterlogging detection model using the predictions of the labeled waterlogging images and the pseudo-labels of the unlabeled waterlogging images, and use the loss function to optimize the network parameters of the large-scale waterlogging detection model to obtain the final large-scale waterlogging detection model.

[0013] In a third aspect, the present invention provides a waterlogging detection method, where the method includes:

[0014] Acquire the waterlogging image to be processed;

[0015] Input the waterlogging image into the final large-scale waterlogging detection model trained by using the above large-scale waterlogging detection model training method to obtain the waterlogging detection result.

[0016] In a fourth aspect, the present invention provides an electronic device, where the electronic device includes:

[0017] At least one processor; and,

[0018] A memory communicatively connected to the at least one processor; where,

[0019] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the above large-scale waterlogging detection model training method.

[0020] In summary, the present application includes the following beneficial technical effects:

[0021] Data annotation is a labor-intensive task that requires a large amount of human and time costs. The present application can make full use of limited labeled datasets and easily accessible unlabeled datasets to train a large-scale urban waterlogging detection model, which can reduce the workload of staff while ensuring the accuracy of the large-scale urban waterlogging detection model. Before calculating the loss function using pseudo-labels, the pseudo-labels are screened using preset rules, and the pseudo-label pixels that do not meet the threshold conditions do not participate in model training. Only the pseudo-labels that meet the conditions are used to update the network parameters of the large-scale urban waterlogging detection model, which can reduce the negative effects of noisy pseudo-labels and improve the accuracy of the large-scale urban waterlogging detection model.

[0022] The present application explicitly constrains the prediction results (the second prediction result and the third prediction result) of two strong perturbation branches (the branches corresponding to the first strongly augmented image and the second strongly augmented image), constrains the consistency of the prediction results of the two strong perturbation branches, promotes the large-scale urban waterlogging detection model to learn and extract discriminative features invariant to perturbations, and helps to improve the generalization and robustness of the large-scale urban waterlogging detection model. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 It is a schematic flowchart of a method for training a large-scale urban waterlogging detection model provided by an embodiment of the present invention;

[0024] Figure 2 It is a network architecture diagram of a large-scale urban waterlogging detection model provided by an embodiment of the present invention;

[0025] Figure 3 It is an overall framework diagram of a semi-supervised algorithm provided by an embodiment of the present invention;

[0026] Figure 4 It is a comparison diagram of the waterlogging area segmentation effects of a large-scale urban waterlogging detection model trained by the method for training a large-scale urban waterlogging detection model provided by an embodiment of the present invention and a large-scale urban waterlogging detection model trained by an existing training method for urban waterlogging images;

[0027] Figure 5 It is a schematic structural diagram of an electronic device for implementing the method for training a large-scale urban waterlogging detection model provided by an embodiment of the present invention.

[0028] Reference numerals: 10, processor; 11, memory; 12, communication bus; 13, communication interface.

[0029] The realization, functional characteristics, and advantages of the objectives of the present invention will be further described in conjunction with the embodiments with reference to the drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention.

[0031] In the description of the present invention, it should be understood that the orientation or positional relationships indicated by the terms "longitudinal", "lateral", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. are based on the orientation or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.

[0032] In the description of the present invention, unless otherwise specified and defined, it should be noted that the terms "installation", "connection", and "connection" should be understood in a broad sense. For example, it may be a mechanical connection or an electrical connection, or it may be the communication inside two elements. It may be directly connected or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.

[0033] Refer to Figure 1 As shown, it is a schematic flow chart of a method for training a large model for waterlogging detection provided by an embodiment of the present invention. In this embodiment, the method for training the large model for waterlogging detection includes:

[0034] S1. Obtain a training data set.

[0035] The training data set includes a number of images related to urban waterlogging scenarios. The training data set includes a labeled data set and an unlabeled data set; the labeled data set includes a number of labeled waterlogging images and the corresponding mask labels of the labeled waterlogging images, and the unlabeled data set includes a number of unlabeled waterlogging images. In this embodiment, the labeled data set is denoted as The unlabeled data set is denoted as Among them, represents the i-th labeled waterlogging image, represents the corresponding true value mask label; represents the i-th unlabeled waterlogging image. During training, each batch consists of B l labeled waterlogging images and B u unlabeled waterlogging images.

[0036] For the labeled dataset, existing datasets can be used (such as the UW-Bench dataset). The UW-Bench dataset contains 5,584 images. Among them, the v2 test set of UW-Bench includes 2,124 images, including 1,069 general samples and 1,055 hard cases. For the unlabeled dataset, it can be collected by oneself or other unlabeled datasets related to the urban waterlogging scenario can be selected, which is not limited in this embodiment. Data annotation is a labor-intensive task that requires a large amount of manpower and time costs. By using limited labeled data and a large amount of easily obtainable unlabeled data, the amount of data annotation can be reduced while ensuring the performance of the large waterlogging detection model.

[0037] S2. Construct the network structure of the large waterlogging detection model.

[0038] Refer to Figure 2 , in this embodiment, the framework of the large waterlogging detection model adopts the SAM2-Adapter model framework. The full English name of SAM2 is Segment Anything in Images and Videos, and its Chinese interpretation is the Segmentation of All Things Model 2. SAM2 includes an encoder, a decoder, and a prompt encoder. SAM 2 has strong zero-shot generalization ability and can accept points, boxes, and masks as prompts (used to indicate the objects to be segmented) to achieve the segmentation of the input image. Zero-shot generalization ability refers to the ability of the model to complete tasks only through given instructions or prompts on specific tasks or datasets that it has never seen before.

[0039] The SAM2-Adapter model framework is further designed based on SAM2. An adapter module is added to the encoder to inject task-specific information into the encoder to achieve the adaptation of the downstream tasks of the SAM2 architecture. In the SAM2-Adapter framework of this embodiment, the task-specific information adopted is the high-frequency information of the image. By extracting the high-frequency information of the input image, the task-specific information is injected into the large model (i.e., the vision foundation model) through several simple MLP (Multi-Layer Perceptron) modules to achieve the fine-tuning of the downstream tasks of the large model. High-frequency information can reflect texture details, edges, etc., which is convenient for the large waterlogging detection model to segment the image.

[0040] Refer to Figure 2 , the processing steps of the SAM2-Adapter model for the input image include:

[0041] S201. Process the input image using multiple adapters respectively. Each adapter extracts task-specific information (high-frequency information) and injects it into the SAM2 image encoder.

[0042] S202. Input the input image into the SAM2 Image Encoder. The SAM2 Image Encoder extracts the features of the input image, and fuses the features of the input image and the task-specific information extracted by multiple Adapters to obtain an image embedding.

[0043] S203. Input the image embedding into the SAM2 Mask Decoder to obtain the final segmentation mask.

[0044] It should be noted that in the overall framework of the SAM2-Adapter model, the Adapter and the SAM2 Mask Decoder parts are trainable, and the image encoder part is frozen.

[0045] S3. Use the labeled dataset and the unlabeled dataset to train the network of the large-scale model for waterlogging detection. In each training, use the large-scale model for waterlogging detection to obtain the prediction of the labeled waterlogging images, and process the unlabeled waterlogging images according to the preset rules to obtain the pseudo-labels of the unlabeled waterlogging images; calculate the loss function of the large-scale model for waterlogging detection using the prediction of the labeled waterlogging images and the pseudo-labels of the unlabeled waterlogging images, and use the loss function to optimize the network parameters of the large-scale model for waterlogging detection to obtain the final large-scale model for waterlogging detection.

[0046] The prediction of the labeled waterlogging images refers to the prediction result of the target objects existing in the labeled waterlogging images after the network of the large-scale model for waterlogging detection analyzes the labeled waterlogging images; the loss related to the labeled waterlogging images of the large-scale model for waterlogging detection can be determined by comparing the prediction of the labeled waterlogging images and the mask labels corresponding to the labeled waterlogging images.

[0047] Specifically, calculating the loss function of the large-scale model for waterlogging detection using the prediction of the labeled waterlogging images and the pseudo-labels of the unlabeled waterlogging images includes:

[0048] S31. Analyze the unlabeled waterlogging images to obtain the prediction results, and generate pseudo-labels according to the prediction results.

[0049] Specifically, the prediction results include multiple prediction categories. The steps of analyzing the unlabeled waterlogging images to obtain the prediction results and generating pseudo-labels according to the prediction results include:

[0050] S311. Process the unlabeled waterlogging images according to the first preset rule to obtain the weakly augmented images of the unlabeled waterlogging images.

[0051] Refer to Figure 3 , in this embodiment, obtain the corresponding weakly augmented image x through the weak augmentation method A w ​w , the weak augmentation methods include random resizing, random cropping, random horizontal flipping, etc., and the weak augmentation method can be determined according to actual requirements; the processing process of the weak augmentation method is expressed by the following formula:

[0052] x w = A w (x u )

[0053] S312. Process the unlabeled waterlogging images according to the second preset rule to obtain the first strongly augmented image and the second strongly augmented image of the unlabeled waterlogging images.

[0054] Referring to Figure 3 , in this embodiment, the corresponding first strongly augmented image s is obtained through the strong augmentation method A and the second strongly augmented image The strong augmentation methods include color jittering, grayscaling, gaussian blurring, etc., and the strong augmentation method can be determined according to actual requirements; the processing process of the strong augmentation method is expressed by the following formula:

[0055]

[0056] S313. Input the weakly augmented images into the teacher model, the teacher model outputs the first prediction result, and generate the first pseudo label according to the first prediction result.

[0057] Referring to Figure 3 , in this embodiment, the teacher model is obtained by using exponential moving average. Set the teacher parameter as θ t , the student model parameter as θ s , and the update method of the teacher model is: θ t ← γ × θ t + (1 - γ) × θ s ,

[0058] where γ is a hyperparameter, and γ is dynamically set as ← means updating the left side with the operation on the right side, that is, the new teacher parameter is calculated and updated from the relevant old parameters on the right side, and the left side of ← is updated by the right side of ←; "iter" means iteration, and the process of training a batch of data is an iteration. The process of obtaining the prediction of the weakly augmented images by using the teacher model can be expressed as This function is denoted as the teacher model and is used to obtain pseudo-labels. Specifically, it is obtained by using EMA (Exponential Moving Average), which has better performance than the student model. For unlabeled images, the prediction results of the teacher model are more reliable. The obtained prediction p w serves as the first pseudo-label. The first pseudo-label p of this embodiment w is post-processed using the argmax operator to become a one-hot label

[0059] S314. Extract the features of the first strongly augmented image to obtain the first intermediate feature, perform feature-level augmentation processing on the first intermediate feature to obtain the first target feature, and determine the second prediction result according to the first target feature.

[0060] S315. Generate the second pseudo-label according to the second prediction result.

[0061] S316. Extract the features of the second strongly augmented image to obtain the second intermediate feature, perform feature-level augmentation processing on the second intermediate feature to obtain the second target feature, and determine the third prediction result according to the second target feature.

[0062] S317. Generate the third pseudo-label according to the third prediction result.

[0063] Referring to Figure 3 , for a strongly augmented image x s , its prediction result p sf , the acquisition process can be expressed as:

[0064] e s = g(x s ), p sf = h(F(e s ))

[0065] where g(.) is the encoder, h(.) is the decoder, e s represents the extracted intermediate feature, F represents feature-level augmentation, that is, complementary channel Dropout. Specifically, for the features extracted from the first strongly augmented image we sample a binary dropout mask M with the same dimension as the first intermediate feature from a binomial distribution with a probability of 0.5, and the values of half of its channels are set to all 1s and the other half to 0s. Use the binary dropout mask M to perform complementary channel Dropout operations on the first intermediate feature and the second intermediate feature respectively. The specific operation expression is:

[0066]

[0067] Among them, ⊙ represents the Hadamard product, ← means that the value on the left is updated by the operation on the right, and × represents the multiplication operation. In this embodiment, the operation of multiplying by 2 is to ensure that the expectation of the processed feature is consistent with the normal feature.

[0068] The expression of the second prediction result is The expression of the third prediction result is h(.) is the decoder.

[0069] Using two strongly augmented images (the first strongly augmented image and the second strongly augmented image) can more fully explore the preset image-level perturbation space; in addition, constraining the two strongly augmented branches to approach the same weakly augmented branch can be regarded as minimizing the distance between these two strongly augmented branches, borrowing the idea of contrastive learning, and can learn more discriminative representations.

[0070] In addition, in this embodiment, only plays a supervisory role. When converting p w to , no gradient backpropagation is performed (i.e., it does not participate in the update of the loss function); the network parameters of the waterlogging detection large model are updated using the second pseudo-label and the third pseudo-label.

[0071] S32. Determine whether the pseudo-label corresponding to the prediction result meets the preset conditions according to the confidence of the predicted category in the prediction result, and obtain the pseudo-labels that meet the preset conditions.

[0072] The screening steps for the pseudo-labels that meet the preset conditions include:

[0073] S321. Obtain the preset value of the confidence.

[0074] S322. Compare the maximum confidence of the predicted category in the prediction result with the preset value of the confidence:

[0075] If the maximum confidence of the predicted category in the prediction result is greater than or equal to the preset value of the confidence, the pseudo-label of the pixel corresponding to the preset result is put into the supervision set, and the pseudo-labels in the supervision set are the pseudo-labels that meet the preset conditions;

[0076] If the maximum confidence of the predicted category in the prediction result is less than the preset value of the confidence, the loss function will not be updated using this pseudo-label.

[0077] S33. Calculate the loss function of the waterlogging detection large model according to the pseudo-labels that meet the preset conditions.

[0078] The loss function of the large-scale model for waterlogging detection includes the loss of labeled waterlogging images and the loss of unlabeled waterlogging images. The loss of unlabeled waterlogging images includes the first supervision loss and the second supervision loss. The first supervision loss is used to quantify the error between the weakly augmented images and the first strongly augmented images, and to quantify the error between the weakly augmented images and the second strongly augmented images. The second supervision loss is used to quantify the error between the first strongly augmented images and the second strongly augmented images.

[0079] Denote the loss function of the large-scale model for waterlogging detection as L, and the expression of the loss function of the large-scale model for waterlogging detection is:

[0080] L = L l + λL u ,

[0081] L l is the loss of labeled waterlogging images, L u is the loss of unlabeled waterlogging images, and λ is a parameter for balancing the effect of unlabeled data.

[0082] The expression of the loss of labeled waterlogging images is:

[0083]

[0084] where B l represents the total number of labeled waterlogging images, is the prediction result of the network of the large-scale model for waterlogging detection for the i-th labeled waterlogging image, is the ground truth mask corresponding to the i-th labeled waterlogging image. H(.) represents the cross-entropy loss.

[0085] In this embodiment, the expression of the loss of unlabeled waterlogging images is

[0086]

[0087] where B u represents the total number of unlabeled waterlogging images, i represents the index of the unlabeled waterlogging image, represents the second prediction result, represents the third prediction result, represents the first pseudo-label, represents the second pseudo-label, represents the third pseudo-label, and H(.) represents the cross-entropy loss.

[0088] represents the maximum confidence in the predicted class corresponding to the first prediction result; τ represents the first confidence preset value, τ s represents the second confidence preset value, and Ⅱ(.) takes the value of 1 when the judgment condition in the parentheses is satisfied; The operator is used to obtain a binary mask of 0 or 1. The positions of the pixels that meet the conditions in the parentheses are assigned 1, otherwise 0, so as to obtain a binary mask. Specifically, the pixels that meet the conditions in the parentheses participate in the loss calculation, and those that do not meet the conditions do not participate in the loss calculation. The Ⅱ(.) operator can avoid the negative effects of noise pseudo-labels. It predefines a confidence threshold (τ / τ s ), and the pseudo-label pixels that do not meet the threshold conditions do not participate in the training (i.e., do not calculate the loss). The specific meaning of the Ⅱ(.) operator is that if the maximum confidence in all categories is greater than the preset threshold, it is multiplied by 1, that is, it participates in the loss calculation of the unlabeled waterlogging images. If the maximum confidence in all categories is less than the preset threshold, it is multiplied by 0, that is, it does not participate in the loss calculation. The value of τ / τ s can be set according to the actual situation, and this application does not make any restrictions.

[0089] This application explicitly constrains the prediction results of the two strong perturbation branches (the second prediction result and the third prediction result), narrows the distance between the two strong perturbation branches, and promotes the large-scale waterlogging detection model to learn and extract discriminative features that are invariant to perturbations, which helps to improve the generalization and robustness of the large-scale waterlogging detection model.

[0090] Refer to Table 1 and Figure 4 , Table 1 is a comparison table of the performance parameters of the large-scale waterlogging detection model trained by the large-scale waterlogging detection model training method of this application and the performance parameters of the model trained by the existing semi-supervised learning algorithm UniMatch V2;

[0091] Figure 4 is a comparison chart of the segmentation effect of the large-scale waterlogging detection model trained by the large-scale waterlogging detection model training method of this application on urban waterlogging images and the segmentation effect of the model trained by the existing semi-supervised learning algorithm UniMatch V2 on urban waterlogging images;

[0092] Table 1 Performance comparison table of the large-scale waterlogging detection model training method of this application and the UniMatch V2 training method

[0093]

[0094]

[0095] It can be seen that the large-scale waterlogging detection model training method corresponding to this application has better performance than UniMatch V2 and can achieve better IoU (i.e., intersection over union) and F1-score (i.e., F1 score). It can be seen that the large-scale waterlogging detection model training method of this application can bring performance gains.

[0096] Based on the same inventive concept, an embodiment of the present invention provides a large-scale waterlogging detection model training device.

[0097] The training device for the large-scale model for waterlogging detection according to the present invention can be installed in an electronic device. According to the functions achieved, the training device for the large-scale model for waterlogging detection includes an acquisition module, a construction module, and a training module. The acquisition module can acquire a labeled data set and an unlabeled data set as the training data set for the large-scale model for waterlogging detection. The labeled data set includes labeled waterlogging images and mask labels corresponding to the labeled waterlogging images. The unlabeled data set includes unlabeled waterlogging images. The construction module can construct the network structure of the large-scale model for waterlogging detection. The training module can use the labeled data set and the unlabeled data set to train the network of the large-scale model for waterlogging detection. During each training, the large-scale model for waterlogging detection is used to obtain predictions of the labeled waterlogging images, and the unlabeled waterlogging images are processed according to a preset rule to obtain pseudo-labels of the unlabeled waterlogging images. The loss function of the large-scale model for waterlogging detection is calculated using the predictions of the labeled waterlogging images and the pseudo-labels of the unlabeled waterlogging images, and the network parameters of the large-scale model for waterlogging detection are optimized using the loss function to obtain the final large-scale model for waterlogging detection.

[0098] The module described in the present invention can also be referred to as a unit, which refers to a series of computer program segments that can be executed by a processor of an electronic device and can complete fixed functions, and are stored in the memory of the electronic device.

[0099] The various variation methods and specific examples in the waterlogging detection large model training method provided in the above embodiments are equally applicable to the waterlogging detection system in this embodiment. Through the detailed description of the waterlogging detection large model training method above, those skilled in the art can clearly know the implementation method of the waterlogging detection system in this embodiment. For the sake of brevity of the specification, it will not be described in detail here.

[0100] Based on the same inventive concept, the present invention also provides a waterlogging detection method, which includes:

[0101] S101. Obtain the waterlogging image to be processed.

[0102] S102. Input the waterlogging image into the final large-scale model for waterlogging detection trained by the above waterlogging detection large model training method to obtain the waterlogging detection result.

[0103] The present application also discloses an electronic device, as Figure 5 shown, which is a schematic structural diagram of an electronic device for the waterlogging detection large model training method or / or the waterlogging detection method provided in an embodiment of the present invention. The electronic device may include at least one processor 10, a memory 11 communicatively connected to the at least one processor, a communication bus 12, and a communication interface 13. It may also include a computer program stored in the memory 11 and executable on the processor 10, such as a waterlogging detection large model training method program.

[0104] Among them, the processor 10 may be composed of an integrated circuit in some embodiments. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple packaged integrated circuits with the same or different functions, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 10 is the control core (Control Unit) of the electronic device, connecting various components of the entire electronic device through various interfaces and circuits, and by running or executing programs or modules stored in the memory 11 (such as executing the training method of the large model for waterlogging detection, etc.), and calling the data stored in the memory 11, to perform various functions of the electronic device and process data.

[0105] The memory 11 includes at least one type of readable storage medium. The readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as: SD or DX memory, etc.), magnetic memory, magnetic disk, optical disc, etc. The memory 11 may be an internal storage unit of the electronic device in some embodiments, such as the mobile hard disk of the electronic device. The memory 11 may also be an external storage device of the electronic device in other embodiments, such as a plug-in mobile hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. equipped on the electronic device. Further, the memory 11 may also include both the internal storage unit and the external storage device of the electronic device. The memory 11 can not only be used to store application software installed on the electronic device and various types of data, such as the code of the training method program of the large model for waterlogging detection, etc., but also be used to temporarily store the data that has been output or will be output.

[0106] The communication bus 12 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. The bus is set to realize the connection and communication between the memory 11 and at least one processor 10, etc.

[0107] The communication interface 13 is used for communication between the above-mentioned electronic device and other devices, including a network interface and a user interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is generally used to establish a communication connection between this electronic device and other electronic devices. The user interface may be a display, an input unit (such as a keyboard), and optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, which is used to display the information processed in the electronic device and to display a visual user interface.

[0108] Figure 5 Only the electronic device with components is shown. Those skilled in the art can understand that Figure 5 The shown structure does not constitute a limitation on the electronic device. It may include fewer or more components than shown, or combine certain components, or have a different component arrangement. For example, although not shown, the electronic device may further include a power source (such as a battery) for supplying power to each component. Preferably, the power source may be logically connected to at least one processor 10 through a power management device, so as to implement functions such as charge management, discharge management, and power consumption management through the power management device. The power source may also include any components such as one or more DC or AC power sources, a recharge device, a power failure detection circuit, a power converter or an inverter, and a power status indicator. The electronic device may also include various sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.

[0109] It should be understood that the embodiments are only for illustration purposes and are not limited by this structure in the scope of the patent application.

[0110] Furthermore, if the integrated module / unit of the electronic device is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. The computer-readable storage medium may be volatile or non-volatile.

[0111] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", "one implementation manner", "one preferred implementation manner" or "some examples", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0112] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the claims and their equivalents.

Claims

1. A training method for an urban waterlogging detection large model, characterized in that The method includes: Obtaining a training data set, which includes a labeled data set and an unlabeled data set; the labeled data set includes labeled waterlogging images and mask labels corresponding to the labeled waterlogging images, and the unlabeled data set includes unlabeled waterlogging images; Constructing the network structure of the large waterlogging detection model; Training the network of the large waterlogging detection model using the labeled data set and the unlabeled data set. In each training, obtaining the prediction of the labeled waterlogging image using the large waterlogging detection model, and processing the unlabeled waterlogging image according to a preset rule to obtain the pseudo-label of the unlabeled waterlogging image; calculating the loss function of the large waterlogging detection model using the prediction of the labeled waterlogging image and the pseudo-label of the unlabeled waterlogging image, and optimizing the network parameters of the large waterlogging detection model using the loss function to obtain the final large waterlogging detection model.

2. The training method of an urban waterlogging detection large model according to claim 1, characterized in that, The calculating the loss function of the large waterlogging detection model using the prediction of the labeled waterlogging image and the pseudo-label of the unlabeled waterlogging image includes: Parsing the unlabeled waterlogging image to obtain a prediction result, and generating a pseudo-label according to the prediction result. The prediction result includes multiple prediction categories; Judging whether the pseudo-label corresponding to the prediction result meets a preset condition according to the confidence of the prediction category in the prediction result, and obtaining the pseudo-label that meets the preset condition; Calculating the loss function of the large waterlogging detection model according to the pseudo-label that meets the preset condition.

3. The training method of the large model for waterlogging detection according to claim 2, characterized in that, The screening step of the pseudo-label that meets the preset condition includes: Obtaining a preset confidence value; Comparing the maximum confidence of the prediction category in the prediction result with the preset confidence value: If the maximum confidence of the prediction category in the prediction result is greater than or equal to the preset confidence value, putting the pseudo-label of the pixel corresponding to the preset result into the supervision set, and the pseudo-labels in the supervision set are the pseudo-labels that meet the preset condition.

4. A method for training a large model for waterlogging detection according to claim 1, characterized in that, The processing the unlabeled waterlogging image according to a preset rule to obtain the pseudo-label of the unlabeled waterlogging image includes: Processing the unlabeled waterlogging image according to a first preset rule to obtain a weakly augmented image of the unlabeled waterlogging image; Processing the unlabeled waterlogging image according to a second preset rule to obtain a first strongly augmented image and a second strongly augmented image of the unlabeled waterlogging image; Inputting the weakly augmented image into a teacher model, the teacher model outputs a first prediction result, and generating a first pseudo-label according to the first prediction result; Extracting the features of the first strongly augmented image to obtain a first intermediate feature, performing feature-level augmentation processing on the first intermediate feature to obtain a first target feature, and determining a second prediction result according to the first target feature; Generating a second pseudo-label according to the second prediction result; Extracting the features of the second strongly augmented image to obtain a second intermediate feature, performing feature-level augmentation processing on the second intermediate feature to obtain a second target feature, and determining a third prediction result according to the second target feature; Generating a third pseudo-label according to the third prediction result.

5. The training method of the large model for waterlogging detection according to claim 4, characterized in that The loss function of the large-scale model for waterlogging detection includes the loss of labeled waterlogging images and the loss of unlabeled waterlogging images. The loss of unlabeled waterlogging images includes the first supervision loss and the second supervision loss. The first supervision loss is used to quantify the error between the weakly augmented images and the first strongly augmented images, and to quantify the error between the weakly augmented images and the second strongly augmented images. The second supervision loss is used to quantify the error between the first strongly augmented images and the second strongly augmented images.

6. The training method of the large model for waterlogging detection according to claim 5, wherein, The expression of the loss of unlabeled waterlogging images is Among them, B u represents the number of unlabeled waterlogging images, and i represents the index of unlabeled waterlogging images. represents the maximum confidence in the predicted category corresponding to the first prediction result; τ represents the first confidence preset value, τ s represents the second confidence preset value, and Ⅱ(.) takes the value of 1 when the judgment condition in the parentheses is satisfied. represents the second prediction result, represents the third prediction result, represents the first pseudo-label, represents the second pseudo-label, represents the third pseudo-label, and H(.) represents the cross-entropy loss.

7. The training method of an urban waterlogging detection large model according to claim 3, characterized in that Update the network parameters of the large-scale model for waterlogging detection using the second pseudo-label and the third pseudo-label.

8. An apparatus for training a large model for waterlogging detection, characterized in that, including: An acquisition module, configured to acquire a training data set, where the training data set includes a labeled data set and an unlabeled data set; The labeled data set includes labeled waterlogging images and mask labels corresponding to the labeled waterlogging images, and the unlabeled data set includes unlabeled waterlogging images; A construction module, configured to construct the network structure of the large-scale model for waterlogging detection; A training module, configured to train the network of the large-scale model for waterlogging detection using the labeled data set and the unlabeled data set. During each training, use the large-scale model for waterlogging detection to obtain predictions of the labeled waterlogging images, and process the unlabeled waterlogging images according to a preset rule to obtain pseudo-labels of the unlabeled waterlogging images. Calculate the loss function of the large-scale model for waterlogging detection using the predictions of the labeled waterlogging images and the pseudo-labels of the unlabeled waterlogging images, and optimize the network parameters of the large-scale model for waterlogging detection using the loss function to obtain the final large-scale model for waterlogging detection.

9. A method for detecting waterlogging, characterized in that, including: Obtain the waterlogging image to be processed; Input the waterlogging image into the final large-scale model for waterlogging detection trained by the method for training a large-scale model for waterlogging detection according to any one of claims 1 to 7 to obtain a waterlogging detection result.

10. An electronic device, characterized in that, The electronic device includes: At least one processor (10); and, A memory (11) communicatively connected to the at least one processor (10); Wherein, the memory (11) stores a computer program executable by the at least one processor (10), and the computer program is executed by the at least one processor (10) so that the at least one processor (10) can execute the method according to claim 1 or 2 or 3 or 4 or 5 or 6 or 7 or 9.

Citation Information

Patent Citations

  • Model training method and device, image processing method and device, equipment and storage medium

    CN113869449A

  • Image processing method and processor

    CN115019218A

  • Semi-supervised semantic segmentation method guided by high-density representative prototype

    CN117437426A

  • Colposcope image segmentation model construction method, image classification method and device

    CN118334336A

  • Emotion recognition intelligent contract construction method based on cross fusion and confidence evaluation

    CN118799948A