River sewage draining exit detection and traceability method and system

By integrating multi-source data to construct an integrated model for sewage outlet detection and source tracing, the problems of low detection efficiency and difficulty in identifying pollution sources in existing technologies have been solved. This model enables efficient and accurate identification of the spatial location of sewage outlets and pollution sources, thereby improving the precision of investigation and source tracing.

CN122024005APending Publication Date: 2026-05-12NORTHWEST NORMAL UNIVERSITY
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NORTHWEST NORMAL UNIVERSITY
Filing Date
2026-01-06
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing technologies are inefficient and labor-intensive in detecting sewage outlets, and they are difficult to identify hidden sewage outlets and their pollution sources in complex environments, thus failing to meet the needs of refined supervision.

Method used

By integrating UAV remote sensing imagery, cross-sectional water quality monitoring data, and shoreline environmental text data, an integrated model for sewage outlet detection and source tracing is constructed. A composite objective function is used for joint optimization training, and the prediction results are combined with environmental physical rules to achieve collaborative identification of the spatial location of sewage outlets and the attributes of pollution sources.

Benefits of technology

In complex riverbank environments, it improves the detection accuracy of concealed sewage outlets and the reliability of pollution source identification, ensuring detection precision and pollution source tracing accuracy, and providing technical support for refined investigation and source tracing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122024005A_ABST
    Figure CN122024005A_ABST
Patent Text Reader

Abstract

The invention discloses a river sewage draining exit detection and tracing method and system. The method comprises the following steps: acquiring remote sensing visual features, pollution source features and environmental semantic features of a to-be-monitored river reach; a sewage draining exit detection and traceability integrated model is constructed, the model comprises a sewage draining exit detection network and a multi-mode traceability network, the sewage draining exit detection network is used for outputting a sewage draining exit space position prediction result, and the multi-mode traceability network is used for outputting a pollution source attribute discrimination result; in the model training process, remote sensing visual features, pollution source features and environment semantic features are input into a model, and the model is trained in combination with a composite objective function; and after model training is completed, inputting the current visible light orthoimage of the to-be-predicted area into the trained model, and outputting a current sewage draining exit space position prediction result and a current pollution source attribute discrimination result. The method can improve the detection accuracy of the hidden river sewage draining exit and the reliability of pollution source discrimination, and can be widely applied to the technical field of environment monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of environmental monitoring technology, and in particular to a method and system for detecting and tracing sewage outlets into rivers. Background Technology

[0002] Sewage outfalls are key points where pollutants enter river waters, and their control directly affects the water environment quality and ecological security of the basin. Currently, the investigation of sewage outfalls still relies mainly on manual on-site inspections. Faced with challenges such as complex shoreline topography and the concealed distribution of sewage outfalls, the traditional manual, grid-like inspection method is inefficient, labor-intensive, and prone to omissions due to environmental and subjective factors, making it difficult to meet the needs of refined supervision.

[0003] With the development of remote sensing and intelligent identification technologies, automated sewage outlet investigation based on remote sensing imagery has been widely adopted. However, these technologies largely rely on visual features in remote sensing imagery to locate and identify sewage outlets. When sewage outlets are obscured by vegetation, located underwater, or discharged through concealed pipes, methods based solely on visual features are difficult to identify effectively. Furthermore, these models typically treat sewage outlets as single targets, focusing primarily on their existence and spatial location, which limits their application in identifying pollution source attributes and accurately tracing pollution sources.

[0004] In summary, the relevant sewage outlet detection technologies still have shortcomings and urgently need improvement. Summary of the Invention

[0005] The embodiments of this application aim to at least partially solve one of the technical problems in the related art. Therefore, the main objective of the embodiments of this application is to propose a method and system for detecting and tracing the source of sewage discharge outlets into rivers, which can improve the speed and accuracy of determining the pollution source attributes of sewage discharge outlets into rivers, and effectively enhance the ability to identify concealed sewage discharge outlets.

[0006] To achieve the above objectives, this application proposes a method for detecting and tracing the source of sewage discharge outlets into rivers, the method comprising the following steps: Acquire UAV remote sensing images of the river section to be monitored, and perform a first feature processing on the UAV remote sensing images to obtain remote sensing visual features. Acquire cross-sectional water quality monitoring data that is synchronized with the UAV remote sensing image in time and space, and perform a second feature processing on the cross-sectional water quality monitoring data to obtain pollution source characteristics; The shoreline environmental text data of the river section to be monitored is obtained, and the shoreline environmental text data is subjected to a third feature processing to obtain environmental semantic features. An integrated model for sewage outlet detection and source tracing is constructed; wherein, the integrated model for sewage outlet detection and source tracing includes a sewage outlet detection network and a multimodal source tracing network, the sewage outlet detection network is used to output the spatial location prediction result of the sewage outlet, and the multimodal source tracing network is used to output the pollution source attribute discrimination result; A composite objective function is constructed, and the integrated model for detecting and tracing sewage outlets is jointly optimized and trained based on the composite objective function, the remote sensing visual features, the pollution source features, and the environmental semantic features until the preset training requirements are met, thus obtaining the trained integrated model for detecting and tracing sewage outlets. The current visible light orthophoto of the area to be predicted is input into the integrated model for detecting and tracing sewage outlets after training, and the current sewage outlet prediction result is output. The current sewage outlet prediction result includes the current sewage outlet spatial location prediction result and the current pollution source attribute discrimination result.

[0007] To achieve the above objectives, another aspect of this application proposes a system for detecting and tracing sewage outfalls into rivers, the system comprising the following modules: The remote sensing visual feature generation module is used to acquire UAV remote sensing images of the river section to be monitored, and to perform a first feature processing on the UAV remote sensing images to obtain remote sensing visual features. The pollution source feature generation module is used to acquire cross-sectional water quality monitoring data that is synchronized with the UAV remote sensing image in time and space, and to perform a second feature processing on the cross-sectional water quality monitoring data to obtain pollution source features. An environmental semantic feature generation module is used to acquire the shoreline environmental text data of the river section to be monitored, and to perform a third featureization process on the shoreline environmental text data to obtain environmental semantic features. The model building module is used to construct an integrated model for sewage outlet detection and source tracing; wherein, the integrated model for sewage outlet detection and source tracing includes a sewage outlet detection network and a multimodal source tracing network, the sewage outlet detection network is used to output the spatial location prediction results of sewage outlets, and the multimodal source tracing network is used to output the pollution source attribute discrimination results; The model training module is used to construct a composite objective function and perform joint optimization training on the integrated model of sewage outlet detection and source tracing based on the composite objective function, the remote sensing visual features, the pollution source features, and the environmental semantic features until the preset training requirements are met, and the trained integrated model of sewage outlet detection and source tracing is obtained. The sewage outlet detection and tracing module is used to input the current visible light orthophoto of the area to be predicted into the integrated sewage outlet detection and tracing model that has been trained, and output the current sewage outlet prediction result; the current sewage outlet prediction result includes the current sewage outlet spatial location prediction result and the current pollution source attribute discrimination result.

[0008] The embodiments of this application include at least the following beneficial effects: This application provides a method and system for detecting and tracing sewage outfalls into rivers. This scheme integrates UAV remote sensing visual features, cross-sectional water quality monitoring features, and shoreline environmental semantic features to achieve collaborative discrimination of the spatial location and pollution source attributes of sewage outfalls into rivers. Compared with related technologies that rely solely on remote sensing visual information, the embodiments of this application can more comprehensively characterize the spatial morphological features, pollution emission characteristics, and environmental background information of sewage outfalls in complex riverside environments, thereby effectively improving the detection accuracy of concealed sewage outfalls and the reliability of pollution source discrimination. In addition, the embodiments of this application introduce a composite objective function in the model training stage to jointly optimize the sewage outfall detection task and the pollution source discrimination task, and combine environmental physical rules to constrain the consistency of model prediction results, effectively suppressing the irrational predictions that may be generated by deep learning models and enhancing the logical rationality and stability of prediction results. Thus, the embodiments of this application can ensure detection accuracy while improving the accuracy and robustness of pollution source discrimination when relying only on visible light orthophoto image input in the prediction stage, providing reliable technical support for the refined investigation and pollution source discrimination of sewage outfalls into rivers. Attached Figure Description

[0009] Figure 1 This is a flowchart illustrating the steps of a method for detecting and tracing sewage outlets into rivers, as provided in an embodiment of this application. Figure 2 This is a flowchart illustrating a method for detecting and tracing sewage outlets into rivers, as provided in an embodiment of this application. Figure 3 This is a visual sample diagram of a sewage outlet into a river provided in an embodiment of this application; Figure 4 This is a schematic diagram of the detection and source tracing results of sewage outlets into rivers during the rainy season in a typical watershed, provided in an embodiment of this application. Figure 5 This is a schematic diagram of the detection and source tracing results of sewage outlets into rivers during the dry season in a typical watershed, provided in an embodiment of this application. Figure 6 This is a schematic diagram of the structure of a sewage outlet detection and tracing system provided in an embodiment of this application; Figure 7 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0010] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of systems and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.

[0011] It is understood that the terms “first,” “second,” etc., used in this application may be used herein to describe various concepts, but unless otherwise stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the words “if,” “when,” or “in response to a determination” as used herein may be interpreted as “when…” or “when…” or “in response to a determination.”

[0012] As used in this application, the terms "at least one", "multiple", "each", "any", etc., "at least one" includes one, two or more, "multiple" includes two or more, "each" refers to each of the corresponding multiples, and "any" refers to any one of the multiples.

[0013] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0014] Before providing a detailed description of the embodiments of this application, some of the nouns and terms involved in the embodiments of this application will be explained first. The nouns and terms involved in the embodiments of this application are subject to the following interpretations.

[0015] (1) Digital Orthophoto Map (DOM) refers to a two-dimensional image product with a uniform scale and true ground projection relationship generated by geometric correction, radiometric correction and topographic correction of visible light images. DOM eliminates the tilt and distortion in photogrammetry and can be used as an accurate geographic base map to extract the morphology, texture and spatial location features of sewage outlets.

[0016] (2) Digital Surface Model (DSM) refers to a rasterized model that provides a three-dimensional description of the elevation of ground features (including vegetation, buildings, structures, and exposed ground features), representing the height field of the highest point on the ground. DSM can be used to analyze the topographic relief, aspect, and bank structure around sewage outlets, providing elevation constraints for spatial geometric interpretation and river outlet morphology identification.

[0017] (3) Thermal Infrared Imagery (TIR) ​​refers to the radiation brightness temperature image of ground objects acquired using the thermal infrared band (generally 8–14 μm). TIR can characterize the temperature field distribution of the ground surface or water body and can be used to identify abnormal thermal features around sewage outlets, such as the temperature difference between polluted water bodies and background water bodies, and the diffusion pattern of thermal plumes.

[0018] (4) Multilayer Perceptron (MLP) is a typical feedforward neural network structure consisting of an input layer, one or more hidden layers, and an output layer. It uses a nonlinear activation function to map and abstract the feature space. MLP is often used for fusion encoding of spatial-spectral features, pollution intensity vectors, etc., or for classification and regression tasks.

[0019] (5) Online Hard Example Mining (OHEM) refers to dynamically selecting high-loss samples (i.e., hard examples) during model training and adding them to gradient updates to improve the model's ability to discriminate complex scenarios such as blurred boundaries, large scale variations, and weak feature discharge outlets. OHEM can improve the robustness and recognition accuracy of the model in remote sensing monitoring.

[0020] (6) Cross-Attention Mechanism (CAM) refers to a class of attention operators that achieve information alignment and complementary enhancement by explicitly modeling the cross-feature dependency relationship between query features and key-value features when they come from different sources or are in different modalities.

[0021] (7) (DEtection Transformer, abbreviated as DETR) is an end-to-end object detection framework based on Transformer. Its core utilizes the cross attention mechanism to perform object query localization and category inference on the global feature map, thereby avoiding anchor box design and post-processing steps.

[0022] (8) (Bidirectional Encoder Representations from Transformers, abbreviated as BERT) is a bidirectional language representation model based on Transformer encoders. It can be pre-trained on large-scale corpora and provides context-sensitive text feature representations for downstream tasks.

[0023] This application provides a method for detecting and tracing sewage discharge outlets into rivers, relating to the field of environmental monitoring technology. This method can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, or vehicle-mounted terminal, but is not limited to these. The server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network. The software can be an application implementing the method for detecting and tracing sewage discharge outlets into rivers, but is not limited to the above forms.

[0024] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0025] Please see Figure 1 , Figure 1 This is an optional flowchart of a method for detecting and tracing sewage outlets into rivers, provided in an embodiment of this application. Figure 1 The method may include, but is not limited to, steps S101 to S106.

[0026] Step S101: Acquire UAV remote sensing images of the river section to be monitored, and perform a first feature processing on the UAV remote sensing images to obtain remote sensing visual features. In some embodiments, step S101 may include: acquiring UAV remote sensing images of the river section to be monitored; wherein the UAV remote sensing images include visible light orthophotos, digital surface model data, and thermal infrared images; performing geometric correction, spatial registration, and scale unification processing on the UAV remote sensing images to obtain spatially aligned multi-source remote sensing data; inputting the spatially aligned multi-source remote sensing data into a feature extraction network for joint feature extraction to generate remote sensing visual features for characterizing the spatial morphology and spatial distribution characteristics of sewage outlets into the river.

[0027] Among them, UAV remote sensing images may include, but are not limited to, visible light orthophotos (DOM), thermal infrared remote sensing images (TIR), digital surface models (DSM), etc.

[0028] In the specific implementation, firstly, UAV remote sensing images of the river section to be monitored are acquired; then, after preprocessing such as geometric correction, spatial registration, and cropping of the UAV remote sensing images, they are input into the deep learning visual feature extraction network ResNet-50 to extract the spatial-spectral interpretation feature matrix representing the physical morphology and spatial distribution characteristics of sewage outlets into the river. The spatial-spectral interpretation feature matrix is ​​then used as the remote sensing visual feature. The specific implementation process is as follows: steps S101a-S101b: Step S101a: Plan the UAV flight path and acquire visible light orthophoto (DOM), digital surface model (DSM), and thermal infrared (TIR) ​​images of the river section to be monitored.

[0029] Step S101b involves preprocessing the multimodal remote sensing data acquired in step S101a. This includes using ground control points to perform geometric correction on the visible light and thermal infrared images, and completing orthorectification. Subsequently, the DOM, DSM, and TIR images are registered from multiple sources and uniformly resampled to the same spatial resolution (0.5 m / pixel) and geographic coordinate system (WGS84). Finally, the registered multimodal images are stacked into a multichannel image tensor. The data is then input into the ResNet-50 deep learning feature extraction network to extract the spatial-spectral interpretation feature matrix representing the physical morphology of sewage outfalls into rivers. (i.e., remote sensing visual features), where the calculation formula for feature extraction is as follows: ; In the formula, The input image tensor after preprocessing. Image row and column number; It is the sum of the number of visible light bands, the number of thermal infrared bands, and the number of elevation data channels, i.e. ; This represents the convolutional layer operation of a feature extraction network; For the extracted spatial-spectral interpretation feature matrix, This represents the size of the downsampled feature map. For feature channel dimensions.

[0030] Step S102: Obtain cross-sectional water quality monitoring data that is synchronized with the UAV remote sensing image in time and space, and perform a second feature processing on the cross-sectional water quality monitoring data to obtain pollution source characteristics; In some embodiments, step S102 may include: acquiring cross-sectional water quality monitoring data that is synchronized with UAV remote sensing images in time and space; performing data standardization processing on the cross-sectional water quality monitoring data to obtain a standardized water quality vector; and inputting the standardized water quality vector into a multilayer perceptron network for linear transformation and nonlinear activation function processing to obtain pollution source features in a high-dimensional feature space.

[0031] The water quality monitoring data at the cross-section may include, but is not limited to, multiple physicochemical indicators such as dissolved oxygen (DO), chemical oxygen demand (COD), ammonia nitrogen (NH3-N), total phosphorus (TP), flow rate (Q), and temperature (T).

[0032] In the specific implementation, firstly, cross-sectional water quality monitoring data synchronized in time and space with the UAV remote sensing image in step S101 is acquired; then, after standardizing the cross-sectional water quality monitoring data, it is input into a feature mapping network to map discrete physicochemical indicators into continuous water chemical pollution feature vectors, and the water chemical pollution feature vectors are used as pollution source features. The specific implementation process is as follows: steps S102a-S102b: Step S102a: First, collect cross-sectional water quality monitoring data that is synchronized in time and space with the multi-channel remote sensing imagery from step S101, including dissolved oxygen (DO), chemical oxygen demand (COD), ammonia nitrogen (NH3-N), total phosphorus (TP), flow rate (Q), and temperature (T). Each index is denoted as a vector. Then, Z-score normalization is used to eliminate the influence of dimensions. For the first... The first sample Item Indicators Its standardized value The calculation formula is as follows: ; In the formula, For the first Historical statistical mean of the indicator; For the first The historical statistical standard deviation of each indicator; after the above processing, a standardized water quality input vector is obtained. .

[0033] Step S102b, the above discrete standardized water quality input vector Inputting a 3-layer MLP (Multi-Layer Perceptron) network, and processing it through linear transformations and nonlinear activation functions in fully connected layers, yields a water chemical pollution feature vector in a high-dimensional feature space. The calculation formulas for the linear transformation and nonlinear activation function processing of fully connected layers are as follows: ; In the formula, The standardized water quality input vector has the following dimensions: , Here is the learnable weight matrix of the feature mapping network, with dimension 1. , For target feature dimensions; Let be the bias vector, with dimension . ; This represents matrix multiplication. It is a non-linear activation function used to enhance the non-linear expressive power of features; The output is a water chemical pollution feature vector with dimension 1. This vector can continuously and densely represent the current biochemical pollution status of the water body at the monitoring section, and can be used as a key-value item when calculating the pollution source classification loss.

[0034] Step S103: Obtain the shoreline environmental text data of the river section to be monitored, and perform third feature processing on the shoreline environmental text data to obtain environmental semantic features. In some embodiments, step S103 may include: determining the geographical extent of the river section to be monitored based on spatial metadata of UAV remote sensing images; obtaining riparian environmental text data corresponding to the river section to be monitored from a pre-constructed riparian geographical environment knowledge database based on the geographical extent; wherein, the riparian geographical environment knowledge database includes unstructured text data carrying riparian geographical tags, and the unstructured text data includes descriptions of riparian land use types, records of historical pollution sources, and information on environmentally sensitive receptors; concatenating the unstructured text data in the riparian environmental text data to obtain a target text sequence, and adding corresponding classification tags to the beginning and end of the target text sequence; mapping the target text sequence to an embedding vector matrix; performing hierarchical feature transformation on the embedding vector matrix through a bidirectional language representation model to obtain a context hidden layer state matrix; extracting the global text feature vector corresponding to the first classification tag in the context hidden layer state matrix, and performing feature space transformation on the global text feature vector to obtain environmental semantic features.

[0035] Among them, the riparian environmental text data may include, but is not limited to, river chief system management records, land use type descriptions, registration information of polluting enterprises along the riparian banks, historical pollution event reports, and hydrogeological survey summaries.

[0036] In the specific implementation, firstly, based on the monitoring range defined in step S1, unstructured text data of the riverbank geographic environment corresponding to that river section is obtained; then, the BERT language model is used to perform semantic understanding and vectorization representation of the text, generating structured geographic environment attribute feature vectors that characterize the attributes of the riverbank environment and the features of human activities, and using the structured geographic environment attribute feature vectors as environmental semantic features. The specific implementation process is as follows: steps S103a-S103c: Step S103a: First, read the spatial metadata of the multi-channel remote sensing image from step S101 to determine the geographic coordinate range of the river section to be monitored. Then, based on the geographical coordinate range... Constructing a knowledge database of coastal geography This database contains geographically tagged unstructured text data, covering descriptions of riparian land use types, historical pollution source records, and information on environmentally sensitive receptors. Spatial indexing algorithms are used, centered on geographic coordinates. Set a preset spatial buffer radius around the center. In the database Perform a spatial query operation to retrieve a set of unstructured text sequences that fall within the specified spatial range. The retrieved text sequence set is concatenated to construct the input text sequence. And add BERT classification tags to the beginning of the sequence. Add to the tail Input text sequence The calculation formula is as follows: ; In the formula, Represents the first in the text sequence Each token This is the preset maximum sequence length.

[0037] Step S103b, the text sequence Mapped to an embedding vector matrix ,in The sequence length contains and By using the multi-head self-attention mechanism within the Transformer encoder of the BERT model and performing hierarchical feature transformation with the feedforward neural network, the context hidden layer state matrix is ​​obtained. Context hidden layer state matrix The calculation formula is as follows: ; in, This represents the hidden layer dimension of the BERT language model.

[0038] Step S103c: Extract the hidden layer state matrix The feature vector corresponding to the first classification token ([CLS] Token) This vector is typically considered a global feature representing the macroscopic semantics of the entire text. To ensure that the dimension of the text features is consistent with the dimension of the water chemical pollution feature vector in step S102 and the key-value terms of the subsequent attention mechanism, a linear projection layer is constructed. Perform feature space transformation to obtain the final geographic environment attribute feature vector. The final geographic environment attribute feature vector The calculation formula is as follows: ; In the formula, The original semantic vector output by the pre-trained model; The weight matrix of the linear projection layer; It is the bias vector; This is the output feature vector of geographic environment attributes. It serves as a unified dimension for the multimodal fusion feature space.

[0039] Step S104: Construct an integrated model for sewage outlet detection and source tracing; wherein, the integrated model for sewage outlet detection and source tracing includes a sewage outlet detection network and a multimodal source tracing network, the sewage outlet detection network is used to output the spatial location prediction result of the sewage outlet, and the multimodal source tracing network is used to output the pollution source attribute discrimination result; The discharge outlet detection network is used to output the spatial location prediction results of discharge outlets based on remote sensing visual features. The multimodal source tracing network is used to model pollution source features and environmental semantic features under the constraint of the spatial location of discharge outlets, guided by remote sensing visual features, and through attention-weighted feature association, and output the pollution source attribute discrimination results.

[0040] Step S105: Construct a composite objective function, and perform joint optimization training on the integrated model for detecting and tracing sewage outlets based on the composite objective function, the remote sensing visual features, the pollution source features, and the environmental semantic features, until the preset training requirements are met, and obtain the trained integrated model for detecting and tracing sewage outlets. In some embodiments, step S105 may include: constructing a composite objective function; wherein the composite objective function includes target detection regression loss, pollution source tracing classification loss, and logical consistency loss, the logical consistency loss being used to constrain the consistency between the pollution source attribute discrimination result and the environmental physical rules based on water flow direction, terrain elevation, and spatial location relationship; inputting remote sensing visual features, pollution source features, and environmental semantic features into the integrated model for discharge outlet detection and source tracing, and calculating the target detection regression loss, pollution source tracing classification loss, and logical consistency loss of the integrated model for discharge outlet detection and source tracing during forward propagation based on the remote sensing visual features, pollution source features, and environmental semantic features; weighting the target detection regression loss, pollution source tracing classification loss, and logical consistency loss to obtain the total loss value of the composite objective function; and optimizing the model parameters of the integrated model for discharge outlet detection and source tracing through backpropagation based on the total loss value of the composite objective function until the preset training requirements are met, thereby obtaining the trained integrated model for discharge outlet detection and source tracing.

[0041] The composite objective function can be understood as the overall loss formed by integrating various loss terms, with each loss term being an input component of the composite objective function. Optionally, the composite objective function is a mathematical expression composed of a weighted average of the target detection regression loss, pollution source tracing classification loss, and logical consistency loss. It is used to output the total loss value for model training, and finally, the model can be optimized based on this composite objective function.

[0042] In step S105, firstly, a composite objective function and preset composite objective function rules are defined for training the integrated model of sewage outlet detection and source tracing. During the training process, remote sensing visual features, pollution source features, and environmental semantic features are input into the integrated model. Then, based on the calculation rules corresponding to each loss term, each loss term is calculated uniformly during the forward propagation process to obtain the target detection regression loss, pollution source tracing classification loss, and logical consistency loss, respectively. Then, the above three loss terms are substituted into the composite objective function for weighted summation to obtain the total loss value of model training. Finally, the model parameters are optimized through backpropagation using this total loss value to obtain the trained integrated model of sewage outlet detection and source tracing.

[0043] In some embodiments, the step of calculating the target detection regression loss may include: extracting a salient feature map representing the temperature anomaly distribution from the thermal infrared imagery in the UAV remote sensing imagery using a shallow convolutional network; performing spatial weighted fusion on the remote sensing visual features based on the salient feature map to obtain enhanced visual features; calculating an elevation gradient field based on digital surface model data in the UAV remote sensing imagery, and generating a binarized topological mask for identifying the land-water interface area based on the elevation gradient field and a preset topological gradient threshold; constraining the spatial sampling range of attention calculation using the binarized topological mask based on a mask attention mechanism, and performing attention calculation on the enhanced visual features to obtain attention-enhanced features; outputting the spatial location prediction result of the sewage outlet based on the attention-enhanced features; and calculating the target detection regression loss based on the deviation between the predicted spatial location of the sewage outlet and the actual location label using an online hard example mining strategy.

[0044] In the specific implementation, DETR is used as the baseline model for target detection. Thermal infrared features are fused at the input end to enhance the discrimination capability of visible light images. At the decoding end, the spatial elevation information of the digital surface model is combined, and the model's attention range for potential sewage discharge areas is constrained through a masked attention mechanism. The output sewage outlet detection bounding box and its confidence score are then calculated using the IoU of the online hard example mining strategy (OHEM). The specific implementation process is as follows: Steps S1051a-S1051c: Step S1051a: At the input end of the DETR model, input the thermal infrared image tensor obtained in step S101 after preprocessing. Input a shallow convolutional network, and extract salient feature maps representing the distribution of temperature anomalies through the shallow convolutional network. Then, the spatial-spectral interpretation feature matrix output in step S101 is used to interpret the saliency feature map. Perform spatial weighting operations to generate enhanced visual features. The specific calculation formula is as follows: ; in, This represents the convolution operation; It is a non-linear activation function used to map salient features to... interval; It is a learnable intensity adjustment coefficient used to adaptively control the enhancement of visible light features by thermal infrared information; This represents the Hadamard product (i.e., element-wise multiplication). Through this operation, the response values ​​corresponding to regions of temperature aberration in the visible light feature map are explicitly amplified.

[0045] Step S1051b introduces spatial sampling constraints into the self-attention mechanism of the DETR model decoder, forcing the model to focus its attention on high-potential sewage discharge areas at the water-land interface. The elevation gradient field is then calculated based on the digital terrain model (DSM) data from step S101. Set the terrain gradient threshold. Construct a binary topological mask for identifying the land-water boundary region. The specific calculation formula is as follows: ; In the formula, For spatial coordinates, This represents the numerical range that characterizes the slope features of a typical shoreline.

[0046] Next, in the self-attention calculation stage of the model, the binarized topological mask is... Superimposed on the attention score matrix, combined with the enhanced visual features The encoded feature vector generated by the encoder of the sewage outlet detection network is input (the encoded features are transformed into the query vector Q, key vector K, and value vector V required by the self-attention mechanism) to apply spatial sampling constraints. The specific calculation formula is as follows: ; In the formula, This represents the attention-enhanced features after spatial sampling constraints and binarized topological mask constraints. These are the query vector, key vector, and value vector, respectively. Because the non-coastal region is... The value in After the Softmax operation, the corresponding attention weights approach 0, thus forcing the model to sample and aggregate only the features of the water-land interface area during the decoding stage.

[0047] Step S1051c, calculate the outlet detection loss: use the attention enhancement features obtained in step S1051b. Input the detection regression branch, output the predicted geographic coordinate bounding box of the sewage outlet. (i.e., the predicted spatial location of the sewage outlet); then calculate the prediction frame. With the actual annotation box The geometric deviation between them generates the entity detection regression loss. To address the challenges of small sample sizes (few sewage outlets) and difficult sample sizes (easily confused and challenging to detect), an online hard case mining (OHEM) strategy is employed to calculate the target detection regression loss. The specific calculation formula is as follows: ; In the formula, The detection location regression loss function (IoU Loss) is used, which is the target detection regression loss. The loss values ​​in the current training batch are ranked first. A set of difficult example samples; This represents the number of difficult examples. This step outputs... This will be used as part of the subsequent construction of the composite objective function to guide the gradient updates of the model.

[0048] In some embodiments, the step of calculating the pollution source tracing classification loss may include: performing feature transformation on remote sensing visual features to construct query features; concatenating pollution source features with environmental semantic features, and performing feature transformation on the environmental context features obtained after feature concatenation to construct key features and value features; employing a scaling dot product attention mechanism to calculate association weights based on query features and key features, and performing weighted summation on value features based on association weights to obtain multimodal fusion features that integrate water quality information and environmental semantic information; inputting the multimodal fusion features into a classification layer to output pollution source attribute discrimination results; and employing a cross-entropy loss function to calculate the pollution source tracing classification loss based on the pollution source attribute discrimination results and the actual pollution source labels.

[0049] In the specific implementation, the remote sensing visual features extracted in step S101 are used as the query vector, and the water chemical pollution feature vector generated in step S102 and the geographic environmental semantic feature vector generated in step S103 are used as the key vector and value vector, respectively. A multimodal cross-attention module is used to achieve interactive fusion of heterogeneous features, and a fully connected classifier is used to output the pollution source attribute discrimination result. The corresponding pollution source tracing classification loss is calculated. The specific implementation process is as follows: steps S1052a-S1052e: Step S1052a, Configuration (Query): Configure the spatial-spectral interpretation feature matrix extracted in step S101. Spatial dimensional flattening and linear projection are performed to construct a visual query sequence. The specific calculation formula is as follows: ; In the formula, , The number of pixel regions representing visual features. This is a learnable query projection matrix.

[0050] Step S1052b, Configure Key-Value Items: This involves configuring the water chemical pollution feature vector output in step S102. The geographic environment attribute feature vector output in step S103 Perform feature concatenation to construct an environmental context feature vector. Subsequently, a key sequence is generated by performing a linear transformation on the environmental context feature vector. AND value sequence The specific calculation formula is as follows: ; In the formula, , This represents the number of context features (here, the length of the concatenated features). This is the corresponding projection matrix.

[0051] Step S1052c involves using a scaled dot product attention mechanism to calculate the association weights of environmental context information with visual features. This mechanism enables the model to adaptively focus on specific regions in the visual feature map based on abnormal water quality indicators or environmental text descriptions. The specific calculation formula is as follows: ; In the formula, A semantic relevance score matrix representing the relationship between visual features and environmental features; This is a scaling factor used to prevent the gradient from vanishing due to excessively large dot product values; The operation is used to normalize the relevance scores and generate the attention weight distribution; A visual feature enhancement sequence that integrates water quality and environmental semantics.

[0052] Step S1052d: Fusing feature decoding and classification output to obtain global feature representation, visual feature enhancement sequence of attention output is applied. Perform global average pooling or extract classification tokens (CLS tokens) to obtain the global multimodal feature vector. The data is then fed into a fully connected classification layer (Classifier Head), and the Softmax activation function outputs the probability distribution vector of the discharge outlet belonging to each pollution source category. The specific calculation formula is as follows: ; In the formula, , This represents the preset total number of pollution source categories (industrial discharge outlets, municipal sewage treatment plant discharge outlets, agricultural discharge outlets, and other discharge outlets). Indicates belonging to the first Confidence level of the class.

[0053] Step S1052e, Pollution source classification loss: During the forward propagation phase of training, based on the real pollution source labels... (One-hot encoding) Calculate pollution source tracing and classification loss using the cross-entropy loss function. The loss value The deviation between the current predictions of the quantification model and the actual pollution attributes is used as a component of the composite objective function.

[0054] In some embodiments, the step of calculating logical consistency loss may include: constructing a logical relationship matrix between environmental scenarios and pollution sources; mapping environmental semantic features to environmental scenario indices; extracting logical constraint mask vectors from the logical relationship matrix based on the environmental scenario indices; and performing consistency verification on the pollution source attribute discrimination results based on the logical constraint mask vectors to obtain the logical consistency loss.

[0055] In the specific implementation, a differentiable logical rule base is constructed based on geographical and environmental physical rules such as water flow direction, topographic elevation gradient, and the relative positional relationship between sewage outlets and pollution sources. During the model forward propagation process, the pollution source attribute discrimination results output in step S105 are checked for logical consistency. If the prediction results conflict with the rule base, the logical consistency loss is calculated using a differentiable logical loss function and used as a rule penalty term to constrain the model. The specific implementation process is as follows: steps S1053a-S1053c: Step S1053a: Based on expert knowledge in the field of environmental science, quantify the compatibility between different coastal environmental backgrounds and various pollution sources, and construct a binary logical relationship matrix between environmental scenarios and pollution sources. The specific calculation formula is as follows: ; In the formula, The total number of predefined environmental scene categories (corresponding to the geographic environment attribute in step S103). The total number of pollution source categories (corresponding to the model output categories in step S1052d); matrix element values (A value of 1.0) represents a strong penalty, meaning that the occurrence of such a sewage outlet in this environmental context is an extremely low probability event that defies common sense; a value of 0 represents permission.

[0056] Step S1053b: In order to determine the logical constraint row vector corresponding to the current sample, the continuous geographic environment attribute feature vector output in step S103 needs to be... This is transformed into a discrete environmental scene index. A lightweight fully connected layer classifier is constructed to... Mapped to a specific index in the set of environment scenes The specific calculation formula is as follows: ; In the formula, An index representing the most likely macroscopic environmental scenario to which the current sample belongs; $b_{scene}$ is a learnable weight matrix used to assist the classifier in mapping feature vectors from a high-dimensional semantic space to a scene category space; $b_{scene}$ is a bias vector.

[0057] According to the index From the logical conflict matrix Extract the first Row vector, denoted as the logical constraint mask vector of the current sample. The specific calculation formula is as follows: ; The position of the non-zero element in the vector in the formula indicates the type of pollution that is prohibited in the current environment.

[0058] Step S1053c: Obtain the pollution source attribute discrimination result output in step S105. (i.e., the probability distribution vector predicted by the model). Using the logical constraint mask vector. With the predicted probability vector Perform dot product operation to generate logical conflict loss. (i.e., logical consistency loss), the specific calculation formula is as follows: ; In the formula, The model predicts that the sewage outlet belongs to the first The probability of a pollution source; For the first element in the mask vector The value of each element. If the model predicts the result... Higher probabilities were assigned to logically mutually exclusive categories (i.e.) Very large, and corresponding ),but This will significantly increase, resulting in a high loss value. During backpropagation, this loss will force the model to suppress the predicted probabilities of mutually exclusive classes, thereby achieving logistic consistency regularization.

[0059] In the specific implementation, a composite objective function is set, which is composed of a weighted average of detection regression loss, source tracing classification loss, and logical consistency loss. The model parameters of the feature extraction network, the sewage outlet detection network, and the multimodal source tracing network are simultaneously optimized through backpropagation until the overall model converges, thus obtaining a fully trained integrated sewage outlet detection and source tracing model. The specific implementation process is as follows: steps S1054a-S1054b: Step S1054a: To balance the spatial positioning accuracy, attribute classification accuracy, and logical compliance of the sewage outlet, a composite objective function is constructed. This function is composed of a weighted sum of the target detection regression loss, the pollution source tracing classification loss, and the logistic consistency loss. The specific calculation formula is as follows: ; in, (Object detection regression loss) is used to quantify the geometric positional deviation between the predicted bounding box and the ground truth bounding box; (Pollution source tracing classification loss) is used to quantify the difference in probability distribution between the predicted category and the true label; (Logical consistency loss) is used to quantify the degree to which prediction results violate common sense about the environment; These are preset hyperparameter weights used to balance the contribution of different tasks in gradient updates (for example, a larger weight can be set in the early stages of training). Prioritize stable target location capabilities.

[0060] Step S1054b: Using the gradient-based optimization algorithm Adam, calculate the total loss for each learnable parameter in the network. gradient and using the learning rate Update parameters Repeat the above forward inference (step S1053c of steps S101-S105) and backpropagation (step S1054a-S1054b of steps S105) processes until the composite objective function is achieved. The model converges to a preset threshold or reaches the maximum number of iterations, resulting in a trained integrated model for sewage outlet detection and tracing with fixed parameters.

[0061] Step S106: Input the current visible light orthophoto of the area to be predicted into the integrated model for detecting and tracing sewage outlets after training, and output the current sewage outlet prediction result; the current sewage outlet prediction result includes the current sewage outlet spatial location prediction result and the current pollution source attribute discrimination result.

[0062] Among them, the current visible light orthophoto image refers to the preprocessed image of the area to be predicted, which conforms to the input specifications of the integrated model for sewage outlet detection and source tracing, after scale normalization, radiometric correction and pixel value normalization.

[0063] In the specific implementation, after the model training is completed, the UAV remote sensing image of the current area to be predicted is first acquired. At this time, only visible light orthophotos are needed for the UAV remote sensing image. Then, the visible light orthophoto is processed by scale normalization, radiometric correction, and pixel value normalization according to the data preprocessing method in step S101 to obtain a preprocessed image that conforms to the input specifications of the integrated sewage outlet detection and source tracing model. Subsequently, the preprocessed visible light orthophoto is input into the fully trained integrated sewage outlet detection and source tracing model obtained in step S105, and forward inference is performed. First, the sewage outlet detection network outputs the current sewage outlet spatial location prediction result, and then the multimodal source tracing network judges the pollution source attribute of the sewage outlet based on the current sewage outlet spatial location prediction result, and outputs the current pollution source attribute judgment result corresponding to the current sewage outlet spatial location prediction result. Specifically, the specific implementation process of step S106 is as follows: steps S106a to S106e: Step S106a: First, acquire the visible light orthophoto image of the area to be predicted. ,in, , These represent the spatial resolution of the image. The image band number is indicated; then, the visible light orthophoto image is subjected to scale normalization, radiometric correction, and pixel value normalization to obtain a preprocessed image that conforms to the input specifications of the integrated model for sewage outlet detection and source tracing. ,in, This is a preprocessing operator.

[0064] Step S106b: The preprocessed image The input is fed into the backbone encoding network of the trained integrated model for sewage outlet detection and source tracing. This backbone encoding network extracts multi-scale feature maps with semantic information awareness using ResNet101. The calculation expression for feature extraction is as follows: ; In the formula, A network for encoding visual features; and This represents the number of layers in the feature pyramid. The lower-level features contain rich spatial details such as edges and textures, which are helpful in locating small sewage outlets; the higher-level features contain strong semantic information (such as whether it is an industrial outlet), which is used to assist in source tracing and identification.

[0065] Step S106c: Convert the multi-scale feature map Input the sewage outlet detection subnetwork. This subnetwork performs location regression and confidence scoring on potential sewage outlet targets, initially generating a candidate set. ,in, Let the coordinates be those of the candidate bounding boxes. The confidence score has a range of values. A higher value indicates a greater likelihood that the model determines the presence of a sewage outlet in that area. Subsequently, non-maximum suppression (NMS) is used to filter out overlapping and low-confidence targets, thus determining the final spatial location of the sewage outlet. .

[0066] Step S106d, based on the predicted spatial location The feature alignment operator ROI Align is used to extract and aggregate features of corresponding regions on multi-scale feature maps, thereby constructing instance-level feature vectors for sewage outlets. ,in, For instance ROI Align feature alignment operator, this vector refines the remote sensing visual representation of a single sewage outlet and its surrounding environment.

[0067] Step S106e: Convert the sewage outlet instance-level feature vector Input the attribute discrimination module of the multimodal source tracing network. It should be noted that this attribute discrimination module, during the training phase, performs joint learning by fusing pollution source features and environmental semantic features, establishing an implicit mapping between visual features and pollution attributes. During the prediction phase, the model only needs to input the instance-level feature vectors of the sewage outlet from the visible light orthophoto image. This mapping relationship allows for the accurate output of the corresponding pollution source attribute identification results. ,in, This represents a classification function indicating the source attributes of pollution. The results are for tracing the source and making a judgment.

[0068] In some embodiments, after step S106, the method may further include: performing a non-maximum suppression operation on the current sewage outlet spatial location prediction result to obtain the target sewage outlet spatial location prediction result; performing a confidence filtering operation on the current pollution source attribute discrimination result to obtain the target pollution source attribute discrimination result; performing visualization processing on the target sewage outlet spatial location prediction result and the target pollution source attribute discrimination result to generate spatialized representation data, and displaying the spatialized representation data on the visualization interface.

[0069] In this embodiment, a non-maximum suppression operation can be performed on the spatial location prediction results of the sewage outlets output in step S106 to eliminate redundant detection boxes, and a confidence screening operation can be performed on the pollution source attribute discrimination results output in step S106 to remove low-confidence discrimination results and eliminate duplicate or conflicting attribute judgments. Then, the retained spatial location prediction results of the sewage outlets and their corresponding pollution source attribute discrimination results are mapped to a unified geographic coordinate reference system. Finally, based on the pollution source attribute categories, spatialized representation data of the sewage outlet investigation and pollution source tracing results are generated according to preset symbolization and coding rules for use in sewage outlet distribution analysis and pollution source tracing decision support.

[0070] Steps S101 to S106 of this embodiment, by fusing UAV remote sensing visual features, cross-sectional water quality monitoring features, and shoreline environmental semantic features, achieve collaborative discrimination of the spatial location and pollution source attributes of sewage outlets into rivers. Compared with related technologies that rely solely on remote sensing visual information, this embodiment can more comprehensively characterize the spatial morphology, pollution emission characteristics, and environmental background information of sewage outlets in complex riverside environments, thereby effectively improving the detection accuracy of concealed sewage outlets and the reliability of pollution source discrimination. Furthermore, this embodiment introduces a composite objective function during the model training stage to jointly optimize the sewage outlet detection task and the pollution source tracing task, and combines environmental physical rules to constrain the consistency of model prediction results, effectively suppressing irrational predictions that may arise from deep learning models and enhancing the logical rationality and stability of the prediction results. Therefore, this embodiment can ensure detection accuracy while improving the accuracy and robustness of pollution source discrimination, even when relying solely on visible light orthophoto imagery input during the prediction stage, providing reliable technical support for the refined investigation and pollution source tracing of sewage outlets into rivers.

[0071] To explain in detail the principles of the technical solution of this application, the overall process of this application will be described below with reference to some specific embodiments. It is easy to understand that the following is an explanation of the technical principles of this application and should not be regarded as a limitation of this application.

[0072] The purpose of this application is to overcome the shortcomings of related technologies in identifying concealed sewage outfalls into rivers, determining pollution source attributes, and improving the reliability of model prediction logic. This application proposes a sewage outfall detection and source tracing method based on multimodal cross-attention and logical rule constraints. During the model training phase, this method integrates multi-source heterogeneous data, jointly utilizing visible light orthophotos (DOM), thermal infrared remote sensing images (TIR), and digital surface models (DSM) acquired by UAVs to extract remote sensing visual features. It also introduces cross-sectional water quality monitoring indicators to extract pollution characterization features and combines riparian geographic environmental text data to extract environmental semantic features. A multimodal cross-attention mechanism guides the association learning between remote sensing visual features, water quality features, and semantic features. Simultaneously, a logical consistency loss function is introduced to constrain the model's inference process with physical rules and environmental knowledge. This allows the model to establish a stable mapping relationship between remote sensing visual features and pollution source attributes during the training phase. During the prediction phase, it can achieve joint determination of the spatial location of sewage outfalls and pollution source attributes solely based on UAV visible light orthophotos (DOM), improving the accuracy, reliability, and engineering feasibility of sewage outfall detection and source tracing in complex riverbank environments.

[0073] Please see Figure 2 , Figure 2 This is a flowchart illustrating a method for detecting and tracing sewage outlets into rivers, as provided in an embodiment of this application. Figure 2 As shown in the figure, the specific implementation process of the method for detecting and tracing sewage outlets into rivers provided in this application embodiment is as follows: steps S1 to S9: Step S1, extract remote sensing visual features of the river section to be monitored: First, collect UAV remote sensing images of the river section to be monitored; then, after performing geometric correction, spatial registration, cropping and other preprocessing on the UAV remote sensing images, input them into the deep learning visual feature extraction network ResNet-50 to extract the spatial-spectral interpretation feature matrix representing the physical morphology of the sewage outlet into the river, and use the spatial-spectral interpretation feature matrix as the remote sensing visual feature. Step S2, extract pollution source features of the river section to be monitored: First, obtain the cross-sectional water quality monitoring data corresponding to the UAV remote sensing image in Step S1; then, after standardizing the cross-sectional water quality monitoring data, input it into the feature mapping network to map the discrete physicochemical indicators into continuous water chemical pollution feature vectors, and use the water chemical pollution feature vectors as pollution source features. Step S3: Extract the environmental semantic features of the river section to be monitored. First, based on the monitoring range defined in Step S1, obtain the unstructured text data of the riverbank geographic environment corresponding to the river section. Then, use the BERT language model to perform semantic understanding and vectorization of the text, generate structured geographic environment attribute feature vectors, and use the structured geographic environment attribute feature vectors as environmental semantic features. Step S4: Construct a fusion mechanism between the sewage outlet detection network and remote sensing visual features: Using DETR as the target detection baseline model, thermal infrared features are fused at the input end to enhance the discrimination capability of visible light images; at the decoding end, the spatial elevation information of the digital surface model is combined, and the model's attention range for potential sewage discharge areas is constrained through a mask attention mechanism, outputting the sewage outlet detection bounding box and its confidence score, and the target detection regression loss is calculated by using the IoU of the online hard example mining strategy (OHEM); Step S5: Construct a multimodal cross-attention source tracing network: Use the remote sensing visual features extracted in step S1 as the query vector, and use the water chemical pollution feature vector generated in step S2 and the geographic environmental semantic feature vector generated in step S3 as the key and value vectors, respectively. The multimodal cross-attention module realizes the interaction and fusion of heterogeneous features, and outputs the pollution source attribute discrimination result based on the fully connected classifier, and calculates the corresponding pollution source tracing classification loss. Step S6: Construct a logical rule constraint module: Based on geographical and environmental physical rules such as water flow direction, topographic elevation gradient, and the relative positional relationship between sewage outlets and pollution sources, construct a differentiable logical rule base; During the forward propagation of the model, perform logical consistency verification on the pollution source category prediction results output in Step S5. If the prediction results conflict with the rule base, calculate the logical consistency loss through a differentiable logical loss function, and use it as a rule penalty term to constrain the model. Step S7: Design a composite loss function for end-to-end joint training: Set a composite objective function consisting of the detection regression loss of step S4, the source tracing classification loss of step S5, and the logical consistency loss of step S6. Simultaneously optimize the model parameters of the feature extraction network, the target detection module, and the multimodal source tracing network through backpropagation until the overall model converges and a fully trained integrated model for sewage outlet detection and source tracing is obtained. Step S8, Joint Inference of Outlet Detection and Pollution Source Tracing: After the model training is completed, only the visible light orthophoto image (DOM) of the area to be predicted is acquired and preprocessed according to step S1; the preprocessed visible light orthophoto image is input into the fully trained integrated model of outlet detection and source tracing obtained in step S7 to perform forward inference, and output the spatial location prediction result of the outlet and its corresponding pollution source attribute discrimination result; wherein, the pollution source attribute discrimination result is based on the cross-modal correlation learned in the training stage, and is generated by the multimodal source tracing network without introducing cross-sectional water quality monitoring data and shoreline text data.

[0074] Step S9, Spatial Representation of Discharge Outlet Detection and Source Tracing Results: Perform non-maximum suppression on the discharge outlet spatial location prediction results and their corresponding pollution source attribute discrimination results output in Step S8 to eliminate redundant detection boxes, and map the retained discharge outlet spatial locations and their pollution source attributes to a unified geographic coordinate reference system; based on the pollution source attribute categories, generate spatial representation data of discharge outlet investigation and pollution source tracing results according to preset symbolization and coding rules, which is used for discharge outlet distribution analysis and pollution source tracing decision support.

[0075] It should be noted that this embodiment is only a brief illustrative description of the overall process of a method for detecting and tracing sewage outlets into rivers. Detailed descriptions of each step can be found in the relevant content of the foregoing embodiments, and will not be repeated here. It is understood that this application does not impose any limitations on this.

[0076] To verify the effectiveness and feasibility of the proposed method for detecting and tracing sewage outfalls into rivers based on multimodal cross-attention and logical rule constraints, a cross-seasonal verification experiment of sewage outfall investigation was conducted in a representative watershed of a province. The experimental procedure followed... Figure 2 Steps S1 to S9 are performed as follows: (1) Data Acquisition and Multimodal Feature Encoding (corresponding to steps S1-S3): Following step S1, acquire UAV remote sensing data for the watershed, including visible light orthophotos (DOM), digital surface model (DSM), and thermal infrared (TIR) ​​images. Simultaneously, following steps S2 and S3, acquire cross-sectional water quality monitoring data and corresponding unstructured text data of the shoreline geographic environment. Please refer to... Figure 3 , Figure 3 This is a visual sample diagram of a sewage outlet into a river provided in an embodiment of this application, such as... Figure 3 As shown, the visual sample set constructed in step S2 of this application covers a diverse range of challenging discharge port scenarios, and is composed of... Figure 3 As can be seen, the sample set includes not only obvious exposed pipe outlets, but also concealed outlets, hidden drainage ditches in farmland, water outlets obscured by vegetation, and complex road culverts. These challenging visual samples are effectively incorporated into the training system through thermal infrared saliency guidance in step S4 and multimodal feature encoding in step S2, ensuring the model's ability to perceive targets with weak visual features.

[0077] (2) Model training and rainy season investigation and verification (corresponding to steps S4-S7), please refer to Figure 4 , Figure 4 This is a schematic diagram illustrating the detection and source tracing results of sewage outfalls into rivers during the rainy season in a typical watershed, as provided in the embodiments of this application. The results of verifying the method of this application with data from the rainy season in the watershed are as follows. Figure 4 As shown in Table 1: Table 1: Statistical Table of Detection and Source Tracing Accuracy of Different Types of Sewage Outfalls into Rivers Using the Method Applied in This Application (Rainy Season)

[0078] Depend on Figure 4 As shown in Table 1, under the conditions of large water flow and complex environmental background during the rainy season, the model successfully achieved high recall of densely distributed sewage outlets into the river in the basin by means of the enhanced detection mechanism in step S4. Figure 4 The study visually illustrates the geospatial distribution of sewage outlets during the rainy season and verifies the effectiveness of the multimodal cross-attention mechanism employed in step S5 in distinguishing complex pollution source types. The model can clearly identify industrial outlets (marked in red in the figure, mainly distributed in main waterways and around towns), and effectively differentiate and locate dispersed pollution sources such as urban domestic sewage outlets, rural domestic sewage outlets, and large-scale livestock and poultry farm outlets (marked in yellow and purple, respectively). In terms of quantitative analysis, the F1 score for all categories in Table 1 is above 92.2%, demonstrating a good balance between precision and recall. Particularly for rural sewage treatment facility outlets, the model achieved the highest F1 score of 96.04% (Precision = 96.15%, Recall = 95.93%), fully demonstrating the efficient discrimination capability of the S5 mechanism, which integrates hydrochemical characteristics and shoreline geographic semantic features, for pollution sources with clearly defined facility characteristics. Furthermore, the model recall for rural domestic sewage outfalls reached 96.43%, indicating that the model has extremely high coverage of these easily overlooked non-point sources / scattered sources. This comprehensively verifies the multi-dimensional advantages of the proposed method in accurately identifying and reliably tracing sewage outfalls under complex hydrological conditions.

[0079] (3) Online detection during the dry season and cross-seasonal application (corresponding to step S8): To verify the model's generalization ability, following step S8 above, the trained model was directly applied to online inference using dry season imagery data from the watershed. Please refer to [link / reference]. Figure 5 , Figure 5 This is a schematic diagram illustrating the detection and source tracing results of sewage outfalls into rivers during the dry season in a typical watershed, provided in an embodiment of this application. (Comparison) Figure 4 (Rainy season) and Figure 5 (Dry season) As can be seen, despite seasonal changes leading to lower water levels, altered riparian vegetation texture, and fluctuations in water quality physicochemical concentrations, the model in this application can still stably locate the coordinates of discharge outlets and infer their attributes. The statistical results of the identification accuracy of various discharge outlets under dry season conditions are shown in Table 2: Table 2: Statistical Table of Detection and Source Tracing Accuracy of Different Types of Sewage Outfalls into Rivers Using the Method Applied in This Application (Dry Season)

[0080] Table 2 shows that the F1 scores for sewage outlets from industrial park wastewater treatment plants, urban wastewater treatment plants, and rural domestic sewage outfalls reached 91.33%, 93.55%, and 93.27% respectively during the dry season, maintaining a high cross-seasonal identification capability. The source tracing and classification accuracy of rural sewage treatment facility outfalls during the dry season was 92.36% Precision, 94.82% Recall, and 93.50% F1, demonstrating the model's strong robustness to environmental changes. This result fully demonstrates that by introducing environmental science logic rules as constraints in step S6, the model effectively overcomes the shortcomings of traditional visual detection methods that are prone to failure during seasonal changes, achieving automated and accurate investigation capabilities for river outfalls across seasons and all weather conditions.

[0081] (4) Comparative experiments with other methods: To further verify the significant advantages of the entire process technology system established in steps S1-S9 of this application in complex scenarios, the general object detection model YOLOv9 and Faster R-CNN were selected as control groups for experiments. The specific comparison results are shown in Table 3: Table 3: Performance Comparison of Different Detection Models in the Task of Investigating Sewage Outfalls into Rivers

[0082] As shown in Table 3, the control group model relies solely on the visible light image features from step S1, lacking the multimodal cross-attention mechanism in step S5 and the logical rule verification in step S6 provided in this application. This makes it difficult to distinguish visually similar interfering targets, with an F1 score of only 70.8% to 78.5% and an IoU of 57.13% to 68.36%. In contrast, the method in this application, through thermal infrared and topological constraints in step S4, semantic injection in step S5, and logical consistency loss in step S6, effectively overcomes the recognition difficulties caused by weakened visual features, increasing the F1 score to 94.7% and the IoU to 88.31%. This fully demonstrates the necessity of multimodal semantic guidance and logical constraint mechanisms in the task of investigating and tracing the source of sewage discharge outlets into rivers.

[0083] In summary, the beneficial effects of the method provided in the embodiments of this application are as follows: (1) By extracting remote sensing visual features from UAV remote sensing images and introducing thermal infrared features for spatial weighting during the sewage outlet detection process, and combining the elevation information of the digital surface model (DSM) to construct topological mask attention constraints, the defects of single visible light images being easily affected by vegetation shading and changes in lighting conditions are effectively overcome, significantly improving the identification accuracy and anti-interference ability of hidden sewage outlets, and providing a foundation for high-precision detection relying solely on UAV remote sensing images in the prediction stage.

[0084] (2) Water chemical pollution features and geographic semantic features based on the BERT model are introduced in the model training stage, and remote sensing visual features are jointly constrained through a multimodal cross-attention mechanism, so that the model can fully learn the intrinsic relationship between pollution source attributes and remote sensing spatial morphology during the training process. Therefore, in the prediction stage, there is no need to introduce cross-sectional water quality monitoring data and environmental semantic data. The pollution source attribute discrimination results of the sewage outlet can be stably output based solely on UAV remote sensing images, which effectively reduces the dependence on external monitoring conditions and improves the feasibility and applicability of the method in actual regulatory scenarios.

[0085] (3) In the model optimization stage, a composite objective function is constructed to jointly train the sewage outlet detection task and the pollution source tracing task. The online hard example mining (OHEM) strategy and the logical consistency constraint based on physical rules and environmental knowledge are combined to globally constrain the model inference results. This design significantly improves the stability and robustness of the prediction results without increasing the complexity of the input information in the prediction stage, suppresses the irrational predictions that may occur under complex environmental conditions, and ensures the consistency between the prediction results and the objective environmental laws.

[0086] Please see Figure 6 This application also provides a system 600 for detecting and tracing sewage outlets into rivers, which can implement the above-mentioned method. The system includes the following modules: The remote sensing visual feature generation module 601 is used to acquire UAV remote sensing images of the river section to be monitored, and to perform a first feature processing on the UAV remote sensing images to obtain remote sensing visual features. The pollution source feature generation module 602 is used to acquire cross-sectional water quality monitoring data that is synchronized with the UAV remote sensing image in time and space, and to perform a second feature processing on the cross-sectional water quality monitoring data to obtain pollution source features. The environmental semantic feature generation module 603 is used to acquire the shoreline environmental text data of the river section to be monitored, and to perform a third feature processing on the shoreline environmental text data to obtain environmental semantic features. The model building module 604 is used to build an integrated model for sewage outlet detection and source tracing; wherein, the integrated model for sewage outlet detection and source tracing includes a sewage outlet detection network and a multimodal source tracing network, the sewage outlet detection network is used to output the spatial location prediction result of the sewage outlet, and the multimodal source tracing network is used to output the pollution source attribute discrimination result; The model training module 605 is used to construct a composite objective function and perform joint optimization training on the integrated model of sewage outlet detection and source tracing based on the composite objective function, the remote sensing visual features, the pollution source features and the environmental semantic features until the preset training requirements are met, and the trained integrated model of sewage outlet detection and source tracing is obtained. The sewage outlet detection and tracing module 606 is used to input the current visible light orthophoto of the area to be predicted into the integrated sewage outlet detection and tracing model that has been trained, and output the current sewage outlet prediction result; the current sewage outlet prediction result includes the current sewage outlet spatial location prediction result and the current pollution source attribute discrimination result.

[0087] It is understood that the content of the above method embodiments is applicable to this system embodiment. The specific functions implemented in this system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0088] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0089] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0090] Please see Figure 7 , Figure 7 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes: The processor 701 can be implemented using a general-purpose CPU (Central Processing Unit), GPU (Graphics Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 702 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 702 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 702 and is called and executed by the processor 701 using the methods described in the embodiments of this application. The input / output interface 703 is used to implement information input and output; The communication interface 704 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 705 transmits information between various components of the device (e.g., processor 701, memory 702, input / output interface 703, and communication interface 704); The processor 701, memory 702, input / output interface 703, and communication interface 704 are connected to each other within the device via bus 705.

[0091] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.

[0092] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0093] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.

[0094] It is understood that the content of the above method embodiments is applicable to the embodiments of this program product. The specific functions implemented by the embodiments of this program product are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0095] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0096] This application provides a method and system for detecting and tracing sewage outfalls into rivers. By integrating UAV remote sensing visual features, cross-sectional water quality monitoring features, and shoreline environmental semantic features, it achieves collaborative discrimination of the spatial location and pollution source attributes of sewage outfalls. Compared with related technologies that rely solely on remote sensing visual information, this application can more comprehensively characterize the spatial morphology, pollution emission characteristics, and environmental background information of sewage outfalls in complex riverside environments, thereby effectively improving the detection accuracy of concealed sewage outfalls and the reliability of pollution source discrimination. Furthermore, this application introduces a composite objective function during the model training stage to jointly optimize the sewage outfall detection and pollution source discrimination tasks. It also incorporates environmental physical rules to constrain the consistency of model prediction results, effectively suppressing irrational predictions that may arise from deep learning models and enhancing the logical rationality and stability of the prediction results. Therefore, this application can ensure detection accuracy while improving the accuracy and robustness of pollution source discrimination, even when relying solely on visible light orthophoto imagery input during the prediction stage, providing reliable technical support for the refined investigation and pollution source discrimination of sewage outfalls into rivers.

[0097] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A method for detecting and tracing the source of sewage discharge outlets into rivers, characterized in that, The method includes the following steps: Acquire UAV remote sensing images of the river section to be monitored, and perform a first feature processing on the UAV remote sensing images to obtain remote sensing visual features. Acquire cross-sectional water quality monitoring data that is synchronized with the UAV remote sensing image in time and space, and perform a second feature processing on the cross-sectional water quality monitoring data to obtain pollution source characteristics; The shoreline environmental text data of the river section to be monitored is obtained, and the shoreline environmental text data is subjected to a third feature processing to obtain environmental semantic features. An integrated model for sewage outlet detection and source tracing is constructed; wherein, the integrated model for sewage outlet detection and source tracing includes a sewage outlet detection network and a multimodal source tracing network, the sewage outlet detection network is used to output the spatial location prediction result of the sewage outlet, and the multimodal source tracing network is used to output the pollution source attribute discrimination result; A composite objective function is constructed, and the integrated model for detecting and tracing sewage outlets is jointly optimized and trained based on the composite objective function, the remote sensing visual features, the pollution source features, and the environmental semantic features until the preset training requirements are met, thus obtaining the trained integrated model for detecting and tracing sewage outlets. The current visible light orthophoto of the area to be predicted is input into the integrated model for detecting and tracing sewage outlets after training, and the current sewage outlet prediction result is output. The current sewage outlet prediction result includes the current sewage outlet spatial location prediction result and the current pollution source attribute discrimination result.

2. The method according to claim 1, characterized in that, The process of acquiring UAV remote sensing images of the river section to be monitored and performing a first feature processing on the UAV remote sensing images to obtain remote sensing visual features includes: The UAV remote sensing images of the river section to be monitored are acquired; wherein the UAV remote sensing images include visible light orthophotos, digital surface model data and thermal infrared images. Geometric correction, spatial registration, and scale unification processing are performed on the UAV remote sensing images to obtain spatially aligned multi-source remote sensing data. The spatially aligned multi-source remote sensing data is input into a feature extraction network for joint feature extraction, generating remote sensing visual features that characterize the spatial morphology and spatial distribution of sewage outfalls into rivers.

3. The method according to claim 1, characterized in that, The acquisition of cross-sectional water quality monitoring data synchronized in time and space with the UAV remote sensing imagery, and the subsequent second feature processing of the cross-sectional water quality monitoring data to obtain pollution source characteristics, includes: Acquire the cross-sectional water quality monitoring data that is synchronized in time and space with the UAV remote sensing imagery; The water quality monitoring data of the cross section are subjected to data standardization processing to obtain a standardized water quality vector; The standardized water quality vector is input into a multilayer perceptron network for linear transformation and nonlinear activation function processing to obtain the pollution source features in a high-dimensional feature space.

4. The method according to claim 1, characterized in that, The process involves acquiring the shoreline environmental text data of the river section to be monitored, and performing a third featureization process on the shoreline environmental text data to obtain environmental semantic features, including: The geographical range of the river section to be monitored is determined based on the spatial metadata of the UAV remote sensing imagery. Based on the geographical range, the shoreline environmental text data corresponding to the river section to be monitored is obtained from a pre-constructed shoreline geographical environment knowledge database; wherein, the shoreline geographical environment knowledge database includes unstructured text data carrying shoreline geographical tags, and the unstructured text data includes shoreline land use type descriptions, historical pollution source records, and environmentally sensitive receptor information; The unstructured text data in the coastal environmental text data is spliced ​​to obtain the target text sequence, and corresponding classification tags are added to the beginning and end of the target text sequence. Map the target text sequence into an embedding vector matrix; The embedding vector matrix is ​​subjected to hierarchical feature transformation using a bidirectional language representation model to obtain the context hidden layer state matrix; Extract the global text feature vector corresponding to the first classification label in the context hidden layer state matrix, and perform feature space transformation on the global text feature vector to obtain the environmental semantic features.

5. The method according to claim 1, characterized in that, The process involves constructing a composite objective function and then jointly optimizing and training the integrated model for sewage outlet detection and source tracing based on the composite objective function, the remote sensing visual features, the pollution source features, and the environmental semantic features, until the preset training requirements are met, resulting in a trained integrated model for sewage outlet detection and source tracing. This includes: Construct the composite objective function; wherein, the composite objective function includes target detection regression loss, pollution source tracing classification loss and logical consistency loss, and the logical consistency loss is used to constrain the consistency between the pollution source attribute discrimination result and the environmental physical rules based on water flow direction, terrain elevation and spatial location relationship; The remote sensing visual features, pollution source features, and environmental semantic features are input into the integrated model for detecting and tracing sewage outlets. Based on the remote sensing visual features, pollution source features, and environmental semantic features, the target detection regression loss, pollution source classification loss, and logical consistency loss of the integrated model for detecting and tracing sewage outlets are calculated during the forward propagation process. The total loss value of the composite objective function is obtained by weighting the target detection regression loss, the pollution source tracing classification loss, and the logical consistency loss. Based on the total loss value of the composite objective function, the model parameters of the integrated sewage outlet detection and source tracing model are optimized through backpropagation until the preset training requirements are met, thus obtaining the trained integrated sewage outlet detection and source tracing model.

6. The method according to claim 5, characterized in that, The method further includes a step of calculating the target detection regression loss, wherein calculating the target detection regression loss includes: A shallow convolutional network is used to extract salient feature maps representing temperature anomaly distribution from the thermal infrared images in the UAV remote sensing images. Based on the salient feature map, spatial weighted fusion is performed on the remote sensing visual features to obtain enhanced visual features; The elevation gradient field is calculated based on the digital surface model data in the UAV remote sensing image, and a binary topological mask for identifying the water-land interface area is generated according to the elevation gradient field and the preset landform gradient threshold. Based on the mask attention mechanism, the spatial sampling range of attention calculation is constrained by the binarized topological mask, and attention calculation is performed on the enhanced visual features to obtain attention-enhanced features; Based on the attention enhancement feature, the spatial location prediction result of the sewage outlet is output; An online hard case mining strategy is adopted to calculate the target detection regression loss based on the deviation between the predicted spatial location of the sewage outlet and the actual location label.

7. The method according to claim 5, characterized in that, The method further includes a step of calculating the pollution source tracing classification loss, wherein calculating the pollution source tracing classification loss includes: The remote sensing visual features are transformed to construct query features; The pollution source features and the environmental semantic features are concatenated, and the environmental context features obtained after feature concatenation are transformed to construct key features and value features. A scaling dot product attention mechanism is adopted to calculate the association weight based on the query feature and the key feature, and then the value feature is weighted and summed according to the association weight to obtain a multimodal fusion feature that integrates water quality information and environmental semantic information. The multimodal fusion features are input into the classification layer, and the pollution source attribute discrimination result is output. The cross-entropy loss function is used to calculate the pollution source tracing classification loss based on the pollution source attribute discrimination results and the actual pollution source labels.

8. The method according to claim 5, characterized in that, The method further includes a step of calculating the logical consistency loss, wherein calculating the logical consistency loss includes: Construct a logical relationship matrix between environmental scenarios and pollution sources; Map the environmental semantic features to an environmental scene index; Extract the logical constraint mask vector from the logical relationship matrix based on the environmental scene index; Based on the logical constraint mask vector, the consistency verification of the pollution source attribute discrimination result is performed to obtain the logical consistency loss.

9. The method according to claim 1, characterized in that, In the integrated model for detecting and tracing sewage outlets, which has been trained by inputting the current visible light orthophoto of the area to be predicted into the model, after outputting the current sewage outlet prediction result, the method further includes: Perform non-maximum suppression operation on the current sewage outlet spatial location prediction result to obtain the target sewage outlet spatial location prediction result; Perform a confidence filtering operation on the current pollution source attribute identification result to obtain the target pollution source attribute identification result; The spatial location prediction results of the target sewage outlet and the target pollution source attribute discrimination results are visualized to generate spatialized representation data, which is then displayed on a visualization interface.

10. A system for detecting and tracing sewage outlets into rivers, characterized in that, The system includes the following modules: The remote sensing visual feature generation module is used to acquire UAV remote sensing images of the river section to be monitored, and to perform a first feature processing on the UAV remote sensing images to obtain remote sensing visual features. The pollution source feature generation module is used to acquire cross-sectional water quality monitoring data that is synchronized with the UAV remote sensing image in time and space, and to perform a second feature processing on the cross-sectional water quality monitoring data to obtain pollution source features. An environmental semantic feature generation module is used to acquire the shoreline environmental text data of the river section to be monitored, and to perform a third featureization process on the shoreline environmental text data to obtain environmental semantic features. The model building module is used to construct an integrated model for sewage outlet detection and source tracing; wherein, the integrated model for sewage outlet detection and source tracing includes a sewage outlet detection network and a multimodal source tracing network, the sewage outlet detection network is used to output the spatial location prediction results of sewage outlets, and the multimodal source tracing network is used to output the pollution source attribute discrimination results; The model training module is used to construct a composite objective function and perform joint optimization training on the integrated model of sewage outlet detection and source tracing based on the composite objective function, the remote sensing visual features, the pollution source features, and the environmental semantic features until the preset training requirements are met, and the trained integrated model of sewage outlet detection and source tracing is obtained. The sewage outlet detection and tracing module is used to input the current visible light orthophoto of the area to be predicted into the integrated sewage outlet detection and tracing model that has been trained, and output the current sewage outlet prediction result; the current sewage outlet prediction result includes the current sewage outlet spatial location prediction result and the current pollution source attribute discrimination result.