River plastic floating object classification method

CN122695369BActive Publication Date: 2026-09-29CHINA JILIANG UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611178805.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-08-05
Publication Date
2026-09-29
Estimated Expiration
2046-08-05

AI Technical Summary

Technical Problem

[0003]然而,可见光影像与多光谱影像的成像方式不同,两类影像在空间分辨率和像素值分布等方面存在差异,不同单波段影像对同一河道场景的成像结果也存在差异

Benefits of technology

[0006]与现有技术相比,本发明对可见光单通道影像和多光谱单波段影像进行处理后再进行特征匹配,建立可见光影像与多光谱影像之间的坐标映射关系,并根据该坐标映射关系将实例掩膜映射至多光谱影像。可以理解,通过减小两类影像的空间分辨率差异和像素值分布差异,有利于提高特征匹配的位置对应精度,进而减小实例掩膜的映射偏差,使形成的多光谱目标区域与相应漂浮物所在区域相匹配,减少非目标区域反射信息对所提取反射率的干扰,从而有利于提高河道塑料漂浮物分类的准确度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122695369B_ABST
    Figure CN122695369B_ABST
Patent Text Reader

Abstract

The application discloses a river plastic floating object classification method. Visible light images and multispectral images of the same river area are obtained, the visible light images are subjected to instance segmentation, and instance masks of each floating object instance are obtained; a spatial correspondence relationship between the two types of images is established, each instance mask is mapped to the multispectral images, and corresponding multispectral target regions are obtained; a plurality of sampling pixels are selected from each multispectral target region, multispectral reflectivity of each sampling pixel is extracted and spectral features are generated, and a plastic material classification model is used to obtain material classification results of each sampling pixel; and according to the material classification results of each sampling pixel in the same multispectral target region, the plastic material category of the corresponding floating object instance is determined. The application is beneficial to reducing the mixing of background reflection information and reducing the influence of single sampling pixel classification deviation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of water environment monitoring technology, specifically to a method for classifying floating plastic debris in rivers. Background Technology

[0002] The detection of floating plastic debris in waterways typically uses visible light imagery to determine the location and outline of the debris, and multispectral imagery to extract the reflectance of the debris in different spectral bands to determine the plastic material. To ensure that the reflectance extraction area in the multispectral image corresponds to the floating debris in the visible light image, it is necessary to establish a coordinate relationship between the two images and map the outline of the floating debris from the visible light image to the multispectral image.

[0003] However, visible light imagery and multispectral imagery differ in their imaging methods, resulting in variations in spatial resolution and pixel value distribution. Furthermore, different single-band images can produce different imaging results for the same river scene. Directly selecting single-band images for feature matching with visible light images can easily lead to deviations in the determined coordinate relationships. Consequently, the outlines of floating objects in the visible light image may deviate from their corresponding locations when mapped to the multispectral image, causing the extracted reflectance to be mixed with the reflection information from the water surface background or nearby floating objects. Summary of the Invention

[0004] To address the aforementioned technical problems, this invention proposes a method for classifying floating plastic debris in river channels. This method improves the accuracy of mapping instance masks to multispectral images, thereby reducing the interference of non-target area reflectance information on the extracted reflectance and improving the accuracy of floating plastic debris classification in river channels.

[0005] To achieve the above objectives, the present invention provides a method for classifying floating plastic debris in waterways, comprising: Acquire visible light and multispectral images corresponding to the same river channel area; perform instance segmentation on the visible light images to obtain multiple floating object instances and instance masks representing the contour range of each floating object instance; convert the visible light images into single-channel images, and select a single-band image from multiple single-band images corresponding to the multispectral images; adjust the spatial resolution of the selected single-band image according to the spatial resolution of the single-channel images, and adjust the pixel value distribution of the selected single-band image according to the pixel value distribution of the single-channel images to obtain a single-band image with adjusted pixel value distribution; perform feature matching on the single-channel images and the single-band image with adjusted pixel value distribution, and... The feature matching results are transformed based on the pixel position correspondence before and after spatial resolution adjustment to obtain the coordinate mapping relationship between the visible light image and the multispectral image. Based on the coordinate mapping relationship, each instance mask is mapped to the multispectral image to obtain the multispectral target region corresponding to each floating object instance. Multiple sampling pixels are selected from each multispectral target region, and the material classification result corresponding to each sampling pixel is obtained based on the reflectance of each sampling pixel in multiple spectral bands. Based on the material classification result corresponding to each sampling pixel belonging to the same multispectral target region, the plastic material category of the floating object instance corresponding to the multispectral target region is determined.

[0006] Compared with existing technologies, this invention processes visible light single-channel images and multispectral single-band images before performing feature matching, establishing a coordinate mapping relationship between the visible light image and the multispectral image, and mapping the instance mask onto the multispectral image based on this coordinate mapping relationship. It can be understood that by reducing the spatial resolution difference and pixel value distribution difference between the two types of images, it is beneficial to improve the positional accuracy of feature matching, thereby reducing the mapping deviation of the instance mask, ensuring that the formed multispectral target area matches the corresponding floating object area, reducing the interference of non-target area reflection information on the extracted reflectivity, and thus improving the accuracy of river plastic floating object classification. Attached Figure Description

[0007] Figure 1 A flowchart of a method for classifying floating plastic debris in river channels is provided in this application; Figure 2 A schematic diagram illustrating the acquisition of visible light and multispectral images provided in this application; Figure 3 (a) in the figure represents an example of an acquired visible light image. Figure 3 (b) in the text indicates that... Figure 3 The diagram in (a) shows the instance segmentation result obtained after performing instance segmentation on the visible light image. Figure 3 (c) in the text indicates that according to Figure 3 (b) shows a schematic diagram of the instance mask generated from the instance segmentation result; Figure 4 (a) in the figure represents an example of an acquired visible light image. Figure 4 (b) in the figure is a schematic diagram showing the result after converting the visible light image shown in 4(a) into a single-channel image; Figure 5 (a) in the figure is a combination diagram showing an example of multiple single-band images corresponding to a multispectral image. Figure 5 Figure (b) in the figure is an example of a single-channel image. Figure 5 (c) in the text indicates that each is for Figure 5 The diagram in (a) shows multiple single-band edge images obtained by edge extraction of multiple single-band images. Figure 5 (d) in the text indicates that... Figure 5 A schematic diagram of a single-channel edge image obtained by edge extraction of the single-channel image shown in (b) of the image. Figure 6 (a) in the image is a comparison chart showing the overlap ratio between edge pixels in each single-band image and edge pixels in a single-channel image. Figure 6 (b) in the text indicates that according to Figure 6 Figure (a) shows an example of a selected single-band image with a determined overlap ratio. Figure 7 Figure (a) in the figure is an example of a single-channel image. Figure 7 Figure (b) in the figure shows an example of a selected single-band image. Figure 7 (c) in the text indicates that... Figure 7 (b) shows a schematic diagram of the spatially resolution-adjusted single-band image obtained after spatial resolution adjustment of the selected single-band image. Figure 7 (d) in the text indicates that according to Figure 7 The pixel value distribution of the single-channel image shown in (a) is a pair Figure 7 The single-band image with adjusted pixel value distribution is shown in (c) above. Figure 8 (a) in the image is a schematic diagram showing the feature points in a single-channel image and a single-band image after pixel value distribution adjustment. Figure 8 (b) in the diagram represents the candidate matching feature point pairs determined based on the feature points. Figure 8 (c) in the diagram represents the matching feature point pairs obtained after filtering the candidate matching feature point pairs. Figure 8 (d) in the figure represents the image overlay result after determining the initial coordinate transformation relationship based on the matching feature point pairs; Figure 9 (a) in the diagram is a schematic representation of an instance mask for a visible light image. Figure 9 (b) in the diagram represents the mapped instance mask obtained by mapping the instance mask to a multispectral image based on the pixel coordinate system of the selected single-band image according to the coordinate mapping relationship. Figure 9 (c) in the diagram represents a plurality of multispectral target regions formed according to the instance mask after the mapping. Figure 10 (a) in the diagram illustrates the selection of multiple sampling pixels from a multispectral target region. Figure 10 (b) in the diagram represents the material classification results corresponding to each sampled pixel. Figure 10 (c) in the diagram represents the determination of the plastic material category of the corresponding floating object instance based on multiple material classification results belonging to the same multispectral target region; Explanation of reference numerals in the attached figures: 100. River channel area; 110. UAV; 120. Visible light imaging device; 130. Multispectral imaging device; 140. Visible light image; 150. Multispectral image; 160. Floating object instance; 170. Instance mask; 180. Single-channel image; 180a. Single-channel edge image; 190. Single-band image; 190a. Single-band edge image; 190b. Selected single-band image; 190c. Single-band image after spatial resolution adjustment; 190d. Single-band image after pixel value distribution adjustment; 200. Feature point; 210. Matched feature point pair; 210a. Candidate matched feature point pair; 220. Mapped instance mask; 230. Multispectral target region; 240. Sampled pixel. Detailed Implementation

[0008] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only used to illustrate the technical solutions of the present invention and should not be construed as limiting the scope of protection of the present invention. Other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are all within the scope of protection of the present invention.

[0009] Current methods for classifying floating plastic debris in waterways typically involve first determining the location and outline of the debris using visible light images, then identifying the corresponding region from multispectral images, and extracting the reflectance of that region in different spectral bands to determine the plastic material of the debris.

[0010] The inventors discovered that when visible light images and multispectral images are acquired by different imaging devices, the two types of images differ in spatial resolution and pixel value distribution. Furthermore, different single-band images produce different imaging results for the same river scene. If the established image coordinate relationship is flawed, the outline of floating objects in the visible light image, after being mapped to the multispectral image, is easily deviated from the corresponding floating object's location. The extracted reflectance may also contain reflection information from the water surface background or nearby floating objects. Therefore, this invention provides an implementation that first determines each floating object instance and its outline range based on the visible light image. Then, it selects a single-band image from multiple single-band images of the multispectral image, adjusts the spatial resolution and pixel value distribution of the selected single-band image, and establishes a coordinate mapping relationship based on the visible light single-channel image and the single-band image with adjusted pixel value distribution. Based on this coordinate mapping relationship, a mask for each instance is mapped to the multispectral image, forming multispectral target regions corresponding to each floating object instance. Finally, the plastic material type of the corresponding floating object instance is determined based on the reflectance of multiple sampled pixels within each target region.

[0011] As one implementation method, see Figures 1 to 10 As shown, this invention provides a method for classifying floating plastic debris in river channels, comprising: acquiring visible light images and multispectral images corresponding to the same river area; segmenting the visible light images to obtain multiple floating debris instances and instance masks representing the contour range of each floating debris instance; converting the visible light images into single-channel images, and determining a selected single-band image from multiple single-band images corresponding to the multispectral images; adjusting the spatial resolution of at least one of the single-channel images and the selected single-band image, and adjusting the pixel value distribution of the selected single-band image according to the pixel value distribution of the single-channel image to obtain a single-band image with adjusted pixel value distribution; and performing pixel value distribution adjustment on the single-channel image and the single-band image with adjusted pixel value distribution. Feature matching is performed on the segmented images, and the feature matching results are transformed according to the pixel position correspondence before and after spatial resolution adjustment to obtain the coordinate mapping relationship between the visible light image and the multispectral image. According to the coordinate mapping relationship, each instance mask is mapped to the multispectral image to obtain the multispectral target region corresponding to each floating object instance. Multiple sampling pixels are selected from each multispectral target region, and the material classification result corresponding to each sampling pixel is obtained according to the reflectance of each sampling pixel in multiple spectral bands. According to the material classification result corresponding to each sampling pixel belonging to the same multispectral target region, the plastic material category of the floating object instance corresponding to the multispectral target region is determined.

[0012] Specifically, see Figure 1 As shown, the method for classifying floating plastic debris in waterways provided in this application includes steps S101 to S106; Step S101: Obtain the visible light image 140 and the multispectral image 150 corresponding to the same river area 100.

[0013] See Figure 2 As shown, the UAV 110 is equipped with a visible light imaging device 120 and a multispectral imaging device 130. The visible light imaging device 120 is used to acquire a visible light image 140 corresponding to the river area 100, and the multispectral imaging device 130 is used to acquire a multispectral image 150 corresponding to the river area 100. The visible light image 140 and the multispectral image 150 can be acquired during the same flight of the UAV 110, or they can be acquired separately during different acquisition processes. The two types of images contain the same river scene range and corresponding scene content, thus providing an image basis for subsequently establishing the coordinate mapping relationship between the two types of images. In one embodiment, the visible light imaging device 120 and the multispectral imaging device 130 are simultaneously installed on the UAV 110, and the visible light image 140 and the multispectral image 150 are acquired respectively while the UAV 110 maintains a predetermined flight attitude; in another embodiment, the visible light imaging device 120 and the multispectral imaging device 130 can also be used to perform fixed-point acquisition of the same river area 100.

[0014] It can be understood that the visible light image 140 and the multispectral image 150 correspond to the same river area 100, which means that the two types of images contain river scenes that can correspond to each other. It does not require that the two types of images have the same image size, spatial resolution or imaging range. In this way, it is beneficial to expand the applicable range of the visible light imaging device 120 and the multispectral imaging device 130, and to reserve enough scene overlap area for subsequent image processing.

[0015] Step S102: Perform instance segmentation on the visible light image 140 to obtain multiple floating object instances 160 and instance masks 170 that respectively characterize the contour range of each floating object instance 160.

[0016] Specifically, see Figure 3 (a) The visible light image 140 includes a water surface background and multiple floating objects distributed on the water surface; the visible light image 140 is segmented into instances to determine the pixel region corresponding to each floating object in the visible light image 140, and a single floating object that can be distinguished from other floating objects is determined as a floating object instance 160, resulting in... Figure 3 (b) shows the instance segmentation result. Based on the pixel region corresponding to each floating object instance 160, an instance mask 170 representing the contour range of each floating object instance 160 is generated, as shown below. Figure 3As shown in (c), the target pixels in instance mask 170 correspond to the pixel positions of the corresponding floating object instance 160 in the visible light image 140, and the pixels outside instance mask 170 correspond to the water background or other non-target areas. In this embodiment, different floating object instances 160 correspond to different instance masks 170, and each instance mask 170 retains the position and contour information of the corresponding floating object instance 160, so that different floating object instances 160 can be distinguished in subsequent processing.

[0017] It is understood that the instance mask 170 defines the target range according to the actual pixel contour of the floating object instance 160, rather than defining the target range according to the rectangular area around the floating object. In this way, it is beneficial to reduce the inclusion of water background and adjacent floating objects in the subsequent spectral data extraction area, thereby improving the correspondence between the extracted reflectance and the corresponding floating object instance 160.

[0018] Step S103: Convert the visible light image 140 into a single-channel image 180, and select a single-band image 190b from multiple single-band images 190 corresponding to the multispectral image 150; adjust the spatial resolution of at least one of the single-channel image 180 and the selected single-band image 190b, and adjust the pixel value distribution of the selected single-band image 190b according to the pixel value distribution of the single-channel image 180 to obtain a single-band image 190d with adjusted pixel value distribution; perform feature matching on the single-channel image 180 and the single-band image 190d with adjusted pixel value distribution, and convert the feature matching result according to the pixel position correspondence before and after spatial resolution adjustment to obtain the coordinate mapping relationship between the visible light image 140 and the multispectral image 150.

[0019] Specifically, see Figure 4 (a) and Figure 4 As shown in (b), the visible light image 140 containing multiple color channels is converted into a single-channel image 180, so that the scene outline, edge and texture information in the visible light image 140 is represented by a single pixel value; the single-channel image 180 serves as a reference image for subsequent comparison with each single-band image 190 and for feature matching with the single-band image 190d after pixel value distribution adjustment.

[0020] Multispectral image 150 includes multiple single-band images 190 corresponding to different spectral bands; in this embodiment, see... Figure 5As shown in (a), the multispectral image 150 includes six single-band images 190 with center wavelengths of 450nm, 555nm, 660nm, 720nm, 750nm, and 840nm, respectively. Different single-band images 190 exhibit different imaging responses to water bodies, floating objects, and scenes surrounding the river channel. Consequently, the outlines and edges of the same scene content differ across different single-band images 190. A comparison is made between the single-channel image 180 and the multiple single-band images 190, and a selected single-band image 190b with a higher degree of scene correspondence to the single-channel image 180 is determined from the multiple single-band images 190. In this embodiment, see... Figure 5 and Figure 6 As shown, a selected single-band image 190b is determined based on the correspondence between scene edges in single-channel image 180 and each single-band image 190; after determining the selected single-band image 190b, the spatial resolution of at least one of the single-channel image 180 and the selected single-band image 190b is adjusted; in this embodiment, see... Figure 7 (a) to Figure 7 As shown in (c), using the spatial resolution of single-channel image 180 as the processing benchmark, the spatial resolution of the selected single-band image 190b is adjusted to obtain the spatially resolution-adjusted single-band image 190c, ensuring that the same river scene has corresponding pixel scales in the two images. Based on the pixel value distribution of single-channel image 180, the pixel value distribution of the spatially resolution-adjusted single-band image 190c is adjusted to obtain... Figure 7 (d) shows the single-band image 190d after pixel value distribution adjustment; the single-band image 190d after pixel value distribution adjustment corresponds to the single-channel image 180 in terms of pixel value variation range and scene brightness distribution.

[0021] It is understandable that spatial resolution adjustment is used to reduce the pixel scale difference of the same scene content in two images, and pixel value distribution adjustment is used to reduce the pixel value response difference caused by different imaging methods. In this way, it is beneficial to improve the comparability of the same scene location in single-channel image 180 and single-band image 190d after pixel value distribution adjustment, which in turn is beneficial to improve the location correspondence accuracy of subsequent feature matching.

[0022] See Figure 8 As shown in (a), feature points 200 are extracted from the single-channel image 180 and the single-band image 190d after pixel value distribution adjustment, respectively. Based on the pixel change information around feature point 200 in the two images, candidate matching feature point pairs 210a are determined, and the candidate matching feature point pairs 210a are filtered to obtain... Figure 8(c) shows the matching feature point pairs 210; based on the pixel positions of each matching feature point pair 210 in the single-channel image 180 and the single-band image 190d after pixel value distribution adjustment, the initial coordinate transformation relationship from the pixel coordinates in the single-channel image 180 to the pixel coordinates in the single-band image 190d after pixel value distribution adjustment is determined. The matching feature point pairs 210 are further filtered according to feature differences and minimum point spacing to obtain matching feature point pairs used to solve the initial coordinate transformation relationship, such as... Figure 8 As shown in (d), based on the pixel positions of the selected matching feature points in the single-channel image 180 and the single-band image 190d after pixel value distribution adjustment, the initial coordinate transformation relationship is determined. Further, based on the pixel position correspondence of the selected single-band image 190b before and after spatial resolution adjustment, the initial coordinate transformation relationship is transformed to the pixel coordinate system of the selected single-band image 190b, resulting in a coordinate mapping relationship from the pixel coordinates in the visible light image 140 to the pixel coordinates in the multispectral image 150. It can be understood that the position of each pixel is not changed when the visible light image 140 is converted to the single-channel image 180, and the pixel coordinate system of the selected single-band image 190b is used as the reference coordinate system for the multispectral image 150. This method facilitates the transformation of the feature matching results obtained in the processed single-band image coordinate system to the original multispectral image coordinate system, ensuring that the mapping position of the instance mask 170 corresponds to the subsequent reflectance extraction position.

[0023] In step S104, according to the coordinate mapping relationship, each instance mask 170 is mapped from the pixel coordinate system of the visible light image 140 to the multispectral image 150 based on the pixel coordinate system of the selected single-band image 190b, so as to obtain the multispectral target region 230 corresponding to each floating object instance 160.

[0024] Specifically, see Figure 9 (a) Instance masks 170 are located in the pixel coordinate system of the visible light image 140, and the positions and contour ranges of the corresponding floating object instances 160 are recorded respectively; according to the coordinate mapping relationship, the pixel positions contained in each instance mask 170 are directly transformed into the multispectral image 150 based on the original pixel coordinate system of the selected single-band image 190b, to obtain... Figure 9 (b) shows the mapped instance mask 220.

[0025] During the mapping process, the correspondence between each instance mask 170 and the corresponding floating object instance 160 is preserved, so that different instance masks 170 still correspond to different floating object instances 160 after mapping; based on the pixel positions belonging to the same floating object instance 160 in the mapped instance mask 220, a corresponding multispectral target region 230 is formed in the multispectral image 150, such as... Figure 9As shown in (c), each multispectral target region 230 corresponds to a floating object instance 160, and the spatial region for subsequent reflectance extraction is defined according to the contour range of the corresponding mapped instance mask 220. It can be understood that the multispectral target region 230 is determined by the coordinate mapping result of the instance mask 170, and its contour corresponds to the contour of the corresponding floating object instance 160. This approach helps to reduce the inclusion of water surface background and reflection information from nearby floating objects into the corresponding multispectral target region 230, thereby improving the target orientation of reflectance data within the multispectral target region 230.

[0026] Step S105: Select multiple sampling pixels 240 from each multispectral target region 230, and obtain the material classification result corresponding to each sampling pixel 240 based on the reflectance of each sampling pixel 240 in multiple spectral bands.

[0027] Specifically, see Figure 10 (a) The pixels contained in each multispectral target region 230 are used as the sampling range of the corresponding floating object instance 160, and multiple sampling pixels 240 are selected from each multispectral target region 230; each sampling pixel 240 maintains a correspondence with its respective multispectral target region 230. For any sampling pixel 240, the reflectance of the sampling pixel 240 in multiple spectral bands is extracted to obtain the multi-band reflectance data corresponding to the sampling pixel 240; the material classification result of the corresponding sampling pixel 240 is determined based on the multi-band reflectance data, such as... Figure 10 As shown in (b).

[0028] In this embodiment, corresponding spectral features are generated based on the reflectance of each sampled pixel 240 in multiple spectral bands, and these spectral features are input into a pre-trained plastic material classification model. The plastic material classification model then outputs the material classification result corresponding to each sampled pixel 240. It can be understood that the multispectral target region 230 corresponding to the same floating object instance 160 includes multiple sampled pixels 240, each reflecting the spectral response at different locations of the floating object instance 160. This approach helps reduce classification bias caused by local reflections, water surface residue, or imaging noise affecting a single pixel, and provides multiple classification results for determining the plastic material category of the floating object instance 160.

[0029] Step S106: Based on the material classification results corresponding to each sampled pixel 240 belonging to the same multispectral target region 230, determine the plastic material category of the floating object instance 160 corresponding to the multispectral target region 230.

[0030] Specifically, according to the multispectral target region 230 to which each sampled pixel 240 belongs, the material classification results corresponding to each sampled pixel 240 are grouped, so that multiple material classification results belonging to the same floating object instance 160 are grouped into the same group; the plastic material category corresponding to the group is determined based on the multiple material classification results in the same group, and the determined plastic material category is used as the plastic material category of the corresponding floating object instance 160, thus obtaining... Figure 10 (c) shows the classification results. For multiple floating object instances 160 present in the visible light image 140, the plastic material category corresponding to each floating object instance 160 is determined, thereby obtaining the position, outline range and plastic material category of each floating object instance 160.

[0031] It can be understood that the material classification result corresponding to the sampled pixel 240 is a pixel-level classification result, and the same multispectral target region 230 corresponds to the same floating object instance 160. In this way, the material classification results corresponding to multiple sampled pixels 240 within the same multispectral target region 230 are combined, which helps to convert the pixel-level classification results into plastic material categories at the level of floating object instance 160, thereby improving the completeness and stability of the classification results of plastic floating objects in the river.

[0032] Through steps S101 to S106, firstly, instance masks 170 representing the contour range of each floating object instance 160 are obtained based on the visible light image 140. Then, the spatial resolution and pixel value distribution of the selected single-band image 190b are adjusted to obtain the single-band image 190d after pixel value distribution adjustment. Feature matching is performed on the single-channel image 180 and the single-band image 190d after pixel value distribution adjustment. The feature matching result is transformed according to the pixel position correspondence of the selected single-band image 190b before and after spatial resolution adjustment to obtain the coordinate mapping relationship between the visible light image 140 and the multispectral image 150. Based on the coordinate mapping relationship, multispectral target regions 230 corresponding to each floating object instance 160 are formed. The plastic material category of the corresponding floating object instance 160 is determined according to the reflectance of multiple sampled pixels 240 in each multispectral target region 230. This approach helps reduce the positional bias caused when instance masks are mapped to multispectral images, reduces the interference of non-target area reflection information on the extracted reflectance, and thus helps improve the accuracy of river plastic floating debris classification.

[0033] Furthermore, as one implementation method, a selected single-band image 190b is determined from multiple single-band images 190 corresponding to the multispectral image 150, including: performing edge extraction on the same scene range in the single-channel image 180 and each single-band image 190 respectively to obtain a single-channel edge image 180a and a single-band edge image 190a corresponding to each single-band image 190; determining the overlap ratio between the edge pixels in each single-band edge image 190a and the edge pixels in the single-channel edge image 180a respectively; and determining the single-band image 190 corresponding to the single-band edge image 190a with the largest overlap ratio as the selected single-band image 190b.

[0034] Specifically, see Figure 5 (a) The multispectral image 150 includes multiple single-band images 190 corresponding to multiple spectral bands. In this embodiment, the center wavelengths of the multiple single-band images 190 are 450nm, 555nm, 660nm, 720nm, 750nm, and 840nm, respectively, and each single-band image 190 contains a river scene range corresponding to the single-channel image 180. The same scene range is determined from the single-channel image 180 and each single-band image 190, and edge extraction is performed on the same scene range. During the edge extraction process, edge pixels are determined based on the pixel value changes between each pixel and its neighboring pixels, thus obtaining... Figure 5 (d) shows the single-channel edge image 180a and Figure 5 (c) shows multiple single-band edge images 190a. It should be noted that for any edge pixel in a single-channel edge image 180a, if there is an edge pixel in the corresponding single-band edge image 190a with the same relative position or a relative position difference not exceeding a preset threshold, then the two are determined to overlap. It can be understood that by comparing the relative positions of each edge pixel, the determination of the degree of edge pixel overlap does not depend on the exact number of pixels in each image, thereby enabling band selection for single-channel images 180 and single-band images 190 acquired at different spatial resolutions.

[0035] In one embodiment, the overlap ratio of edge pixels corresponding to the i-th single-band edge image 190a is determined according to the following formula: Where, N i N represents the number of edge pixels in the i-th single-band edge image 190a that overlap with the edge pixels of the single-channel edge image 180a; v This represents the total number of edge pixels in the single-channel edge image 180a within the overlapping range.

[0036] Using the same edge extraction method and overlap judgment criteria, the edge pixel overlap ratios corresponding to the 450nm, 555nm, 660nm, 720nm, 750nm, and 840nm single-band images 190 were determined respectively. In this embodiment, the edge pixel overlap ratios corresponding to the six single-band images 190 are 61%, 68%, 65%, 72%, 77%, and 86%, respectively. For specific comparison results, please refer to [link to relevant documentation]. Figure 6 As shown in (a), the edge pixel overlap ratio of the single-band image with a center wavelength of 840nm is the largest at 86%. Therefore, the single-band image 190 with a center wavelength of 840nm is selected as the single-band image 190b. Figure 6 As shown in (b).

[0037] It is understandable that the overlap ratio of edge pixels is used to characterize the degree to which each single-band image and single-channel image jointly present the contour of the same river scene. The larger the overlap ratio, the higher the positional correspondence between the scene contour in the corresponding single-band image and the scene contour in the single-channel image. In this way, it is beneficial to determine the selected single-band image 190b suitable for subsequent processing from multiple single-band images 190, thereby improving the positional correspondence accuracy of subsequent matching feature point pairs.

[0038] Furthermore, as one implementation method, the spatial resolution of at least one of the single-channel image 180 and the selected single-band image 190b is adjusted, and the pixel value distribution of the selected single-band image 190b is adjusted according to the pixel value distribution of the single-channel image 180. This includes: redetermining the position and pixel value of each pixel in the selected single-band image 190b according to the spatial resolution of the single-channel image 180, to obtain a single-band image 190c with adjusted spatial resolution; determining the cumulative pixel ratio corresponding to different pixel values ​​in the single-channel image 180 and the single-band image 190c with adjusted spatial resolution, respectively; and adjusting the pixel value of each pixel in the single-band image 190c with adjusted spatial resolution according to the pixel values ​​corresponding to the same cumulative pixel ratio in the single-channel image 180 and the single-band image 190c with adjusted spatial resolution, to obtain a single-band image 190d with adjusted pixel value distribution.

[0039] Specifically, see Figure 7 (a) and Figure 7 (b) In this embodiment, a single-channel image 180 is used as a reference for image processing, and the spatial resolution of the single-channel image 180 is [missing information]. ρ v The spatial resolution of the selected single-band image 190b determined in step S103 is: ρ m ;according to ρ v and ρ mThe pixel grid of the selected single-band image 190b is re-established according to the ratio between them, so that the actual scene distances corresponding to adjacent pixels in the new pixel grid are the same as the actual scene distances corresponding to adjacent pixels in the single-channel image 180.

[0040] For any pixel position in the new pixel grid ( u ′, v The position of the pixel in the selected single-band image 190b is determined according to the following formula (′), u , v ): When the corresponding position ( u , v When a pixel is located between multiple pixels in a selected single-band image 190b, the new pixel position is determined based on the pixel values ​​of the surrounding pixels at the corresponding position. u ′, v The pixel value of the new pixel position is obtained by weighting the distance between each neighboring pixel and the corresponding position according to the distance between each neighboring pixel and the corresponding position. This completes the bilinear interpolation processing of the selected single-band image 190b. Figure 7 (c) shows the spatial resolution adjusted single-band image 190c.

[0041] Based on the pixel grid size and pixel starting position of the selected single-band image 190b and the spatially resolution-adjusted single-band image 190c, determine the pixel position correspondence between the two. Let the homogeneous pixel coordinates in the selected single-band image 190b be... ρ m The corresponding homogeneous pixel coordinates in the single-band image 190c after spatial resolution adjustment are: ρ r Then the two satisfy: in, S The second coordinate transformation relationship represents the transformation from pixel coordinates in the selected single-band image 190b to pixel coordinates in the single-band image 190c after spatial resolution adjustment. When the spatial resolution adjustment only includes pixel grid scaling, the scale parameter in the second coordinate transformation relationship is determined according to the pixel grid size of the images before and after adjustment. When the spatial resolution adjustment also includes scene cropping, the second coordinate transformation relationship also includes the corresponding pixel position offset parameter.

[0042] It is understandable that the pixel value distribution adjustment only changes the pixel value of each pixel in the single-band image 190c after spatial resolution adjustment, without changing the position of each pixel. Therefore, the single-band image 190c after spatial resolution adjustment and the single-band image 190d after pixel value distribution adjustment use the same pixel coordinate system. In this way, it is beneficial to convert the feature matching results obtained in the coordinate system of the single-band image 190d after pixel value distribution adjustment to the original pixel coordinate system of the selected single-band image 190b based on the pixel position correspondence before and after spatial resolution adjustment.

[0043] In this embodiment, the spatial resolution of a single-channel image is 180. ρ v Approximately 2cm, with a selected spatial resolution of 190b for a single-band image. ρ m The pixel size of the selected single-band image 190b is 839×457. After re-determining the pixel position and pixel value, the pixel size of the single-band image 190c after spatial resolution adjustment is 1024×558, and its spatial resolution is basically consistent with that of the single-channel image 180. The spatial resolution adjustment in this embodiment does not include scene cropping. The horizontal scale parameter and the vertical scale parameter in the second coordinate transformation relationship are determined according to the actual pixel grid size before and after adjustment.

[0044] Specifically, the horizontal scale parameter is determined based on the ratio of 1024 to 839, and the vertical scale parameter is determined based on the ratio of 558 to 457; the spatial resolutions of approximately 2cm and approximately 2.4cm are used to describe the actual scene scale of the two types of images, and the scale parameter used for coordinate transformation is based on the actual pixel grid size before and after adjustment.

[0045] It is understandable that spatial resolution represents the actual scene range corresponding to a single pixel in an image. Adjusting spatial resolution is not just about scaling the image display size, but about redetermining the pixel position and pixel value of the selected single-band image 190b based on the spatial resolution of the two types of images. In this way, it is beneficial to reduce the pixel scale difference of the same river scene in single-channel image 180 and single-band image 190.

[0046] After spatial resolution adjustment, the number of pixels corresponding to each pixel value in both the single-channel image 180 and the spatially resolution-adjusted single-band image 190c is counted. The cumulative pixel ratio corresponding to different pixel values ​​is then determined according to the pixel values ​​in ascending order. For pixel value g in single-channel image 180, its cumulative pixel ratio F... v (g) Determined according to the following formula: For the pixel value g in the single-band image 190c after spatial resolution adjustment, its cumulative pixel ratio Fm(g) is determined according to the following formula: Among them, h v (k) represents the number of pixels with a value of k in the single-channel image 180, and hm(k) represents the number of pixels with a value of k in the single-band image 190c after spatial resolution adjustment; N V and N m These represent the total number of pixels in single-channel image 180 and single-band image 190c after spatial resolution adjustment, respectively; g min This represents the minimum pixel value in the corresponding image.

[0047] For any pixel value g in the single-band image 190c after spatial resolution adjustment m Determine the target pixel value g from the pixel values ​​of a single-channel image of 180. v Make the target pixel value g v The corresponding cumulative pixel ratio and pixel value g m The difference between the corresponding cumulative pixel ratios is the smallest, that is: The pixel value in the 190c single-band image after spatial resolution adjustment is g. m The pixel value is adjusted to the target pixel value g. v And in the same manner, adjust the other pixel values ​​in the single-band image 190c after spatial resolution adjustment to obtain... Figure 7 (d) shows the single-band image 190d after pixel value distribution adjustment; in this embodiment, the pixel value correspondence between the single-band image 190c with spatial resolution adjustment and the single-channel image 180 is established according to this processing method, and the cumulative pixel ratio distribution results of the two images after processing are F v (64) = 0.096, F v (128) = 0.870, F v (192) = 0.977, the single-band image 190d after pixel value distribution adjustment is F m (64) = 0.096, F m (128) = 0.876, F m (192) = 0.976.

[0048] Using pixel values ​​g=0,1,…,255 as independent variables, the mean absolute deviation is used to characterize the difference in pixel value distribution between a single-channel image 180 and a single-band image. The mean absolute deviation is determined according to the following formula: in, This represents the cumulative pixel proportion corresponding to pixel value g in a single-channel image 180; when calculating the average absolute deviation of pixel value distribution before adjustment... This represents the cumulative pixel proportion corresponding to pixel value g in the single-band image 190c after spatial resolution adjustment; when calculating the average absolute deviation of pixel value distribution after adjustment... This represents the cumulative pixel ratio corresponding to pixel value g in the single-band image 190d after pixel value distribution adjustment. The smaller the average absolute deviation, the closer the cumulative pixel ratio distributions of the two images are. In this embodiment, the average absolute deviation before pixel value distribution adjustment is approximately 14.30%, and the average absolute deviation after pixel value distribution adjustment is approximately 0.45%.

[0049] It can be understood that the pixels corresponding to the single-channel image 180 and the spatially resolution-adjusted single-band image 190c under the same cumulative pixel ratio represent scene pixels with similar statistical distribution positions in the two images respectively; the pixel values ​​of the spatially resolution-adjusted single-band image 190c are adjusted according to the cumulative pixel ratio of the single-channel image 180, so that the pixel value distribution of the pixel value distribution of the single-band image 190d after pixel value distribution adjustment is transformed into the pixel value distribution of the single-channel image 180.

[0050] The single-band image 190d, with adjusted pixel value distribution, is used only for feature matching with the single-channel image 180 and to determine the initial coordinate transformation relationship. After transforming the initial coordinate transformation relationship according to the pixel position correspondence before and after spatial resolution adjustment, the instance mask 170 is mapped to the pixel coordinate system of the original multispectral image 150. The reflectance of the sampled pixels 240 within the multispectral target region 230 is extracted from the original multispectral image 150 or the radiometrically corrected multispectral image 150, without using the single-band image 190c with adjusted spatial resolution or the single-band image 190d with adjusted pixel value distribution as the reflectance data source.

[0051] It is understood that both spatial resolution adjustment and pixel value distribution adjustment serve image feature matching, and the spectral data in the original multispectral image 150 are not changed by the processing. In this way, it is beneficial to maintain the original spectral characteristics of the reflectance data corresponding to each sampled pixel 240 while improving the accuracy of feature matching position correspondence.

[0052] This approach helps to reduce the differences in pixel scale and pixel value response between the single-channel image 180 and the single-band image 190d after pixel value distribution adjustment. This allows the riverbank outline, floating object outline, and other corresponding structures in the river scene to have similar scale and pixel value variation characteristics in the two images, thereby improving the accuracy of subsequent feature point extraction and matching.

[0053] Furthermore, as one implementation method, feature matching is performed on the single-channel image 180 and the single-band image 190d after pixel value distribution adjustment, including: detecting candidate feature points in the single-channel image 180 and the single-band image 190d after pixel value distribution adjustment at multiple image scales respectively; based on the pixel value changes in the neighborhood of each candidate feature point, removing candidate feature points with pixel value contrast lower than a preset value and candidate feature points located on a single edge, to obtain feature points 200 corresponding to the single-channel image 180 and the single-band image 190d after pixel value distribution adjustment respectively; generating feature description data corresponding to each feature point 200 based on the gradient direction and gradient magnitude of the pixels in the neighborhood of each feature point 200; and determining candidate matching feature point pairs 210a based on the feature description data corresponding to the feature points 200 in the single-channel image 180 and the feature points 200 in the single-band image 190d after pixel value distribution adjustment.

[0054] Specifically, see Figure 8 (a) Single-channel image 180d and single-band image 190d after pixel value distribution adjustment are used as the images to be detected. In this embodiment, the images to be detected are smoothed and scaled stepwise to construct multiple image scales, and candidate feature points are determined by the pixel value differences between adjacent image scales. Let the image to be detected be I(x,y), and its scale space image L(x,y,σ) at scale σ and the difference image D(x,y,σ) between adjacent scale space images are determined according to the following formulas: in, The scale parameter is represented as Gaussian function, This represents the convolution operation. Indicates the scale ratio between adjacent image scales; [This refers to] the difference image. Any pixel in the image is compared with its neighboring pixels at its own image scale and adjacent image scales. When the pixel is the maximum or minimum value in the neighborhood, it is determined as a candidate feature point.

[0055] For each candidate feature point, a local fitting is performed on the scale-space response within its neighborhood, and the response with an absolute value lower than the contrast threshold is selected. T c Candidate feature points are eliminated; simultaneously, a Hessian matrix is ​​constructed based on the changes in the second-order pixel values ​​at the locations of the candidate feature points. H D : in, , and These represent the second-order changes of the difference image in the corresponding directions; According to the Hessian matrix The trace and determinant determine the edge response ratio : When the edge response ratio When the pixel value exceeds a preset edge threshold, it indicates that the pixel value change of the candidate feature point in one direction is significantly greater than that in the other direction, and its position is prone to shift along a single edge. Therefore, the candidate feature point is removed. After contrast filtering and edge response filtering, feature points 200 corresponding to the single-channel image 180 and the single-band image 190d with adjusted pixel value distribution are obtained respectively. In this embodiment, the number of feature points retained in the single-channel image 180 is 320, and the number of feature points retained in the single-band image 190d with adjusted pixel value distribution is 320. The specific extraction results are as follows: Figure 8 As shown in (a).

[0056] For each feature point 200, the gradient magnitude and gradient direction of each pixel in the neighborhood of feature point 200 are determined, and the principal gradient direction of feature point 200 is used as the neighborhood direction reference. In this embodiment, a 16×16 pixel neighborhood is selected centered on feature point 200, and this neighborhood is divided into 4×4 sub-regions. The gradient magnitude corresponding to 8 directional intervals is counted in each sub-region, thereby forming 128-dimensional feature description data. It can be understood that the feature description data is used to characterize the local pixel changes of the riverbank boundary, fixed structures on the bank, and other stable scene structures around feature point 200, and the principal gradient direction is used as the direction reference. In this way, it is beneficial to reduce the impact of scale differences, directional differences, and local pixel value changes between the two types of images on the correspondence of feature points.

[0057] Based on the feature description data corresponding to each feature point 200 in the single-channel image 180 and the single-band image 190d after pixel value distribution adjustment, the feature difference between feature points 200 in the two images is calculated, and two feature points 200 whose feature differences meet the preset matching conditions are combined into candidate matching feature point pairs 210a. The candidate matching results are as follows. Figure 8 As shown in (b), this approach helps to establish an initial positional correspondence by utilizing the local structures in the river scene that can be stably presented in both types of images, providing a data foundation for subsequent removal of incorrect matches and determination of coordinate mapping relationships.

[0058] Further, as one implementation, determining the matching feature point pair 210 from the candidate matching feature point pair 210a includes: for each feature point 200 in the single-channel image 180, determining the first feature difference between the feature point 200 and the first candidate feature point in the single-band image 190d after pixel value distribution adjustment, and the second feature difference between the feature point 200 and the second candidate feature point, wherein the first feature difference is less than the second feature difference; calculating the ratio between the first feature difference and the second feature difference, and eliminating candidate matching feature point pairs 210 whose ratio does not meet a preset condition. a; Select a preset number of candidate matching feature point pairs 210a from the remaining candidate matching feature point pairs 210a multiple times, and determine the candidate coordinate transformation relationship according to the candidate matching feature point pairs 210a selected each time; determine the position error of the remaining candidate matching feature point pairs 210a after transformation by each candidate coordinate transformation relationship, and count the number of candidate matching feature point pairs 210a with position errors less than the preset error; determine the candidate matching feature point pair 210a with the largest number of candidate coordinate transformation relationships and position errors less than the preset error as the matching feature point pair 210.

[0059] Specifically, for the first 180 in a single-channel image There are 200 feature points, and their feature description data is . The 190d single-band image after pixel value distribution adjustment The feature description data corresponding to each feature point 200 is: The differences in characteristics between the two Determined according to the following formula: in, This indicates the dimension of the feature description data. and These respectively represent the first in the corresponding feature description data. Each component; according to the ascending order of feature differences, the first candidate feature point and the second candidate feature point are determined from the single-band image 190d after pixel value distribution adjustment, and the corresponding first feature difference and second feature difference are denoted as follows: and The ratio between the two Determined according to the following formula: when Less than the preset ratio threshold At that time, retain the first 180th image in a single channel. A candidate matching feature point pair 210a is formed by the feature point 200 and the first candidate feature point; when Not less than the preset ratio threshold At that time, the corresponding candidate matching feature point pair 210a is removed. In this embodiment, a preset ratio threshold is used. The ratio is 0.78. After ratio screening, 123 candidate matching feature point pairs (210a) are retained.

[0060] Even after ratio screening, candidate matching feature point pairs 210a may still contain positional errors caused by local texture similarity. Therefore, four candidate matching feature point pairs 210a are randomly selected from the remaining candidate matching feature point pairs 210a, and the candidate coordinate transformation relationship is determined based on the selected candidate matching feature point pairs 210a. For the remaining candidate matching feature point pairs 210a, the first... Pairs of points, based on candidate coordinate transformation relationships The location of feature points in a single-channel image 180 The predicted location is obtained by transforming the image to the pixel coordinate system of the single-band image 190d after pixel value distribution adjustment. Its positional error Determined according to the following formula: in, This indicates the actual location of the feature points in the candidate matching feature point pair 210a in the single-band image 190d after pixel value distribution adjustment. This represents converting homogeneous coordinates to two-dimensional pixel coordinates; when the position error... Less than the preset error threshold At that time, the corresponding candidate matching feature point pair 210a is denoted as the candidate coordinate transformation relationship. The corresponding valid point pairs.

[0061] The process of selecting candidate matching feature point pairs 210a, determining candidate coordinate transformation relationships, and counting the number of valid point pairs is repeated, and the candidate coordinate transformation relationship with the largest number of valid point pairs is selected. Candidate matching feature point pairs 210a with a position error less than a preset error threshold Te under this candidate coordinate transformation relationship are determined as matching feature point pairs 210. In this embodiment, the preset error threshold Te is 3 pixels. After filtering, 28 sets of matching feature point pairs 210 are obtained, as shown in the following figure. Figure 8 As shown in (c).

[0062] Furthermore, the 28 sets of matching feature point pairs 210 are sorted from smallest to largest according to the feature differences between their corresponding feature description data, and a minimum point spacing is set between the matching feature point pairs. d min The minimum point spacing is determined according to the following formula: in, H and WThese represent the height and width of the image to be matched, respectively. β This represents the dot spacing adjustment coefficient, in this embodiment, β Approximately 15. Matching feature point pairs 210 are selected sequentially according to the sorting results, and matching feature point pairs 210 that satisfy the minimum point spacing constraint are retained until there are no more matching feature point pairs 210 that can be retained. In this embodiment, after this processing, 18 sets of matching feature point pairs 210 are obtained, and the initial coordinate transformation relationship is solved using the 18 sets of matching feature point pairs 210.

[0063] It is understandable that ratio filtering is used to eliminate fuzzy matches where the first and second candidate feature points are similar in degree, and the statistical analysis of effective point pairs corresponding to the candidate coordinate transformation relationship is used to eliminate positional errors that do not conform to the overall spatial transformation law of the river scene. In this way, it is beneficial to ensure that the retained matching feature point pairs 210 simultaneously satisfy the local feature similarity condition and the overall positional transformation condition, thereby improving the accuracy of solving the coordinate mapping relationship.

[0064] Further, as one implementation method, obtaining the coordinate mapping relationship based on the matching feature point pairs 210 includes: obtaining the pixel coordinates of each matching feature point pair 210 in the single-channel image 180 and the single-band image 190d after pixel value distribution adjustment; establishing a coordinate transformation model from the pixel coordinates in the single-channel image 180 to the pixel coordinates in the single-band image 190d after pixel value distribution adjustment based on the pixel coordinates of each matching feature point pair 210; adjusting the parameters of the coordinate transformation model so that the sum of the position errors of each matching feature point pair 210 after coordinate transformation meets a preset condition; determining the coordinate transformation model after parameter adjustment as the initial coordinate transformation relationship; and transforming the initial coordinate transformation relationship to the pixel coordinate system of the selected single-band image 190b and the single-band image 190c after spatial resolution adjustment based on the position correspondence of each pixel, thereby obtaining the coordinate mapping relationship from the pixel coordinates in the visible light image 140 to the pixel coordinates in the multispectral image 150.

[0065] Specifically, let the first The homogeneous pixel coordinates of the matched feature point pair 210 in the single-channel image 180 are: The homogeneous pixel coordinates in the single-band image 190d after pixel value distribution adjustment are: In this embodiment, a homography coordinate transformation model is used to represent the coordinate relationship between the two images: in, This represents a 3×3 homography transformation matrix. Representing the proportional relationship in homogeneous coordinates; transforming the homography matrix Represented as: The pixel coordinates in a single-channel image 180 Predicted pixel coordinates after conversion to 190d single-band image with pixel value distribution adjustment satisfy: Substitute the pixel coordinates of each matching feature point pair of 210 into the coordinate transformation model to establish a homography transformation matrix. The linear equations of each parameter are obtained, and the linear equations corresponding to multiple matching feature point pairs 210 are combined to form ;in, The homography transformation matrix The parameters in the matrix are arranged in a preset order to form a parameter vector. (Regarding the coefficient matrix...) Perform singular value decomposition and select the right singular vector corresponding to the minimum singular value as the parameter vector. The initial solution is obtained, and the homography transformation matrix is ​​further adjusted according to the position error of each matching feature point to 210. , so that the objective function To obtain the minimum value: Homography transformation matrix after parameter adjustment H This serves as the initial coordinate transformation relationship for converting pixel coordinates in the single-channel image 180 to pixel coordinates in the single-band image 190d after pixel value distribution adjustment. Let the homogeneous pixel coordinates in the single-channel image 180 be... p v The corresponding homogeneous pixel coordinates in the 190d single-band image after pixel value distribution adjustment are: p r Then the two satisfy: Since the single-band image 190c with adjusted spatial resolution and the single-band image 190d with adjusted pixel value distribution use the same pixel coordinate system, based on the pixel position correspondence between the selected single-band image 190b and the single-band image 190c with adjusted spatial resolution, the corresponding homogeneous pixel coordinates in single-band image 190b are selected. p m satisfy: Therefore, by combining the inverse transformation of the pixel position correspondence before and after spatial resolution adjustment with the initial coordinate transformation relationship, we obtain: in, TThis represents the coordinate mapping relationship from pixel coordinates in visible light image 140 to pixel coordinates in multispectral image 150. The pixel coordinate system of single-band image 190b is selected as the reference coordinate system of multispectral image 150. In this embodiment, the initial coordinate transformation relationship is obtained using 18 sets of matching feature point pairs 210. The average position error of the matching feature point pairs 210 after transformation by the initial coordinate transformation relationship is 0.45 pixels, and the root mean square position error is 0.57 pixels. The 18 sets of matching feature point pairs 210 are as follows: Figure 8 As shown in (d).

[0066] It can be understood that the initial coordinate transformation relationship is used to represent the positional correspondence between the single-channel image 180 and the single-band image 190d after pixel value distribution adjustment. After transforming the initial coordinate transformation relationship according to the pixel positional correspondence of the selected single-band image 190b before and after spatial resolution adjustment, the resulting coordinate mapping relationship is used to represent the positional correspondence between the visible light image 140 and the multispectral image 150 based on the pixel coordinate system of the selected single-band image 190b. In this way, it is beneficial to transform the matching result in the processed single-band image coordinate system to the original multispectral image coordinate system, thereby helping to keep the mapping position of the instance mask 170 corresponding to the subsequent reflectance extraction position. When the spatial change between the two images can be characterized by affine transformation, the coordinate transformation model can also be set as an affine transformation model, and the corresponding parameters can be solved based on the matching feature points 210.

[0067] Furthermore, as one implementation, each instance mask 170 is mapped to a multispectral image 150 based on the pixel coordinate system of the selected single-band image 190b according to the coordinate mapping relationship. This includes: setting different instance labels for the instance masks 170 corresponding to different floating object instances 160, and setting background labels for the areas outside the instance masks 170; determining the output area after each instance mask 170 is mapped to the multispectral image 150; for each pixel in the output area, determining the corresponding position of the pixel in the instance mask 170 according to the coordinate mapping relationship; selecting the instance mask pixel closest to the corresponding position from the instance mask pixels around the corresponding position, and assigning the instance label or background label corresponding to the selected instance mask pixel to the corresponding pixel in the output area; and forming a multispectral target area 230 corresponding to each floating object instance 160 according to the pixels assigned different instance labels.

[0068] Specifically, see Figure 9 (a) Combining the instance masks 170 to form a single-channel shaping mask corresponding to the size of the visible light image 140. The label value of the background pixel is set to 0, the first... The label value of the instance mask pixel corresponding to each floating object instance 160 is set to This results in pixels within the same instance mask 170 having the same label, while pixels within different instance masks 170 have different labels.

[0069] Based on the coordinate mapping relationship T, the boundary corners of the mask Mv are transformed into the multispectral image 150, which is based on the pixel coordinate system of the selected single-band image 190b. The output area of ​​the mapped instance mask 220 is determined according to the transformed boundary range and the image range of the multispectral image 150. For any pixel in the output area... According to the coordinate mapping relationship T The inverse transform determines the pixel in the mask. M v The corresponding position in : Corresponding position When the coordinates are non-integer, the integer pixel positions are determined according to the nearest neighbor principle. : When integer pixel position Located in the mask When the location is within the effective range, assign the label value corresponding to that location to the pixel in the mapped instance mask 220. When the integer pixel position When the pixel is outside the valid range, Assign the value to background label 0, that is: in, This indicates that the instance mask after mapping is 220. Indicates mask The effective pixel range; after traversing each pixel within the output area, we obtain... Figure 9 (b) shows the mapped instance mask 220.

[0070] It is understood that instance labels are used to represent the floating object instance 160 to which a pixel belongs, and there is no continuous numerical relationship between label values. Using the nearest neighbor method to obtain labels from the instance mask 170 avoids the generation of new label values ​​between adjacent instance labels by continuous numerical interpolation methods such as bilinear interpolation. This approach helps maintain the label boundaries and correspondences between different floating object instances 160 during coordinate transformation, thereby facilitating the formation of the multispectral target region 230 corresponding to each floating object instance.

[0071] Based on the pixels with the same instance label in the mapped instance mask 220, a corresponding multispectral target region 230 is formed; The multispectral target region corresponding to 160 floating object instances Determined according to the following formula: In this embodiment, the visible light image 140 includes 17 floating object instances 160, which are mapped to form 17 multispectral target regions 230. Each multispectral target region 230 maintains a correspondence with its corresponding floating object instance 160, as shown in the specific results. Figure 9 As shown in (c), this method helps to define the extraction location of multispectral reflectance according to the outline range of the floating object instance, reducing the mixing of water surface background and adjacent floating object reflectance information into the corresponding multispectral target area 230.

[0072] Furthermore, as an implementation method, before selecting multiple sampling pixels 240 from each multispectral target region 230, the single-band images corresponding to different spectral bands in the multispectral image 150 are registered to the same spatial coordinate system; the known reflectance of the reference plate and the pixel value of the reference plate in each spectral band are obtained; based on the known reflectance of the reference plate, the pixel value of the reference plate in the corresponding spectral band, and the pixel value of the pixel to be tested in the corresponding spectral band, the reflectance of the pixel to be tested in the corresponding spectral band is determined; multiple sampling pixels 240 are randomly selected from the pixels contained in each multispectral target region 230, and each sampling pixel 240 is associated with its respective multispectral target region 230.

[0073] Specifically, the multispectral imaging device 130 acquires single-band images corresponding to multiple spectral bands. In this embodiment, the multispectral image 150 includes six single-band images with center wavelengths of 450nm, 555nm, 660nm, 720nm, 750nm, and 840nm, respectively. The original pixel coordinate system of the selected single-band image 190b is used as the reference coordinate system of the multispectral image 150, wherein the center wavelength of the selected single-band image 190b is 840nm. According to the internal calibration parameters of the multispectral imaging device 130 or the positional correspondence of stable scene structures in each single-band image, the remaining single-band images are registered to the original pixel coordinate system of the selected single-band image 190b, and the registered single-band images are combined according to spectral bands to form a multi-band image, so that the same river scene location corresponds to the same pixel position in the six single-band images.

[0074] The mapped instance mask 220, the multispectral target region 230, and the multiband image adopt the same pixel coordinate system based on the original pixel coordinate system of the selected single-band image 190b. Therefore, the reflectance of the same river scene location in multiple spectral bands can be extracted based on the position of each sampled pixel 240 in the multispectral target region 230.

[0075] It can be understood that the coordinate mapping relationship is used to map the instance mask 170 to the reference coordinate system of the multispectral image 150, and the band registration is used to make the other single-band images correspond to the reference coordinate system. In this way, it is beneficial to keep the instance mask mapping result, the sampling pixel position and the reflectance under multiple spectral bands spatially corresponding.

[0076] Obtain the pixel values ​​of a reference plate with known reflectivity in each spectral band; for the ... Given a spectral band, let the pixel value of the pixel to be measured or the radiance value after dark current correction be [value]. The pixel values ​​of the reference board or the radiance value after dark current correction are... The known reflectance of the reference plate in this spectral band is The reflectance of the pixel under test in this spectral band is... Determined according to the following formula: Reflectance conversion is performed on each of the six spectral bands, so that any valid pixel location in the registered multispectral image 150 corresponds to the reflectance under each of the six spectral bands. It can be understood that band registration is used to ensure that the reflectance under different spectral bands originates from the same river scene location, and reference plate correction is used to convert pixel values ​​affected by the imaging device response and the acquired illumination into comparable reflectance. In this way, it is beneficial to improve the spatial consistency between the multi-band reflectances corresponding to the same sampled pixel and improve the comparability of reflectance data of different floating object samples.

[0077] See Figure 10 (a), for the first A multispectral target region 230 is defined, and all valid pixels within this region are used to form a sampling set. Assuming a total of 778 valid pixels, the sampling is performed according to a preset sampling ratio. The number of sampled pixels is determined to be 78. Randomly select samples from the sample set without repetition. Each pixel is used as a sampling pixel 240, and the multispectral target region 230 to which each sampling pixel 240 belongs is recorded; in this embodiment, a preset sampling ratio is used. It is 10%, the first The multispectral target region 230 contains a total of 778 effective pixels, and the number of selected sampling pixels 240 is 78. In this way, it is beneficial to obtain multi-band reflectivity from different positions within the outline range of the floating object instance, reducing the impact of local reflection at a single position, surface moisture, or pixel noise on material judgment.

[0078] Furthermore, as one implementation method, the material classification result corresponding to each sampled pixel 240 is obtained according to the reflectance of each sampled pixel 240 in multiple spectral bands, including: extracting the reflectance of each sampled pixel 240 in multiple spectral bands respectively; generating the spectral features corresponding to the same sampled pixel 240 based on the reflectance of the same sampled pixel 240 in multiple spectral bands; inputting the spectral features corresponding to each sampled pixel 240 into a pre-trained plastic material classification model to obtain the material classification result corresponding to each sampled pixel 240; and associating the material classification result corresponding to each sampled pixel 240 with the multispectral target region 230 to which the sampled pixel 240 belongs.

[0079] Specifically, in this embodiment, the reflectance of any sampling pixel 240 in six spectral bands—450nm, 555nm, 660nm, 720nm, 750nm, and 840nm—is extracted and recorded sequentially as follows: , , , , and ;Will to Each of the six spectral bands is used as a single-band reflectance characteristic, and the six spectral bands are combined in pairs. For the first... The spectral band and the first For each spectral band, the reflectance difference characteristics are determined according to the following formula. Reflectivity ratio characteristics Features of the difference between normalized reflectance : Six spectral bands form 15 band combinations, resulting in 15 reflectance difference features, 15 reflectance ratio features, and 15 normalized reflectance difference features. Combining the six single-band reflectance features with the 45 combined features yields 51-dimensional spectral features corresponding to 240 sampled pixels. When the absolute value of the denominator is less than a preset non-zero threshold, the corresponding denominator is replaced by the preset non-zero threshold, thus maintaining the numerical stability of the ratio features and the normalized reflectance difference features.

[0080] The plastic material classification model is pre-trained based on sampled pixel training samples with plastic material category labels. Each sampled pixel training sample includes the spectral features formed by the reflectance of the corresponding sampled pixel in six spectral bands, as well as the plastic material category label corresponding to the plastic sample from which the sampled pixel originates. In this embodiment, the plastic material categories include PET plastic, PS plastic, PP plastic and PE plastic. The number of training samples is 1000, and the training set and test set are divided in an 8:2 ratio.

[0081] In one embodiment, before training the plastic material classification model, a density-based clustering method is used to identify anomalous samples whose spectral feature distribution deviates from the main distribution of the corresponding material samples, and a recursive feature elimination method is used to filter the 51-dimensional spectral features. A support vector machine classification model is then trained based on the filtered spectral features, and the model's classification result is determined using a test set. In this embodiment, the final number of retained spectral features is 5, and the trained plastic material classification model achieves a classification accuracy of 94.6% on the test set.

[0082] The spectral features corresponding to each sampled pixel 240 are input into the trained plastic material classification model. The plastic material classification model outputs the material classification result corresponding to the sampled pixel 240. Based on the correlation between the sampled pixel 240 and the multispectral target region 230, the multispectral target region 230 to which the material classification result belongs is recorded. Specific results are as follows: Figure 10 As shown in (b), the six single-band reflectance features are used to characterize the direct reflection response of plastics in each spectral band, while the band combination features are used to characterize the relative changes between the reflection responses of different bands. In this way, it is beneficial to characterize the material differences of plastic floating objects in the river from two aspects: multiple spectral bands and the reflection relationship between bands, thereby improving the accuracy of the material classification results of the sampled pixels.

[0083] Furthermore, as one implementation method, based on the material classification results corresponding to each sampled pixel 240 belonging to the same multispectral target region 230, the plastic material category of the floating object instance 160 corresponding to the multispectral target region 230 is determined, including: grouping the material classification results corresponding to each sampled pixel 240 according to the multispectral target region 230 to which each sampled pixel 240 belongs; counting the occurrence frequency of each plastic material category in the material classification results belonging to the same group; and determining the plastic material category with the most occurrence frequency as the plastic material category of the floating object instance 160 corresponding to the multispectral target region 230 of that group.

[0084] Specifically, see Figure 10 (c), for the first 230 multispectral target regions, including... The sampling pixel is 240, the first... The material classification result corresponding to each sampled pixel 240 is: For any candidate plastic material category This plastic material category is in the first Number of occurrences in 230 multispectral target regions Determined according to the following formula: in, This indicates an indicator function that takes a value of 1 when the condition within the parentheses is true, and 0 otherwise; the most frequently occurring plastic material category is determined as the [number]. The plastic material category of 160 floating objects corresponding to 230 multispectral target regions. : In this embodiment, 78 sampling pixels 240 are selected within the 11th multispectral target region 230. Among them, the number of sampling pixels for PET plastic, PS plastic, PP plastic and PE plastic are 8, 8, 62 and 0 respectively. Among them, PP appears the most frequently, so PP is determined as the plastic material category of the floating object instance 160 corresponding to the multispectral target region 230.

[0085] When two or more plastic material categories appear the same number of times, the number of sampling pixels 240 selected from the corresponding multispectral target region 230 is increased. The newly added sampling pixels 240 are classified into different materials, and the occurrence frequency of each plastic material category is recounted. It can be understood that multiple sampling pixels 240 within the same multispectral target region 230 come from different locations of the same floating object instance 160. In this way, it is beneficial to use the material classification results at multiple pixel levels to jointly determine the plastic material category of the floating object instance 160, reduce the impact of local reflection anomalies or classification deviations of individual sampling pixels on the final material category, and thus improve the stability of the classification results of floating plastic objects in the river.

[0086] Through the above implementation, feature points 200 are first extracted from the single-channel image 180 and the single-band image 190d after pixel value distribution adjustment, candidate matching feature point pairs 210a are determined, and matching feature point pairs 210 are selected from the candidate matching feature point pairs 210a; the initial coordinate transformation relationship from the pixel coordinates in the single-channel image 180 to the pixel coordinates in the single-band image 190d after pixel value distribution adjustment is determined according to the matching feature point pairs 210; and according to the pixel position correspondence of the selected single-band image 190b before and after spatial resolution adjustment, the initial coordinate transformation relationship is converted into a coordinate mapping relationship from the pixel coordinates in the visible light image 140 to the pixel coordinates in the multispectral image 150; then, the instance mask 170 is converted into a multispectral target region 230 using the coordinate mapping relationship, and multiple sampling pixels 240 are selected from each multispectral target region 230 for reflectivity extraction and material classification; finally, the plastic material category of the corresponding floating object instance 160 is determined according to the material classification results of multiple sampling pixels 240 in the same multispectral target region 230. This approach helps maintain the correspondence between the outline range of floating object instances, the reflectivity extraction area, and the material classification results, reducing the interference of water surface background and reflection information of nearby floating objects on the classification results.

[0087] The embodiments described above are merely examples of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications without departing from the concept of the present invention, and these modifications all fall within the protection scope of the present invention.

Claims

1. A method for classifying floating plastic debris in river channels, characterized in that, include: Acquire visible light and multispectral images of the same river channel area; The visible light image is segmented to obtain multiple floating object instances and instance masks that characterize the outline range of each floating object instance. The visible light image is converted into a single-channel image, and a selected single-band image is determined from multiple single-band images of the multispectral image. Determining the selected single-band image from multiple single-band images of the multispectral image includes: extracting edges from the same scene range in both the single-channel image and each single-band image to obtain a single-channel edge image and single-band edge images corresponding to each single-band image; determining the overlap ratio between edge pixels in each single-band edge image and edge pixels in the single-channel edge image; determining the single-band image corresponding to the single-band edge image with the largest overlap ratio as the selected single-band image; adjusting the spatial resolution of the selected single-band image according to the spatial resolution of the single-channel image, and adjusting the pixel value distribution of the selected single-band image according to the pixel value distribution of the single-channel image to obtain a pixel value distribution. The process involves adjusting the spatial resolution of a single-band image, specifically by re-determining the position and pixel value of each pixel in the selected single-band image based on the spatial resolution of the single-channel image, resulting in a spatially resolution-adjusted single-band image. The cumulative pixel ratios corresponding to different pixel values ​​in both the single-channel image and the spatially resolution-adjusted single-band image are then determined. Based on the pixel values ​​corresponding to the same cumulative pixel ratio in both the single-channel image and the spatially resolution-adjusted single-band image, the pixel values ​​of each pixel in the spatially resolution-adjusted single-band image are adjusted to obtain a pixel value distribution-adjusted single-band image. Feature matching is then performed on both the single-channel image and the pixel value distribution-adjusted single-band image, and the feature matching results are transformed based on the pixel position correspondence before and after spatial resolution adjustment to obtain the coordinate mapping relationship between the visible light image and the multispectral image. Based on the coordinate mapping relationship, each instance mask is mapped onto the multispectral image to obtain the multispectral target region corresponding to each floating object instance. Multiple sampling pixels are selected from each multispectral target region, and the material classification results corresponding to each sampling pixel are obtained according to the reflectance of each sampling pixel in multiple spectral bands. Based on the material classification results of each sampled pixel belonging to the same multispectral target region, the plastic material category of the corresponding floating object instance is determined.

2. The method for classifying floating plastic debris in river channels according to claim 1, characterized in that, Feature matching is performed on the single-channel image and the single-band image with adjusted pixel value distribution, including: Candidate feature points in the single-channel image and the single-band image with adjusted pixel value distribution are detected at multiple image scales. Based on the gray-level changes in the neighborhood of each candidate feature point, candidate feature points with gray-level contrast below a preset value and candidate feature points located on a single edge are removed to obtain the feature points corresponding to the single-channel image and the single-band image with adjusted pixel value distribution. Based on the gray-level gradient direction and gray-level gradient magnitude of the pixels in the neighborhood of each feature point, feature description data corresponding to each feature point is generated. Based on the feature description data corresponding to the feature points in the single-channel image and the feature points in the single-band image with adjusted pixel value distribution, candidate matching feature point pairs are determined.

3. The method for classifying floating plastic debris in river channels according to claim 2, characterized in that, Determining matching feature point pairs from the candidate matching feature point pairs includes: For each feature point in the single-channel image, a first feature difference between the feature point and a first candidate feature point in the single-band image after pixel value distribution adjustment, and a second feature difference between the feature point and a second candidate feature point are determined, wherein the first feature difference is less than the second feature difference. Calculate the ratio between the first feature difference and the second feature difference, and remove candidate matching feature point pairs whose ratio does not meet a preset condition; From the remaining candidate matching feature point pairs, a preset number of candidate matching feature point pairs are selected multiple times, and the candidate coordinate transformation relationship is determined according to each selected candidate matching feature point pair. Determine the position error of the remaining candidate matching feature point pairs after transformation by each candidate coordinate transformation relationship, and count the number of candidate matching feature point pairs whose position error is less than a preset error; The candidate matching feature point pair with the largest number of corresponding candidate coordinate transformation relationships and a position error less than the preset error is determined as the matching feature point pair.

4. The method for classifying floating plastic debris in river channels according to claim 3, characterized in that, The coordinate mapping relationship is obtained based on the matched feature point pairs, including: Obtain the pixel coordinates of each matching feature point pair in the single-channel image and the single-band image after pixel value distribution adjustment. Based on the pixel coordinates of each matching feature point pair, establish a coordinate transformation model that transforms the pixel coordinates in the single-channel image to the pixel coordinates in the single-band image after pixel value distribution adjustment. Adjust the parameters of the coordinate transformation model so that the sum of the position errors of each matching feature point pair after coordinate transformation meets a preset condition. Then, determine the coordinate transformation model after parameter adjustment as the initial coordinate transformation relationship. Based on the positional correspondence of each pixel in the selected single-band image and the single-band image after spatial resolution adjustment, the initial coordinate transformation relationship is converted to the pixel coordinate system of the selected single-band image to obtain the coordinate mapping relationship from the pixel coordinates in the visible light image to the pixel coordinates in the multispectral image. The pixel coordinate system of the selected single-band image is the reference coordinate system of the multispectral image.

5. The method for classifying floating plastic debris in river channels according to claim 4, characterized in that, Mapping each instance mask to the multispectral image according to the coordinate mapping relationship includes: Different instance labels are set for instance masks corresponding to different floating object instances, and background labels are set for areas outside the instance masks. The output area of ​​each instance mask after being mapped to the multispectral image is determined. For each pixel in the output area, the corresponding position of the pixel in the instance mask is determined according to the coordinate mapping relationship. From the instance mask pixels surrounding the corresponding position, select the instance mask pixel that is closest to the corresponding position, and assign the instance label or background label corresponding to the selected instance mask pixel to the corresponding pixel in the output area; Based on the pixels assigned different instance labels, a multispectral target region corresponding to each floating object instance is formed.

6. The method for classifying floating plastic debris in river channels according to any one of claims 1-5, characterized in that, Before selecting multiple sampling pixels from each multispectral target region, the process also includes: The single-band images corresponding to different spectral bands in the multispectral image are registered to the same spatial coordinate system to form a multi-band image. The known reflectance of the reference plate and the pixel value of the reference plate in each spectral band are obtained. Based on the known reflectance of the reference plate, the pixel value of the reference plate in the corresponding spectral band, and the pixel value of the pixel to be tested in the corresponding spectral band, the reflectance of the pixel to be tested in the corresponding spectral band is determined. Multiple sampling pixels are randomly selected from the pixels contained in each multispectral target region, and each sampling pixel is associated with its respective multispectral target region.

7. The method for classifying floating plastic debris in river channels according to claim 6, characterized in that, Based on the reflectance of each sampled pixel in multiple spectral bands, the material classification results corresponding to each sampled pixel are obtained, including: The reflectance of each sampled pixel in multiple spectral bands is extracted. Based on the reflectance of the same sampled pixel in multiple spectral bands, the spectral features corresponding to the sampled pixel are generated. The spectral features corresponding to each sampled pixel are input into a pre-trained plastic material classification model to obtain the material classification results corresponding to each sampled pixel. The material classification results corresponding to each sampled pixel are associated with the multispectral target region to which the sampled pixel belongs.

8. The method for classifying floating plastic debris in river channels according to claim 7, characterized in that, Based on the material classification results corresponding to each sampled pixel belonging to the same multispectral target region, the plastic material category of the floating object instance corresponding to that multispectral target region is determined, including: Based on the multispectral target region to which each sampled pixel belongs, the material classification results corresponding to each sampled pixel are grouped. The frequency of occurrence of each plastic material category in the material classification results of the same group is counted. The plastic material category with the most occurrences is determined as the plastic material category of the floating object instance corresponding to the multispectral target region of that group.

Citation Information

Patent Citations

  • Beach plastic garbage unmanned aerial vehicle monitoring method based on map characteristics

    CN120047744A

  • KR20200078723A