Method and device for separating human body and water body based on spectral reflectance characteristics
By acquiring multispectral image data based on spectral reflectance characteristics and using lightweight model segmentation, the problems of detection accuracy and computational complexity in the water conservancy and water affairs industry have been solved, achieving efficient and real-time segmentation of pedestrian skin, clothing, and water bodies.
Patent Information
- Application Number
- CN202211617455.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-13
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2042-12-13
AI Technical Summary
Existing intelligent video surveillance systems, when used in river and lake management scenarios within the water conservancy and water affairs industry, are easily affected by changes in the apparent information of the target scene, resulting in poor detection accuracy, high computational complexity, and low operating efficiency.
By employing a method based on spectral reflectance characteristics, and through multispectral image data acquisition and lightweight model segmentation, combined with a dual-stream lightweight model that integrates qualitative and quantitative feature spectral band selection and spectral and spatial feature fusion, efficient segmentation of pedestrian skin, clothing, and water bodies is achieved.
It improves the accuracy and efficiency of detection, reduces false detections and missed detections, and achieves real-time processing.
Smart Images

Figure CN115830326B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method and apparatus for separating human bodies and water bodies based on spectral reflectance characteristics, belonging to the field of video surveillance. Background Technology
[0002] With the continuous upgrading of social security needs, intelligent video surveillance systems have been widely used in many fields, such as intelligent transportation systems, safe campus construction, security and prevention, and the water conservancy and water affairs industry. Research on intelligent video surveillance systems is of great significance and has promising application prospects. Water conservancy and water affairs supervision and management is a crucial part concerning people's lives and property safety, enterprise production, and environmental protection. In recent years, relevant departments have attached great importance to water safety and river and lake management and protection, and have successively issued a number of policies to guide and promote river and lake supervision and management. Pedestrians are a key target in river and lake management scenarios, making the detection of pedestrian skin and clothing essential.
[0003] Currently, most intelligent video surveillance systems use traditional monitoring equipment to capture visible light video data of the target scene and combine it with visible light visual processing algorithms to extract appearance information of the target scene to achieve specific visual tasks, such as pedestrian skin and clothing detection and water segmentation. While visible light data of the target scene can provide important appearance information for visual tasks such as pedestrian skin and clothing detection, in the context of river and lake management in the water conservancy industry, it is necessary to detect some common violations, such as illegal fishing and swimming. This scenario involves various complex situations, and existing video surveillance systems suffer from the following problems:
[0004] 1) The detection accuracy is poor due to changes in the appearance information of the target scene. In the context of river and lake management in the water conservancy and water affairs industry, the following complex situations exist: pedestrians are easily obscured by trees; pedestrians' skin or clothing is similar in appearance to the surrounding environment, such as camouflage clothing against trees, or pedestrians on the shore against strong reflections on the water; the shape of the part of a person swimming above the water is irregular. In the above complex scenarios, the appearance information of objects is prone to significant changes, which can easily lead to false detections and missed detections in related visual detection tasks, resulting in false alarms.
[0005] 2) The computational task is complex and the operating efficiency is low. In existing video surveillance systems, skin and clothing detection and water segmentation based on object appearance information are actually a process of appearance feature similarity matching. This requires calculating a large range of appearance features such as shape, texture, and color before matching and classifying. The computational load is large, and it is difficult to achieve real-time processing in embedded devices. Summary of the Invention
[0006] This invention proposes a method and apparatus for human body and water body segmentation based on spectral reflectance characteristics. According to a first aspect of the invention, a method for human body and water body segmentation based on spectral reflectance characteristics is provided, comprising: a qualitative and quantitative feature spectral band selection method and an optimization algorithm for skin and clothing detection and water body segmentation in complex environments of river and lake management. According to a second aspect of the invention, a device for human body and water body segmentation based on spectral reflectance characteristics is provided, comprising: a data acquisition module, a data transmission module, a module for pedestrian skin and clothing, a module for water body segmentation, and a data storage module. The device for human body and water body segmentation based on spectral reflectance characteristics is then designed and implemented using a multispectral camera.
[0007] To address the issue of poor detection accuracy due to susceptibility to changes in the appearance of the target scene, this invention introduces multispectral image data into the human body and water body segmentation task and proposes a feature spectral band selection method that combines qualitative and quantitative approaches. First, feature bands are selected based on the changing trends of the spectral reflectance curves of skin, clothing, and common water bodies, as well as their separability from the spectral curves of other materials. Finally, quantitative calculations are performed for verification.
[0008] To address the issues of high computational complexity and low operational efficiency, this invention designs a lightweight skin, clothing, and water segmentation model that fully leverages spectral and spatial image features to achieve efficient segmentation of skin, clothing, and water.
[0009] To address the various challenges faced by existing video surveillance systems in the water conservancy and water affairs industry, this invention proposes a human body and water body segmentation device based on spectral reflectance characteristics. Based on a development platform with multispectral image data acquisition capabilities, it captures multispectral image data of the scene in real time, performs model inference and related processing, segments pedestrian skin, clothing and water bodies in the scene, and displays the results in real time.
[0010] According to one aspect of the present invention, a method for separating the human body from water based on spectral reflectance characteristics is provided, characterized by comprising the following steps:
[0011] A) Select the characteristic bands of the acquired hyperspectral image data based on the changing trends of the spectral reflectance curves of skin, clothing, and common water bodies, as well as the separability of the spectral curves of other materials.
[0012] B) Verify the separability of the feature bands selected in step A by combining multiple distance measurement methods;
[0013] C) Use a multispectral camera to acquire multispectral image data of a scene including a body of water, the multispectral image data containing or containing the characteristic bands selected in step A;
[0014] D) The multispectral image data obtained in step C is preprocessed using a normalization method that preserves spectral and spatial image features in order to reduce the influence of ambient light.
[0015] E) Input the preprocessed multispectral image data from step D into a dual-stream lightweight model that fuses spectral features and spatial image features to segment pedestrian skin and clothing as well as water bodies in the scene.
[0016] F) Perform morphological post-processing (image erosion and dilation) on the segmentation results from step E to remove false detections caused by image noise and correct overexposed areas in the scene, thereby generating bounding boxes based on the segmentation results of pedestrian skin and clothing.
[0017] in:
[0018] Step D) includes:
[0019] D1) Perform spectral normalization to eliminate differences in the absolute values of the spectral curves of the same target object, while preserving the spectral reflectance characteristics of skin, clothing, and water, and highlighting the relative differences between the response values of each band of the target object's spectral curve. The spectral normalization formula is as follows:
[0020]
[0021] In the above formula, I(x, y, i) is the spectral response value of the original multispectral image data, x and y are the spatial location indices of the pixels, n is the number of selected bands, and i is the index of the spectral band. This indicates that the maximum value of the spectral response is obtained by traversing all spectral bands at each pixel position. The subscript spectrumNorm represents spectral direction normalization.
[0022] D2) To prevent spectral orientation normalization preprocessing from disrupting the spatial image features of multispectral image data, such as color and texture, spatial orientation normalization is performed simultaneously on the original multispectral image data. This preserves both the spectral and spatial image features of the multispectral image data, ensuring that these features can be fully utilized in subsequent algorithms. The spatial orientation normalization formula is as follows:
[0023] I spaceNorm (x,y,i)=(I(x,y,i)-μ i ) / σ i
[0024] In the above formula, I(x, y, i) is the spectral response value of the original multispectral image data, x and y are the spatial location indices of the pixels, i is the index of the spectral band, and μ i Let σ be the mean of the i-th band.i Let be the standard deviation of the i-th band, and let the subscript spaceNorm represent spatial orientation normalization.
[0025] Step E) includes:
[0026] Modeling of the dual-stream lightweight model described in E1)
[0027] A two-stream architecture is employed, involving the extraction of spectral features using a 1x1 convolutional kernel and the extraction of spatial image features using a kxk convolutional kernel. These two kernels are then fused, and a classifier is used for classification.
[0028] The training of the dual-stream lightweight model described in E2) includes:
[0029] First, the multispectral image data is preprocessed as described in step D to obtain two preprocessed multispectral image data sets.
[0030] The multispectral image is then segmented into multiple image patches, each with the same size and a category corresponding to the category of its central pixel. The image patches, after spectral orientation normalization, are input into the branch for extracting spectral features.
[0031] The image patch, after spatial orientation normalization, is input into the branch that extracts spatial image features. The final output of the model is the category to which the center pixel of the image patch belongs. Then, the model is learned through backpropagation.
[0032] E3) The testing of the dual-stream lightweight model includes:
[0033] First, the multispectral image data is preprocessed as described in step D to obtain two preprocessed multispectral image data sets.
[0034] Then, the multispectral image data normalized by spectral direction is input into the branch for extracting spectral features, and the multispectral image data normalized by spatial direction is input into the branch for extracting spatial image features.
[0035] Sliding convolution prediction is performed to obtain the segmentation result of the entire image.
[0036] According to another aspect of the present invention, a human body and water separation device based on spectral reflectance characteristics is provided, characterized in that it comprises:
[0037] The data acquisition module is deployed in the multispectral image data acquisition system. After the multispectral image data acquisition system is started, it automatically establishes a command connection with the host system. When it receives the start command transmitted from the host, the multispectral image data acquisition system begins to acquire multispectral image data.
[0038] The data transmission module is primarily used for transmitting multispectral image data and control commands between the multispectral image data acquisition system and the host system. All transmissions utilize TCP connections to ensure data accuracy. Control commands include: establishing and disconnecting multispectral image data transmission connections, and controlling camera exposure time.
[0039] The pedestrian skin and clothing segmentation module, as well as the water body segmentation module, specifically includes image normalization preprocessing, model inference, and segmentation result post-processing. Image normalization preprocessing includes spectral direction normalization and spatial direction normalization to simultaneously preserve the spectral and spatial image features of the data. The specific execution method is described in step D of the first aspect of this invention. The model inference part utilizes a GPU (Graphics Processing Unit) device and TensorRT model deployment technology to accelerate inference and prediction on the preprocessed multispectral image data. The model used is the trained dual-stream lightweight model described in step E of the first aspect of this invention. The execution flow is described in the dual-stream lightweight model testing described in step E of the first aspect of this invention. The segmentation result post-processing is consistent with step F of the first aspect of this invention, including the following parts: First, overexposed areas of the image are corrected; then, morphological operations such as erosion and dilation are performed on the segmentation results of the model to remove false detections caused by image noise; finally, an outer bounding box is generated based on the segmentation results of the pedestrian skin and clothing.
[0040] The data storage module allows the host system to save each complete frame of multispectral image data it receives locally for training or testing, in the original RAW and HDR file formats. Attached Figure Description
[0041] Figure 1 These are example graphs showing the spectral reflectance curves of pedestrian skin, dark cotton clothing, and green plants in different scenarios.
[0042] Figure 2 These are example graphs showing the spectral reflectance curves of water bodies in different water areas.
[0043] Figure 3 It is a spectral and spatially normalized visible light map.
[0044] Figure 4 This is a structural diagram of a dual-flow lightweight model design.
[0045] Figure 5 This is a comparison chart of the post-processing effects of the segmentation results.
[0046] Figure 6 This is a schematic diagram of the effect of a human body and water separation device based on spectral reflectance characteristics. Detailed Implementation
[0047] A method for separating the human body from water based on spectral reflectance characteristics according to an embodiment of the present invention includes the following steps:
[0048] A) Collect hyperspectral image data and select characteristic bands based on the changing trends of the spectral reflectance curves of skin, clothing, and common water bodies, as well as the separability of the spectral curves of other materials.
[0049] B) Verify the separability of the feature bands selected in step A by combining multiple distance measurement methods;
[0050] C) Acquire multispectral image data including water bodies using a multispectral camera, the multispectral image data containing or containing characteristic bands close to those selected in step A;
[0051] D) The multispectral image data obtained in step C is preprocessed using a normalization method that preserves spectral characteristics and spatial image features, in order to reduce the influence of ambient light.
[0052] E) Input the preprocessed multispectral image data from step D into the established dual-stream lightweight model that fuses spectral features and spatial image features to segment pedestrian skin, clothing, and water bodies in the scene.
[0053] F) Perform morphological post-processing of image erosion and dilation on the segmentation results in step E to remove false detections caused by image noise and correct overexposed areas in the scene, thereby generating bounding boxes based on the segmentation results of pedestrian skin and clothing.
[0054] A) Selection of characteristic bands
[0055] This invention utilizes hyperspectral imaging equipment to acquire hyperspectral image data in river and lake management scenarios. The acquisition process comprehensively considers factors such as shooting time, weather conditions, and shooting distance. Then, it selects clear hyperspectral images without motion blur, focus distortion, or overexposure. These hyperspectral images must uniformly include pedestrian skin and common clothing (such as black cotton or polyester clothing), common backgrounds (such as roads, greenery, and building facades), different water bodies (urban rivers, lakes, and streams), different weather conditions (sunny, cloudy, and light rain), and different time periods (morning, noon, and afternoon). The invention then observes the trends in the spectral reflectance curves of pedestrian skin, common clothing, and common water bodies, noting the presence of peaks or troughs in certain bands, or significant differences in absolute values between two bands. Simultaneously, the spectral reflectance curves of common background objects, such as greenery and concrete roads, are compared to ensure that the selected characteristic bands have good separability for skin, common clothing, water bodies, and common background objects. The spectral reflectance curves of pedestrian skin, dark cotton clothing, and green plants in different scenarios are as follows: Figure 1As shown, the spectral reflectance curves of different water bodies are as follows: Figure 2 As shown. In Figure 1 In (a), (c), and (e), the horizontal axis represents the hyperspectral band number, ranging from 0 to 127, and the vertical axis represents the raw response intensity value of the pixel acquired by the hyperspectral imaging device, ranging from 0 to 65536; Figure 1 In (b), (d), and (f), the meaning of the horizontal axis is the same as... Figure 1 The same as (a), the vertical axis represents the normalized response intensity value of the pixel acquired by the hyperspectral imaging device, with a value range of 0-1; in Figure 2 In each coordinate, the horizontal axis represents the wavelength of the spectral band, ranging from 449nm to 956nm, and the vertical axis has the same meaning as... Figure 1 (a) is the same. Finally, 12 feature bands were selected to characterize the spectral features of the spectral image data for subsequent processing.
[0056] B) Separability verification
[0057] To verify the separability of the selected bands in step A for pedestrian skin, common clothing, common water bodies, and common backgrounds, this invention selects a variety of distance metrics to further calculate the separability between classes. Specifically, the following methods can be selected, but are not limited to: Jeffries-Matusita distance, mean spectral divergence, and optimal index factor.
[0058] The formula for the JM (Jeffries-Matusita) distance is as follows:
[0059]
[0060]
[0061] In the above formula, m i m j Let represent the mean of each channel for categories i and j, respectively, as a vector, ∑ i ,∑ j Let |∑_{i=1}^{j} represent the covariance matrices between each channel of category i and j, respectively. The superscript T denotes the transpose of the matrix, and the superscript -1 denotes the inverse of the matrix. i |、|∑ j | respectively represent finding the matrix ∑ i ,∑ j The determinant value of , where e is the base of the natural logarithm, ln is the logarithm with the constant e as the base, and d ij JM represents the Bach distance between category i and category j. i,j This represents the JM distance between category i and category j.
[0062] The formula for calculating Mean Spectral Divergence (MSD) is as follows:
[0063]
[0064] D SKL (B i ||B j ) = D KL (B i ||B j )+D KL (B j ||B i )
[0065]
[0066] In the above formula, p(x) i ) and q(x i Let ) represent the statistical probability distributions of the responses of bands p and q, respectively, and let N represent the maximum quantization level of the response value of band p. i D represents the i-th quantization level. KL (p||q) represents the Kullback-Leibler divergence between bands p and q, D SKL (B i ||B j ) indicates band B i Band B j The symmetric Kullback-Leibler divergence, where k represents the number of selected bands, and MSD(B) represents the final average spectral divergence magnitude.
[0067] The formula for calculating the Optimum Index Factor (OIF) is as follows:
[0068]
[0069] In the above formula, S i R is the standard deviation of the i-th band. ij is the correlation coefficient between band i and band j, and n represents the number of selected bands.
[0070] Among them, the JM distance is a similarity-based measure that characterizes the separability of the selected bands between classes, with a value ranging from 0 to 2. A larger value indicates better separability between classes; generally, a JM distance greater than 1.8 is considered to indicate good separability. The average spectral divergence and the optimal index are information-based measures that characterize the amount of information contained in the selected bands. A smaller average spectral divergence indicates more redundant selected bands; a larger optimal index indicates less correlation between selected bands and richer information content. This invention calculated an average JM distance of 1.86396, an average spectral divergence of 0.77090, and an optimal index of 1092.02864, indicating that the selected bands in step A have good separability for pedestrian skin, common clothing, common water bodies, and common backgrounds.
[0071] D) Spectral data preprocessing
[0072] During data acquisition, it was found that the spectral curves of the same target were generally similar under different lighting intensities, but there were significant differences in the absolute values of the spectral responses. Therefore, spectral normalization was performed to eliminate the differences in the absolute values of the spectral curves of the same target object, while preserving the spectral reflectance characteristics of skin, clothing, and water. The focus was on the relative differences between the response values of different bands of the target object's spectral curve. The spectral normalization formula is as follows:
[0073]
[0074] In the above formula, I(x, y, i) is the spectral response value of the original multispectral image data, x and y are the spatial location indices of the pixels, n is the number of selected bands, and i is the index of the spectral band. This indicates that the maximum value of the spectral response is obtained by traversing all spectral bands at each pixel location. The subscript `spectrumNorm` represents spectral direction normalization. A comparison of the spectral curves before and after normalization is shown below. Figure 1 As shown.
[0075] However, spectral normalization can disrupt the spatial image features of multispectral image data, such as color and texture. Figure 3 As shown in (a) and (b), spatial orientation normalization is performed simultaneously on the original multispectral image data. This preserves both the spectral and spatial image features of the multispectral image data, ensuring that these features can be fully utilized in subsequent algorithms. This allows for an attempt to address the problem of different objects sharing the same spectrum. The spatial orientation normalization formula is as follows.
[0076] I spaceNorm (x,y,i)=(I(x,y,i)-μ i ) / σ i
[0077] In the above formula, I(x, y, i) is the spectral response value of the original multispectral image data, x and y are the spatial location indices of the pixels, i is the index of the spectral band, and μ i Let σ be the mean of the i-th band. i Let be the standard deviation of the i-th band, and the subscript spaceNorm represents spatial orientation normalization.
[0078] E) Modeling of a two-stream lightweight model
[0079] Multispectral image data not only characterizes the spectral reflectance properties of scene targets but also contains spatial image feature information of the scene. To maximize the utilization of the spectral and spatial image feature information contained in multispectral image data, a lightweight two-stream model that fuses spectral and spatial image features is designed to efficiently segment pedestrian skin, clothing, and water bodies. The overall model design structure is as follows: Figure 4 As shown.
[0080] The spectral feature extraction branch takes multispectral image data normalized by spectral orientation as input. It primarily uses a 1x1 convolutional kernel to extract spectral features and includes a spectral band attention extraction module. This module provides multiplicative weighting of spatial image features based on spectral band attention, representing the importance of texture features in different spectral bands in the prediction task. Similarly, the spatial image feature extraction branch takes spatially normalized multispectral image data as input. It primarily uses a kxk convolutional kernel (k can be 3, 5, etc.) to extract spatial image features and includes a spatial attention extraction module. This module provides multiplicative weighting of spectral features based on spatial attention, representing the importance of spectral features at adjacent spatial locations in the prediction task. Finally, the spectral and spatial image features are fused and fed into the final classifier to achieve the final category prediction.
[0081] For the aforementioned two-stream model, during the training phase, the image is first preprocessed as described in step D to obtain two preprocessed multispectral image data sets. Then, the image is segmented into multiple image patches, each with the same size and its category determined by its central pixel. The image patch with spectral orientation normalization is input to the branch extracting spectral features, while the image patch with spatial orientation normalization is input to the branch extracting spatial image features. The model's final output is the category of the central pixel of each image patch, which is then learned through backpropagation. During the testing phase, the image is again preprocessed as described in step D to obtain two preprocessed multispectral image data sets. The image with spectral orientation normalization is input to the branch extracting spectral features, and the image with spatial orientation normalization is input to the branch extracting spatial image features. Sliding convolution prediction is then performed to obtain the final segmentation result for the entire image.
[0082] F) Post-processing of segmentation results
[0083] Due to factors such as camera imaging noise and environmental noise, the segmentation results of algorithms are prone to false detections in small, noisy regions. This invention uses morphological operations involving opening to post-process the segmentation results and remove some of these false detections. Simultaneously, due to excessively strong sunlight or excessively long camera exposure times, the spectral response values of some regions in the multispectral image data approach the upper limit of the sensor's response, resulting in severe distortion of the target object's spectral curve. The final segmentation result for these regions becomes meaningless. Therefore, this invention calculates the maximum value I in the spectral direction of the multispectral image data. max (x, y) is used to extract overexposed or highly reflective areas, and then the detection results are corrected. The calculation formula is as follows.
[0084]
[0085]
[0086] In the above formula, I(x, y, i) is the spectral response value of the original multispectral image data, x and y are the spatial location indices of the pixels, n is the number of selected bands, and i is the index of the spectral band. This indicates that the maximum value of the spectral response is obtained by traversing all spectral bands at each pixel location. threshold I is a threshold slightly smaller than the maximum value of the multispectral camera response. overexposed (x, y) represents the extraction result of overexposed or highly reflective areas, with a value of 1 indicating overexposed or highly reflective areas.
[0087] Finally, a bounding rectangle can be generated based on the segmentation results of the pedestrian's skin and clothing. The overall post-processing effect is as follows: Figure 5 As shown in (a) and (b).
[0088] A schematic diagram of a human body and water separation device based on spectral reflectance characteristics according to an embodiment of the present invention is shown below. Figure 6 As shown, it includes:
[0089] (1) The host system is connected to the multispectral image data acquisition system via a network cable. It first logs into the multispectral image data acquisition system via SSH (Secure Shell), starts the program that listens for the command transmission connection, and then starts the main program on the host side. It will automatically establish a command transmission connection with the multispectral image data acquisition system, send control commands to the multispectral image data acquisition system through the data transmission module, start the acquisition of multispectral image data, and set the camera exposure time.
[0090] (2) A multispectral image data acquisition system, which captures multispectral image data of the scene through a data acquisition module, and then transmits the data to the host system via TCP communication through a data transmission module;
[0091] (3) Pedestrian skin and clothing and water body segmentation module, used to process the multispectral image data received by the host, including first performing normalization preprocessing of spectral direction and spatial direction to obtain two multispectral image data, then using the GPU (Graphics Processing Unit) device to use TensorRT model deployment technology to accelerate inference prediction of the preprocessed multispectral image data, then performing morphological opening operation, overexposure area correction and outer rectangle generation postprocessing, and finally rendering and displaying the postprocessing results on the front-end interface;
[0092] (4) Data storage module, used to save the original RAW and HDR image files locally for data analysis, training or testing;
[0093] The data acquisition module, implemented using the software development kit (SDK) of the multispectral camera, includes a camera initialization section and an image data acquisition section. The camera initialization section calls the SDK's Calibration() function to load the local configuration file for camera calibration, then calls the ConnectCamera() interface function to connect to the multispectral camera, and finally calls StartAcquisition() to start the multispectral camera's data acquisition process, completing the multispectral camera's initialization. Afterwards, the image data acquisition section calls the GetNextRawImage() function to obtain the multispectral image data acquired by the multispectral camera at the current moment, and then uses the data transmission module to transmit the multispectral image data to the host system for subsequent processing by the host system's segmentation or data storage module.
[0094] The data transmission module is used to transmit multispectral image data and control commands using a TCP connection to ensure the correctness of the transmitted data. The transmitted control commands include: establishing a multispectral image data transmission connection, disconnecting a multispectral image data transmission connection, and controlling the camera exposure time.
Claims
1. A method for human and water body segmentation based on spectral reflectance characteristics, characterized in that Comprising the following steps: A) Selecting the characteristic waveband of the collected hyperspectral image data according to the variation trend of the spectral reflectance curve of the skin, clothing and common water body, and the separability from other material spectral curves; B) Verifying the separability of the characteristic waveband selected in step A by comprehensively using multiple distance measurement methods; C) Collecting multispectral image data of a scene including a water body using a multispectral camera, the multispectral image data containing or including the characteristic waveband selected in step A; D) Preprocessing the multispectral image data obtained in step C using a normalization processing method that preserves spectral features and spatial image features to weaken the influence of ambient light; E) Inputting the preprocessed multispectral image data in step D into a dual-flow lightweight model that fuses spectral features and spatial image features to segment the skin, clothing and water body in the scene; F) Performing morphological post-processing such as image erosion and dilation on the segmentation result in step E to remove false detections caused by image noise and correct overexposed areas in the scene, thereby generating a bounding rectangle according to the segmentation result of the skin and clothing of the pedestrian, Wherein: The step D) comprises: D1) Performing spectral direction normalization to eliminate differences in absolute values of spectral curves of the same target object while preserving the spectral reflection characteristics of the skin, clothing and water body and highlighting the relative differences between response values of each waveband of the spectral curve of the target object, wherein the spectral direction normalization formula is: In the above equation, I(x, y, i) is the spectral response value of the original multispectral image data, x, y are the spatial position indices of the pixel, n is the number of selected wavebands, and i is the index of the spectral band, represents the maximum value of the spectral response obtained by traversing all spectral bands at each pixel position, and the subscript spectrumNorm represents the spectral direction normalization, D2) To prevent spectral direction normalization preprocessing from disturbing the spatial image features of the multispectral image data such as color, texture and other image features, simultaneously perform spatial direction normalization on the original multispectral image data to preserve both the spectral features and spatial image features of the multispectral image data to ensure that the spectral features and spatial image features of the multispectral image data can be fully utilized in subsequent algorithms, wherein the spatial direction normalization formula is: I spaceNorm (x, y, i) = (I(x, y, i) - μ i ) / σ i In the above formula, I(x, y, i) is the spectral response value of the original multispectral image data, x, y are the spatial position indexes of the pixels, i is the index of the spectral band, μ i is the mean value of the i-th band, σ i is the standard deviation of the i-th band, and the subscript spaceNorm represents spatial direction normalization, The step E) comprises: E1) Modeling of the dual-flow lightweight model Using a dual-flow structure, including using a 1*1 convolution kernel to extract spectral features, using a k*k size convolution kernel to extract spatial image features, then fusing the two, and then using a classifier for classification, E2) Training of the dual-flow lightweight model comprises: Firstly, preprocessing the multispectral image data in step D to obtain two preprocessed multispectral image data, Then divide the multispectral image into multiple image blocks, each image block has the same size and the class is the class of the center pixel, wherein the image block normalized in the spectral direction is input to the branch that extracts spectral features, The image block normalized in the spatial direction is input to the branch that extracts spatial image features, and the final output of the model is the class of the center pixel of the image block, and then learning through back propagation of the model.
2. The method of claim 1, wherein the method is based on spectral reflectance properties. The step E further comprises: E3) Testing the dual-flow lightweight model, comprising: Firstly, preprocessing the multispectral image data in step D to obtain two preprocessed multispectral image data, Then the multi-spectral image data subjected to the spectral direction normalization is input to a branch for extracting spectral features, and the multi-spectral image data subjected to the spatial direction normalization is input to a branch for extracting spatial image features, Sliding convolution prediction is performed, and finally the segmentation result of the whole image is obtained. 3.The method of claim 1, wherein Further comprising: Collecting the hyperspectral image data.
4. The method for human and water body segmentation based on spectral reflectance characteristics according to claim 1, characterized in that: k takes a value selected from 3 and 5.
5. A computer readable storage medium storing a computer program, which can enable a processor to execute the method according to any one of claims 1-4.
6. A human body and water body segmentation device based on spectral reflectance characteristics, characterized by Comprising: A) a feature band selection part for selecting the feature bands of the collected hyperspectral image data according to the variation trend of the spectral reflectance curves of the skin, clothes and common water bodies, and the separability from the spectral curves of other materials; B) a verification part for verifying the separability of the selected feature bands by comprehensively using multiple distance measurement methods; C) a multi-spectral camera for collecting multi-spectral image data of a scene including water bodies, the multi-spectral image data containing or including the selected feature bands; D) a preprocessing part for pre-processing the multi-spectral image data by using normalization processing methods for maintaining spectral characteristics and spatial image features, for weakening the influence of environmental light, wherein the pre-processed multi-spectral image data is input to a dual-flow lightweight model for fusing spectral features and spatial image features, to segment the skin, clothes and water bodies in the scene; F) a part for performing morphological post-processing of image erosion and expansion on the segmentation result, for removing false detections caused by image noise, and correcting overexposed areas in the scene, to generate a circumscribed rectangle according to the segmentation result of the skin and clothes of the pedestrian, Wherein: The preprocessing part comprises: D1) a part for spectral direction normalization processing, for eliminating the differences in the absolute values of the spectral curves of the same target objects, while retaining the spectral reflectance characteristics of the skin, clothes and water bodies, and highlighting the relative differences between the response values of the spectral curves of the target objects, wherein the normalization formula for the spectral direction is: In the above equation, I(x, y, i) is the spectral response value of the original multispectral image data, x, y are the spatial position indices of the pixel, n is the number of selected wavebands, and i is the index of the spectral band, represents the maximum value of the spectral response obtained by traversing all spectral bands at each pixel position, and the subscript spectrumNorm represents the spectral direction normalization. D2) a part for spatial direction normalization processing, for preventing the normalization preprocessing in the spectral direction from disturbing the spatial image features of the multi-spectral image data, such as color, texture and other image features, and for retaining the spectral features and spatial image features of the multi-spectral image data at the same time, to ensure that the spectral features and spatial image features of the multi-spectral image data can be fully utilized in the subsequent algorithms, wherein the normalization formula for the spatial direction is: I spaceNorm (x, y, i) = (I(x, y, i) - μ i ) / σ i In the above formula, I(x, y, i) is the spectral response value of the original multispectral image data, x, y are the spatial position indexes of the pixel, i is the index of the spectral band, μ i is the mean value of the i-th band, σ i is the standard deviation of the i-th band, and the subscript spaceNorm represents spatial direction normalization, The modeling of the dual-flow lightweight model comprises: using a dual-flow structure, including using a 1*1 convolution kernel to extract spectral features, using a k*k size convolution kernel to extract spatial image features, then fusing the two, and using a classifier for classification, The training of the dual-flow lightweight model comprises: First, the multi-spectral image data is pre-processed according to the preprocessing in step D, to obtain two pre-processed multi-spectral image data, Then the image is divided into multiple image blocks, each image block has the same size, and the category is the category to which the center pixel belongs, wherein the image block subjected to the spectral direction normalization is input into the branch for extracting spectral features, The image block subjected to the spatial direction normalization is input into the branch for extracting spatial image features, and the final output of the model is the category to which the center pixel of the image block belongs, and then learning is performed through the back propagation of the model.
7. The human body and water body segmentation apparatus based on spectral reflectance characteristics according to claim 6, characterized in that The double-flow lightweight model is tested, including: First, the multispectral image data is subjected to the preprocessing in step D to obtain two preprocessed multispectral image data, Then the multispectral image data subjected to the spectral direction normalization is input into the branch for extracting spectral features, and the multispectral image data subjected to the spatial direction normalization is input into the branch for extracting spatial image features, Sliding convolution prediction is performed to finally obtain the segmentation result of the whole image.
8. The human body and water body segmentation apparatus based on spectral reflectance characteristics according to claim 6, characterized in that Further comprising: Part of the hyperspectral image data is collected.
9. The human body and water body segmentation device based on spectral reflectance characteristics according to claim 6, characterized in that: k takes a value selected from 3 and 5.
Citation Information
Patent Citations
Self-adaptive sample selection-based hyperspectral urban water body detection method
CN108734122A
Hydrological remote sensing image target identification method based on deep semantic model
CN112084842A