An intelligent urine component detection method based on spectral recognition and deep learning

CN121190854BActive Publication Date: 2026-08-07SHIJIAZHUANG DEVOKANG BIOTECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHIJIAZHUANG DEVOKANG BIOTECHNOLOGY CO LTD
Filing Date
2025-09-22
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

在实际高光谱检测任务中,现有方法普遍缺乏对空间-光谱联合建模的结构,且尚未针对尿液成分的分类与浓度预测双重任务构建完整的端到端学习架构

Benefits of technology

[0065]首先,本发明设计了具备空间-光谱联合建模能力的三维卷积特征提取模块,,采用多尺度卷积路径提取局部空间-光谱特征,有效提升了高光谱尿液图像在空间与光谱维度上的表征精度。其次,本发明引入多头谱空间注意力网络模块,采用光谱通道注意力分支与空间位置注意力分支的并行结构,分别建模光谱域的通道依赖关系与空间域的位置依赖关系,并融合频谱感知嵌套卷积结构与可学习的位置编码机制,增强了对复杂空间-光谱特征的感知能力与表示深度。此外,本发明构建的特征融合模块引入短连接残差路径与通道压缩机制,能够实现多尺度特征融合与特征维度统一,有效提升了全局判别特征的表达能力;多任务预测模块并行配置了分类与回归分支,实现尿液成分类别与浓度的联合输出,并结合结构化检测数据生成机制,提升了检测结果的规范性与实用性。综上所述,本发明通过构建具有深度融合结构的尿液成分识别网络模型,形成了从高光谱图像输入到检测结果输出的端到端智能检测流程,在特征建模能力、任务泛化能力和检测结构完整性方面具有显著优势,具备工程应用价值与推广潜力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121190854B_ABST
    Figure CN121190854B_ABST
Patent Text Reader

Abstract

The application discloses an intelligent urine component detection method based on spectrum recognition and deep learning, comprising the following steps: S1, acquiring hyperspectral urine image data through a collecting device; S2, generating standardized urine data by preprocessing the hyperspectral urine image data; S3, constructing a urine component detection network model and performing spatial-spectrum joint feature analysis; S4, performing spatial-spectrum joint feature extraction through a three-dimensional convolution feature extraction module; S5, modeling the spectral domain and the spatial domain through a multi-head spectral spatial attention network module; S6, performing residual connection and dimension reconstruction through a feature fusion module; S7, obtaining urine component classification and concentration values and generating structured detection data through a multi-task prediction module; and S8, optimizing the model through a joint optimization mechanism. The application realizes automatic recognition and quantitative detection of multiple components in urine samples and is suitable for popularization and application in a clinical scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of spectral imaging analysis and intelligent medical detection technology, and in particular to an intelligent urine component detection method based on spectral recognition and deep learning. Background Technology

[0002] With the development of intelligent medical testing technology, urine component analysis, as a non-invasive and convenient clinical auxiliary diagnostic method, has been widely used in scenarios such as chronic disease screening, kidney function assessment, and early warning of metabolic abnormalities. Traditional urine testing mainly relies on test strips, colorimetric methods, or chemical analysis instruments. Although these methods are relatively mature, they still generally suffer from problems such as limited dimensions of detection indicators, reliance on manual intervention in the operation process, and significant influence of subjective factors on the stability of results. In addition, these methods are difficult to accurately model and quantify multiple components in complex samples, and cannot meet the requirements of intelligent health monitoring for high-throughput processing, multi-indicator fusion, and a balance between accuracy and efficiency.

[0003] In recent years, hyperspectral imaging technology, as a sensing method that integrates spatial information of images with spectral reflectance characteristics, has been increasingly applied in the field of medical testing. This technology enables non-destructive analysis of the microscopic components within a sample by acquiring spectral reflectance data across multiple continuous bands. Existing research has begun to apply hyperspectral technology to urine sample analysis, primarily employing statistical modeling-based methods to identify and discriminate specific target components. However, hyperspectral data suffers from high dimensionality, significant information redundancy, and strong spatial-spectral coupling. Traditional feature engineering methods are ineffective in uncovering deep-level feature structures, leading to insufficient model generalization ability and poor stability of detection results.

[0004] Deep learning, particularly convolutional neural networks (CNNs), has made significant progress in image classification and spectral recognition. However, directly applying two-dimensional CNN structures to hyperspectral urine images makes it difficult to simultaneously capture the continuous distribution characteristics of the spectral dimension and the local correlation information of the spatial dimension. In contrast, while Transformer-based detection models have advantages in global modeling, their model parameters are large and they are insensitive to small sample data. In practical hyperspectral detection tasks, existing methods generally lack structures for joint spatial-spectral modeling and have not yet built a complete end-to-end learning architecture for the dual tasks of urine component classification and concentration prediction.

[0005] Therefore, how to provide an intelligent urine component detection method based on spectral recognition and deep learning is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0006] One objective of this invention is to propose an intelligent urine component detection method based on spectral recognition and deep learning. This invention fully utilizes hyperspectral imaging technology, three-dimensional convolutional modeling, spectral spatial attention mechanism and multi-task learning structure to construct an end-to-end intelligent urine detection method oriented towards spatial-spectral joint modeling, which has the advantages of high modeling accuracy, strong feature extraction capability and high degree of structure in detection results.

[0007] A method for intelligent urine component detection based on spectral recognition and deep learning according to an embodiment of the present invention includes the following steps:

[0008] S1. Scan urine samples using a hyperspectral image acquisition device to obtain hyperspectral urine image data;

[0009] S2. Preprocess the hyperspectral urine image data to obtain standardized urine data;

[0010] S3. Construct a urine component detection network model that includes a three-dimensional convolutional feature extraction module and a multi-head spectral spatial attention network module, and perform spatial-spectral joint feature analysis on standardized urine data;

[0011] S4. Perform spatial-spectral joint feature extraction on standardized urine data through the three-dimensional convolutional feature extraction module, and output local spatial-spectral features;

[0012] S5. Local features are modeled in the spectral and spatial domains using a multi-head spectral spatial attention network module to generate deep semantic features.

[0013] S6. The feature fusion module performs residual connections and dimension reconstruction on deep semantic features to generate global discriminative features.

[0014] S7. The global discrimination features are used to obtain the component categories and concentration values ​​in the urine sample through the multi-task prediction module, and structured detection data is generated.

[0015] S8. The urine component identification and detection model is updated and optimized through a joint optimization mechanism.

[0016] Optionally, the hyperspectral urine image data is a hyperspectral data cube containing spatial and spectral dimensions, having a three-dimensional structure consisting of image height, image width, and number of spectral bands, with each image pixel corresponding to a set of spectral reflectance intensity values ​​on a continuous spectral band.

[0017] Optionally, the preprocessing includes image calibration, background removal, spectral band filtering, and data normalization, and step S2 specifically includes:

[0018] S21. Perform geometric distortion correction and pixel registration on the hyperspectral urine image to obtain a calibrated urine image with consistent spatial dimensions.

[0019] S22. Based on the set light intensity threshold, mask image or segmentation algorithm, remove the background area in the calibrated urine image that does not contain a urine sample;

[0020] S23. By evaluating the signal-to-noise ratio of each band and calculating the information entropy, bands with low signal-to-noise ratio, high information redundancy, or weak discrimination ability are obtained as effective spectral band sets.

[0021] S24. Perform numerical standardization on the spectral vector of each pixel in the effective spectral channel set to unify the feature scale of each channel and generate standardized urine data. The standardized urine data has a size of H×W×B, where H represents the image height, W represents the image width, and B represents the number of retained effective spectral channels.

[0022] Optionally, the urine component detection network model in step S3 includes a three-dimensional convolutional feature extraction module, a multi-head spectral spatial attention network module, a feature fusion module, and a multi-task prediction module, and the urine component identification and detection model is dynamically updated through a joint optimization mechanism;

[0023] The three-dimensional convolutional feature extraction module adopts a joint modeling approach across spectral and spatial dimensions, introducing continuously stacked three-dimensional convolutional kernels to encode local features in standardized urine data and output local spatial-spectral features.

[0024] The multi-head spectral spatial attention network module adopts a parallel spectral channel attention and spatial position attention structure, and combines spectral sensing nested convolution and learnable position encoding to perform cross-dimensional modeling of local spatial-spectral features and output deep semantic features.

[0025] The feature fusion module introduces short connection residual paths and channel compression mechanisms to perform multi-scale feature combination and dimensional unification processing on deep semantic features, and outputs global discriminative features.

[0026] The multi-task prediction module is configured with independent classification and regression branches, and constructs a category discrimination function and a concentration mapping function based on global discriminant features to obtain the component categories and concentration values ​​in the urine sample.

[0027] Optionally, step S4 specifically includes:

[0028] S41. Input the standardized urine data X into the three-dimensional convolutional feature extraction module to perform spatial-spectral joint feature extraction;

[0029] S42. Set up two sets of parallel 3D convolutional branch structures. The first branch uses a convolutional kernel size of 3×3×5, and the second branch uses a convolutional kernel size of 5×5×7. Perform joint convolution operations of spatial dimension and spectral dimension on the input tensor respectively to extract local spatial-spectral features at different scales.

[0030] S43. After each convolutional branch, a spectral domain-aware gating mechanism is connected. The spectral domain-aware gating mechanism is as follows: first, a global average pooling operation is performed on each spectral band channel in the spatial dimension to obtain a spectral description vector. Then, the spectral description vector is used to generate spectral channel attention weight coefficients through a two-layer fully connected network. The output of the convolutional branch is then adjusted channel by channel to generate spectral response features.

[0031] S44. The spectral response characteristics are integrated through a weighted fusion method to obtain the fused feature X. ′ ;

[0032] S45, By fusing feature X ′ A gated residual structure is constructed using standardized urine data X. The response strength of the main branch and the residual branch is adjusted using the channel weight vector α to form the residual output X. res :

[0033] X res =α·X ′ +(1-α)·X;

[0034] Where α∈[0,1] is a learnable channel attention factor;

[0035] S46. Output the residual X res Perform a 2×2 max pooling operation on the image's height and width spatial dimensions, keeping the spectral band dimensions unchanged, and output the local spatial-spectral feature F.

[0036] Optionally, step S4 specifically includes:

[0037] S51. Input the local spatial-spectral feature F into the multi-head spectral spatial attention network module, which includes parallel spectral channel attention branches and spatial location attention branches, as well as a gated interactive fusion module.

[0038] S52. In the spectral channel attention branch, the input local spatial-spectral feature F is divided into h self-attention heads along the spectral dimension. Each attention head embeds a spectral sensing nested convolutional structure. The nested structure is composed of an outer 1×1×k1 convolutional kernel and an inner 1×1×k2 convolutional kernel, where k1>k2. After performing two-level spectral domain convolution operations on the input feature, the elements are combined in a weighted manner to generate the spectral domain response feature F. s :

[0039] F s =σ(W1*F)+σ(W2*F);

[0040] Where * represents a 3D convolution operation, W1 and W2 are the convolution kernel scales, and σ represents the Gaussian error linear unit activation function;

[0041] S53. In the spatial location attention branch, three sets of asymmetric convolutional paths are constructed with kernel sizes of 1×3, 3×1, and 3×3. These paths perform convolution operations with the local spatial-spectral features F along the horizontal and vertical directions and with the two-dimensional neighborhood, respectively. A learnable location encoding vector P is introduced at each spatial location and summed element-wise with the convolution output to generate the spatial response feature F. x :

[0042] F x =F conv +P;

[0043] Among them, F conv Indicates the convolution output;

[0044] S54, Spectral domain response characteristics F s Spatial response characteristics F x The input is fed into the gated interaction fusion module, which fuses the channel responses of the two branches using a learnable channel gating factor β∈[0,1], and outputs deep semantic features F. out :

[0045] F out =β·F s +(1-β)·F x ;

[0046] Where β is a weight vector that is consistent with the number of channels.

[0047] Optionally, step S6 specifically includes:

[0048] S61, The deep semantic features F out The input is fed into the feature fusion module, which includes a main branch and an auxiliary branch;

[0049] The main branch includes a channel compression mechanism and a channel weighting structure. The channel compression mechanism is as follows: three sets of parallel 1×1 convolutions are used to perform channel mapping operations with different compression ratios on the deep semantic features to generate intermediate features at three scales.

[0050] The channel weighting structure is configured with three sets of learnable channel response coefficients, which weight each intermediate feature according to the channel dimension and perform element-wise summation to output the main branch fusion feature F. m ;

[0051] S62, Incorporate deep semantic features F out The synchronous input is fed into an auxiliary branch with short-connection residual paths, and the deep semantic features F are processed by a set of 1×1 convolutions. out Align the channels according to their dimensions and output auxiliary features F. a ;

[0052] S63, Merge main branch feature F m With auxiliary feature F a Perform residual connections along the channel direction and global average pooling in the spatial dimension to generate global discriminative features G:

[0053] G = GAP(μ·F) m +(1-μ)·F a )

[0054] Where μ is the residual fusion factor and μ∈[0,1], and GAP(·) represents the global average pooling operation.

[0055] Optionally, step S7 specifically includes:

[0056] S71. Input the global discriminative feature G into the multi-task prediction module, and generate the shared prediction feature Z through a fully connected transformation structure;

[0057] S72. Input the shared prediction feature Z into two structurally separated component classification branches and concentration regression branches, respectively. The component classification branch consists of two layers of nonlinear transformation units and a softmax output node, outputting a component category probability vector. The concentration regression branch consists of two layers of nonlinear transformation units and a linear output node, outputting the concentration estimation vectors corresponding to each component.

[0058] S73, Convert the component category probability vector With concentration estimation vector The data is combined to generate structured detection data, which includes the name of the detected component, its category, estimated concentration value, concentration unit, reference range, deviation level, feature contribution score, and sample number, and is presented in a field-based format to form a visualized detection result.

[0059] Optionally, step S8 specifically includes:

[0060] S81, Based on component category probability vector With concentration estimation vector Construct the multi-task loss function L:

[0061]

[0062] Among them, L cL represents the cross-entropy loss function for component classification. r The mean squared error loss function for concentration regression, y c With y r To correspond to the true label, λ is the task loss weighting factor, which represents the target label in the component classification task and the concentration regression task, respectively;

[0063] S82. Based on the multi-task loss function L, perform end-to-end model parameter updates on the urine component detection network model. The model parameters include the convolution kernel weights and gating parameters in the three-dimensional convolution feature extraction module, the attention mapping parameters and position encoding weights in the multi-head spectral spatial attention network module, the channel compression and residual fusion parameters in the feature fusion module, and the weights and bias terms of the fully connected layers in the multi-task prediction module.

[0064] The beneficial effects of this invention are:

[0065] First, this invention designs a 3D convolutional feature extraction module with spatial-spectral joint modeling capabilities. It employs multi-scale convolutional paths to extract local spatial-spectral features, effectively improving the representation accuracy of hyperspectral urine images in both spatial and spectral dimensions. Second, this invention introduces a multi-head spectral spatial attention network module, using a parallel structure of spectral channel attention branches and spatial location attention branches to model channel dependencies in the spectral domain and location dependencies in the spatial domain, respectively. It also integrates a spectrum-aware nested convolutional structure and a learnable location encoding mechanism, enhancing the perception and representation depth of complex spatial-spectral features. Furthermore, the feature fusion module constructed in this invention introduces short-connection residual paths and channel compression mechanisms, enabling multi-scale feature fusion and feature dimension unification, effectively improving the expressive power of global discriminative features. The multi-task prediction module is configured with parallel classification and regression branches, achieving joint output of urine component categories and concentrations. Combined with a structured detection data generation mechanism, it improves the standardization and practicality of the detection results. In summary, this invention constructs a urine component recognition network model with a deep fusion structure, forming an end-to-end intelligent detection process from hyperspectral image input to detection result output. It has significant advantages in feature modeling capability, task generalization capability, and detection structure integrity, and has engineering application value and promotion potential. Attached Figure Description

[0066] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0067] Figure 1 This is a schematic diagram of an intelligent urine component detection method based on spectral recognition and deep learning proposed in this invention;

[0068] Figure 2 This is a schematic diagram of the integration of three-dimensional convolution and spectral spatial attention modules in this invention;

[0069] Figure 3 This is a schematic diagram of multi-task prediction and model optimization in this invention. Detailed Implementation

[0070] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0071] refer to Figure 1-3 A smart urine component detection method based on spectral recognition and deep learning includes the following steps:

[0072] S1. Scan urine samples using a hyperspectral image acquisition device to obtain hyperspectral urine image data;

[0073] S2. Preprocess the hyperspectral urine image data to obtain standardized urine data;

[0074] S3. Construct a urine component detection network model that includes a three-dimensional convolutional feature extraction module and a multi-head spectral spatial attention network module, and perform spatial-spectral joint feature analysis on standardized urine data;

[0075] S4. Perform spatial-spectral joint feature extraction on standardized urine data through the three-dimensional convolutional feature extraction module, and output local spatial-spectral features;

[0076] S5. Local features are modeled in the spectral and spatial domains using a multi-head spectral spatial attention network module to generate deep semantic features.

[0077] S6. The feature fusion module performs residual connections and dimension reconstruction on deep semantic features to generate global discriminative features.

[0078] S7. The global discrimination features are used to obtain the component categories and concentration values ​​in the urine sample through the multi-task prediction module, and structured detection data is generated.

[0079] S8. The urine component identification and detection model is updated and optimized through a joint optimization mechanism.

[0080] Optionally, the hyperspectral urine image data is a hyperspectral data cube containing spatial and spectral dimensions, having a three-dimensional structure consisting of image height, image width, and number of spectral bands, with each image pixel corresponding to a set of spectral reflectance intensity values ​​on a continuous spectral band.

[0081] In this embodiment, the preprocessing includes image calibration, background removal, spectral band filtering, and data normalization. Step S2 specifically includes:

[0082] S21. Perform geometric distortion correction and pixel registration on the hyperspectral urine image to obtain a calibrated urine image with consistent spatial dimensions.

[0083] S22. Based on the set light intensity threshold, mask image or segmentation algorithm, remove the background area in the calibrated urine image that does not contain a urine sample;

[0084] S23. By evaluating the signal-to-noise ratio of each band and calculating the information entropy, bands with low signal-to-noise ratio, high information redundancy, or weak discrimination ability are obtained as effective spectral band sets.

[0085] S24. Perform numerical standardization on the spectral vector of each pixel in the effective spectral channel set to unify the feature scale of each channel and generate standardized urine data. The standardized urine data has a size of H×W×B, where H represents the image height, W represents the image width, and B represents the number of retained effective spectral channels.

[0086] In this embodiment, the urine component detection network model in step S3 includes a three-dimensional convolutional feature extraction module, a multi-head spectral spatial attention network module, a feature fusion module, and a multi-task prediction module, and the urine component identification and detection model is dynamically updated through a joint optimization mechanism;

[0087] The three-dimensional convolutional feature extraction module adopts a joint modeling approach across spectral and spatial dimensions, introducing continuously stacked three-dimensional convolutional kernels to encode local features in standardized urine data and output local spatial-spectral features.

[0088] The multi-head spectral spatial attention network module adopts a parallel spectral channel attention and spatial position attention structure, and combines spectral sensing nested convolution and learnable position encoding to perform cross-dimensional modeling of local spatial-spectral features and output deep semantic features.

[0089] The feature fusion module introduces short connection residual paths and channel compression mechanisms to perform multi-scale feature combination and dimensional unification processing on deep semantic features, and outputs global discriminative features.

[0090] The multi-task prediction module is configured with independent classification and regression branches, and constructs a category discrimination function and a concentration mapping function based on global discriminant features to obtain the component categories and concentration values ​​in the urine sample.

[0091] In this embodiment, step S4 specifically includes:

[0092] S41. Input the standardized urine data X into the three-dimensional convolutional feature extraction module to perform spatial-spectral joint feature extraction;

[0093] S42. Set up two sets of parallel 3D convolutional branch structures. The first branch uses a convolutional kernel size of 3×3×5, and the second branch uses a convolutional kernel size of 5×5×7. Perform joint convolution operations of spatial dimension and spectral dimension on the input tensor respectively to extract local spatial-spectral features at different scales.

[0094] S43. After each convolutional branch, a spectral domain-aware gating mechanism is connected. The spectral domain-aware gating mechanism is as follows: first, a global average pooling operation is performed on each spectral band channel in the spatial dimension to obtain a spectral description vector. Then, the spectral description vector is used to generate spectral channel attention weight coefficients through a two-layer fully connected network. The output of the convolutional branch is then adjusted channel by channel to generate spectral response features.

[0095] S44. The spectral response characteristics are integrated through a weighted fusion method to obtain the fused feature X. ′ ;

[0096] S45, By fusing feature X ′ A gated residual structure is constructed using standardized urine data X. The response strength of the main branch and the residual branch is adjusted using the channel weight vector α to form the residual output X. res :

[0097] X res =α·X ′ +(1-α)·X;

[0098] Where α∈[0,1] is a learnable channel attention factor;

[0099] S46. Output the residual X res Perform a 2×2 max pooling operation on the image's height and width spatial dimensions, keeping the spectral band dimensions unchanged, and output the local spatial-spectral feature F.

[0100] In this embodiment, step S5 specifically includes:

[0101] S51. Input the local spatial-spectral feature F into the multi-head spectral spatial attention network module, which includes parallel spectral channel attention branches and spatial location attention branches, as well as a gated interactive fusion module.

[0102] S52. In the spectral channel attention branch, the input local spatial-spectral feature F is divided into h self-attention heads along the spectral dimension. Each attention head embeds a spectral sensing nested convolutional structure. The nested structure is composed of an outer 1×1×k1 convolutional kernel and an inner 1×1×k2 convolutional kernel, where k1>k2. After performing two-level spectral domain convolution operations on the input feature, the elements are combined in a weighted manner to generate the spectral domain response feature F. s :

[0103] F s =σ(W1*F)+σ(W2*F);

[0104] Where * represents a 3D convolution operation, W1 and W2 are the convolution kernel scales, and σ represents the Gaussian error linear unit activation function;

[0105] S53. In the spatial location attention branch, three sets of asymmetric convolutional paths are constructed with kernel sizes of 1×3, 3×1, and 3×3. These paths perform convolution operations with the local spatial-spectral features F along the horizontal and vertical directions and with the two-dimensional neighborhood, respectively. A learnable location encoding vector P is introduced at each spatial location and summed element-wise with the convolution output to generate the spatial response feature F. x :

[0106] F x =F conv +P;

[0107] Among them, F conv Indicates the convolution output;

[0108] S54, Spectral domain response characteristics F s Spatial response characteristics F x The input is fed into the gated interaction fusion module, which fuses the channel responses of the two branches using a learnable channel gating factor β∈[0,1], and outputs deep semantic features F. out :

[0109] F out =β·F s +(1-β)·F x ;

[0110] Where β is a weight vector that is consistent with the number of channels.

[0111] In this embodiment, step S6 specifically includes:

[0112] S61, The deep semantic features F out The input is fed into the feature fusion module, which includes a main branch and an auxiliary branch;

[0113] The main branch includes a channel compression mechanism and a channel weighting structure. The channel compression mechanism is as follows: three sets of parallel 1×1 convolutions are used to perform channel mapping operations with different compression ratios on the deep semantic features to generate intermediate features at three scales.

[0114] The channel weighting structure is configured with three sets of learnable channel response coefficients, which weight each intermediate feature according to the channel dimension and perform element-wise summation to output the main branch fusion feature F. m ;

[0115] S62, Incorporate deep semantic features F out The synchronous input is fed into an auxiliary branch with short-connection residual paths, and the deep semantic features F are processed by a set of 1×1 convolutions. out Align the channels according to their dimensions and output auxiliary features F. a ;

[0116] S63, Merge main branch feature F m With auxiliary feature F a Perform residual connections along the channel direction and global average pooling in the spatial dimension to generate global discriminative features G:

[0117] G = GAP(μ·F) m +(1-μ)·F a )

[0118] Where μ is the residual fusion factor and μ∈[0,1], and GAP(·) represents the global average pooling operation.

[0119] In this embodiment, step S7 specifically includes:

[0120] S71. Input the global discriminative feature G into the multi-task prediction module, and generate the shared prediction feature Z through a fully connected transformation structure;

[0121] S72. Input the shared prediction feature Z into two structurally separated component classification branches and concentration regression branches, respectively. The component classification branch consists of two layers of nonlinear transformation units and a softmax output node, outputting a component category probability vector. The concentration regression branch consists of two layers of nonlinear transformation units and a linear output node, outputting the concentration estimation vectors corresponding to each component.

[0122] S73, Convert the component category probability vector With concentration estimation vector The data is combined to generate structured detection data, which includes the name of the detected component, its category, estimated concentration value, concentration unit, reference range, deviation level, feature contribution score, and sample number, and forms a visual detection result set in a field-based format.

[0123] In this embodiment, step S8 specifically includes:

[0124] S81, Based on component category probability vector With concentration estimation vector Construct the multi-task loss function L:

[0125]

[0126] Among them, L c L represents the cross-entropy loss function for component classification. r The mean squared error loss function for concentration regression, y c With y r To correspond to the true label, λ is the task loss weighting factor, which represents the target label in the component classification task and the concentration regression task, respectively;

[0127] S82. Based on the multi-task loss function L, perform end-to-end model parameter updates on the urine component detection network model. The model parameters include the convolution kernel weights and gating parameters in the three-dimensional convolution feature extraction module, the attention mapping parameters and position encoding weights in the multi-head spectral spatial attention network module, the channel compression and residual fusion parameters in the feature fusion module, and the weights and bias terms of the fully connected layers in the multi-task prediction module.

[0128] Example 1:

[0129] To verify the practical application effect of this invention, it was applied to the peak physical examination scenario of the laboratory department of a tertiary hospital in a certain city from March to June 2025. This department receives over 1000 urine samples daily. Traditional semi-automatic test strip methods rely on manual solution preparation, colorimetric interpretation, and result entry, limiting the detection capabilities and requiring six laboratory technicians working in three shifts to complete the entire workload. After adopting the method of this invention, the testing process is simplified to three steps: hyperspectral image acquisition, model reasoning analysis, and structured result generation. The overall processing time for a single sample is reduced to approximately 4 seconds.

[0130] This invention acquires multi-scale local features through a 3D convolutional feature extraction module that uses spatial-spectral joint modeling. Subsequently, in a multi-head spectral spatial attention network module, spectral channel attention structures and spatial location attention structures are constructed respectively to achieve joint modeling of information from different dimensions. After processing by a feature fusion module, global discriminative features are generated and input into a multi-task prediction module. These features are processed in parallel by classification and regression branches, simultaneously outputting category labels and corresponding concentration values ​​for fourteen common urine components. During the testing process, only one technician is needed to complete sample numbering and result verification; the rest of the process is fully automated, significantly reducing the manual workload.

[0131] The method of this invention is deployed on a GPU server within a hospital's local area network environment. The accompanying hyperspectral image acquisition equipment operates in the 400–1000 nm band with a spectral resolution of 3 nm, and a total of 200 bands. The hospital randomly selected 500 urine samples from physical examinations as a blind test set, and simultaneously retested eight key components using liquid chromatography-tandem mass spectrometry (LC-MS / MS) as reference data. Furthermore, to verify the effectiveness of this invention in deep learning structure design, a two-dimensional convolutional neural network (2D-CNN) method and a Transformer method were selected as baseline models for comparison. Table 1 shows the comparison results between the method of this invention and the three aforementioned comparative methods in terms of component identification accuracy, concentration prediction error, and average processing time.

[0132] Table 1. Performance comparison of the method of the present invention with different comparative methods in urine component detection tasks.

[0133]

[0134]

[0135] As shown in Table 1, the method of this invention achieves a component category identification accuracy of 96.4%, significantly outperforming the 88.2% of the 2D-CNN method and the 90.5% of the Transformer method. While maintaining high accuracy, its performance is close to that of the standard laboratory LC-MS / MS method. Regarding concentration prediction accuracy, the mean absolute error of the method of this invention is 0.028 mmol / L, better than the 0.081 mmol / L of 2D-CNN and the 0.064 mmol / L of Transformer, and essentially close to the 0.017 mmol / L of LC-MS / MS, verifying the effectiveness of this method in quantitative estimation. In terms of processing efficiency, the single-sample processing time of this invention is approximately 3.9 seconds, significantly faster than the 30-minute timeframe of the LC-MS / MS method, while balancing speed and accuracy compared to the 2D-CNN and Transformer methods. The Kappa coefficient of this invention reaches 0.93, superior to other comparative methods, demonstrating higher consistency in cross-sample detection and stronger stability and generalization ability. This method supports the identification and concentration prediction of up to 14 components, significantly outperforming the 8 indicators of LC-MS / MS and the 12 indicators of 2D-CNN methods, demonstrating higher throughput and adaptability. Furthermore, this method maintains high sensitivity when identifying low-concentration samples, while the identification performance of comparative methods shows a significant decline. Regarding detection stability, the coefficient of variation of this method is 2.1%, close to the 1.5% of laboratory methods, demonstrating good data consistency and robustness.

[0136] In summary, the tabular data fully demonstrates the advantages of the method of the present invention in hyperspectral data processing, feature modeling capabilities, and detection efficiency. It is superior to existing methods in terms of accuracy and stability, and better meets the actual application needs of clinical practice for high-throughput and intelligent urine component detection.

[0137] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. An intelligent urine component detection method based on a spectrum recognition and deep fusion network, characterized in that, Includes the following steps: S1. Scan urine samples using a hyperspectral image acquisition device to obtain hyperspectral urine image data; S2. Preprocess the hyperspectral urine image data to obtain standardized urine data; S3. Construct a urine component detection network model and perform spatial-spectral joint feature analysis on standardized urine data; the urine component detection network model includes a three-dimensional convolutional feature extraction module, a multi-head spectral spatial attention network module, a feature fusion module, and a multi-task prediction module, and dynamically updates the urine component identification and detection model through a joint optimization mechanism; The three-dimensional convolutional feature extraction module adopts a joint modeling approach across spectral and spatial dimensions, introducing continuously stacked three-dimensional convolutional kernels to encode local features in standardized urine data and output local spatial-spectral features. The multi-head spectral spatial attention network module adopts a parallel spectral channel attention and spatial position attention structure, and combines spectral sensing nested convolution and learnable position encoding to perform cross-dimensional modeling of local spatial-spectral features and output deep semantic features. The feature fusion module introduces short connection residual paths and channel compression mechanisms to perform multi-scale feature combination and dimensional unification processing on deep semantic features, and outputs global discriminative features. The multi-task prediction module is configured with independent classification and regression branches, and constructs a category discrimination function and a concentration mapping function based on global discriminant features to obtain the component categories and concentration values ​​in the urine sample. S4. The standardized urine data undergoes spatial-spectral joint feature extraction using a 3D convolutional feature extraction module, outputting local spatial-spectral features, specifically including: S41. Standardize urine data The input is fed into the 3D convolutional feature extraction module for joint spatial-spectral feature extraction; S42. Set up two sets of parallel 3D convolutional branch structures. The first branch uses a convolutional kernel size of... The second branch uses a convolution kernel size of The input tensor is subjected to joint convolution operations of spatial and spectral dimensions to extract local spatial-spectral features at different scales. S43. After each convolutional branch, a spectral domain-aware gating mechanism is connected. The spectral domain-aware gating mechanism is as follows: first, a global average pooling operation is performed on each spectral band channel in the spatial dimension to obtain a spectral description vector. Then, the spectral description vector is used to generate spectral channel attention weight coefficients through a two-layer fully connected network. The output of the convolutional branch is then adjusted channel by channel to generate spectral response features. S44. The spectral response characteristics are integrated through a weighted fusion method to obtain the fused features; S45. By fusing features and standardized urine data, a gated residual structure is constructed. The response intensity of the main branch and the residual branch is adjusted by using the channel weight vector to form the residual output. S46. Output the residual in the image height and width spatial dimensions and perform a kernel size of... The max pooling operation maintains the spectral band dimension unchanged and outputs local spatial-spectral features; S5. Local features are modeled in both the spectral and spatial domains using a multi-head spectral spatial attention network module to generate deep semantic features, specifically including: S51. Input the local spatial-spectral features into the multi-head spectral spatial attention network module, which includes parallel spectral channel attention branches and spatial location attention branches, as well as a gated interactive fusion module. S52. In the spectral channel attention branch, the input local spatial-spectral features are divided along the spectral dimension into... Each attention head contains a self-attention head, and each attention head embeds a spectral sensing nested convolutional structure. The nested convolutional structure consists of an outer layer... Convolution kernel and inner layer Convolutional kernels are stacked together, where After performing two-level spectral domain convolution operations on the input features, the elements are combined in a weighted manner to generate spectral domain response features; S53. In the spatial location attention branch, construct three sets of asymmetric convolutional paths with kernel size of [size missing]. , and Local spatial-spectral features along the horizontal and vertical directions and the two-dimensional neighborhood, respectively. Perform convolution operations and introduce learnable location encoding vectors at each spatial location. Then, sum them element-wise with the convolution output to generate spatial response features. S54. Input the spectral response features and spatial response features into the gated interactive fusion module, and fuse the channel responses of the two branches through learnable channel gating factors to output deep semantic features. S6. The feature fusion module performs residual connections and dimension reconstruction on deep semantic features to generate global discriminative features. S7. The global discrimination features are used to obtain the component categories and concentration values ​​in the urine sample through the multi-task prediction module, and structured detection data is generated. S8. The urine component identification and detection model is updated and optimized through a joint optimization mechanism.

2. The intelligent urine component detection method based on spectral recognition and deep fusion network according to claim 1, characterized in that, The hyperspectral urine image data is a hyperspectral data cube containing spatial and spectral dimensions. It has a three-dimensional structure consisting of image height, image width, and the number of spectral bands. Each image pixel corresponds to a set of spectral reflectance intensity values ​​on a continuous spectral band.

3. The intelligent urine component detection method based on spectral recognition and deep fusion network according to claim 1, characterized in that, The preprocessing includes image calibration, background removal, spectral band filtering, and data normalization. Step S2 specifically involves: S21. Perform geometric distortion correction and pixel registration on the hyperspectral urine image to obtain a calibrated urine image with consistent spatial dimensions. S22. Based on the set light intensity threshold, mask image or segmentation algorithm, remove the background area in the calibrated urine image that does not contain a urine sample; S23. By evaluating the signal-to-noise ratio of each band and calculating the information entropy, bands with low signal-to-noise ratio, high information redundancy, or weak discrimination ability are obtained as effective spectral band sets. S24. Perform numerical standardization on the spectral vector of each pixel in the effective spectral channel set to unify the feature scale of each channel and generate standardized urine data. The size of the standardized urine data is [size missing]. ,in Indicates the image height. Indicates the image width. This indicates the number of valid spectral channels retained.

4. The intelligent urine component detection method based on spectral recognition and deep fusion network according to claim 1, characterized in that, Step S6 specifically includes: S61. Input the deep semantic features into the feature fusion module, which includes a main branch and an auxiliary branch; The main branch includes a channel compression mechanism and a channel weighting structure. The channel compression mechanism is as follows: through three sets of parallel... Convolution performs channel mapping operations with different compression ratios on deep semantic features to generate intermediate features at three scales; The channel weighting structure is configured with three sets of learnable channel response coefficients, which weight each intermediate feature according to the channel dimension and perform element-wise summation to output the main branch fusion feature. S62. Synchronously input the deep semantic features into the auxiliary branch with short-connection residual paths, through a set of Convolution aligns the channel dimensions of deep semantic features and outputs auxiliary features; S63. Perform residual connection on the main branch fusion feature and auxiliary feature according to the channel direction, and perform global average pooling in the spatial dimension to generate global discriminative features.

5. The intelligent urine component detection method based on spectral recognition and deep fusion network according to claim 1, characterized in that, Step S7 specifically includes: S71. Input the global discriminative features into the multi-task prediction module, and generate shared prediction features through a fully connected transformation structure; S72. Input the shared prediction features into two structurally separated component classification branches and concentration regression branches respectively. The component classification branch consists of two layers of nonlinear transformation units and a softmax output node, and outputs the component category probability vector. The concentration regression branch consists of two layers of nonlinear transformation units and a linear output node, and outputs the concentration estimation vector corresponding to each component. S73. Combine the component category probability vector with the concentration estimation vector to generate structured detection data. The structured detection data includes the detected component name, category, estimated concentration value, concentration unit, reference range, deviation level, feature contribution score, and sample number, and form a visualized detection result in a field-based format.

6. The intelligent urine component detection method based on spectral recognition and deep fusion network according to claim 1, characterized in that, Step S8 specifically includes: S81. Construct a multi-task loss function based on the component category probability vector and the concentration estimation vector. The multi-task loss function includes the cross-entropy loss of the component category probability vector and the mean square error loss of the concentration estimation vector. S82. Perform end-to-end model parameter updates on the urine component detection network model based on the multi-task loss function. The model parameters include the convolution kernel weights and gating parameters in the three-dimensional convolution feature extraction module, the attention mapping parameters and position encoding weights in the multi-head spectral spatial attention network module, the channel compression and residual fusion parameters in the feature fusion module, and the weights and bias terms of the fully connected layer in the multi-task prediction module.

Citation Information

Patent Citations

  • Crop pesticide residue detection method and system based on convolutional neural network

    CN116884512A

  • Soil pollutant identification and route tracking method and system based on artificial intelligence

    CN118397376A

  • 4D millimeter wave radar space occupation probability estimation method based on laser radar supervision

    CN120178229A

  • Construction method of adaptive gated spectrum-space-graph collaborative fusion network

    CN120472243A