Multi-mode high-throughput detection system and method based on mobile terminal
By integrating colorimetric, spectral, and fluorescence information into a multimodal high-throughput detection system, the problem of single-modal detection in intelligent mobile terminal biological detection devices has been solved, achieving high-precision, multi-index parallel detection, reducing the risk of misdiagnosis and improving detection efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- MINZU UNIVERSITY OF CHINA
- Filing Date
- 2026-01-15
- Publication Date
- 2026-05-26
AI Technical Summary
Existing biological detection devices based on smart mobile terminals suffer from problems such as single-modal detection, sensitivity to ambient light interference, high complexity, high cost, high risk of misdiagnosis, and limited parallel detection capabilities. Furthermore, they lack internal cross-validation mechanisms and cannot meet the needs of high-precision multi-index parallel detection.
A multimodal high-throughput detection system based on mobile terminals is adopted, which integrates colorimetric, spectral and fluorescence information. Through microfluidic detection chip and intelligent control module, cross-modal mutual verification and result confidence quantification are realized. Combined with the local surface plasmon resonance characteristics of gold nanoparticles, rapid and accurate quantitative detection of multiple samples or multiple indicators is performed.
This technology enables the simultaneous acquisition of colorimetric, spectral, and fluorescence information in a single test, reducing the risk of false positives/false negatives, increasing the throughput per unit time, reducing the amount of sample used, and improving the reliability and efficiency of the test.
Smart Images

Figure CN122084579A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of optical detection and intelligent biosensing technology, and in particular to a multimodal high-throughput detection system and method based on a mobile terminal. Background Technology
[0002] In recent years, point-of-care testing (POCT) technology has demonstrated enormous application potential in fields such as clinical diagnosis, environmental monitoring, and food safety. Smart mobile terminals, with their high-performance image sensors, powerful processing capabilities, and convenient data transmission functions, provide an ideal hardware foundation for building low-cost, portable biological testing systems. However, existing biological testing devices based on smart mobile terminals still face numerous technical bottlenecks, making it difficult to meet the clinical demands for high-precision, multi-indicator parallel testing.
[0003] Currently, point-of-care testing systems primarily employ single-modality detection methods such as colorimetry or fluorescence. While colorimetry is simple, it is susceptible to interference from ambient light and the white balance of mobile devices, resulting in poor repeatability. Fluorescence methods offer higher sensitivity but require additional optical components, increasing system complexity and cost, and are prone to false positives due to background fluorescence interference. More importantly, these single-modality systems lack internal cross-validation mechanisms, making it impossible to assess the reliability of test results and leading to a risk of misdiagnosis. Furthermore, existing smart mobile terminal testing devices are typically designed for a "one sample, one indicator" testing mode, which results in cumbersome, time-consuming, and reagent-intensive processes when performing multi-indicator testing, failing to meet the needs of rapid clinical screening. Although multi-indicator detection chips based on microarrays exist, their reading devices are mostly laboratory-grade scanners, incompatible with mobile terminal platforms.
[0004] The surface plasmon resonance (LSPR) properties of gold nanoparticles (AuNPs) have attracted much attention in the field of biosensing, with their peak redshift being correlated with analyte concentration. However, the redshift range caused by typical immunoassays is extremely narrow, thus requiring a spectrometer with a spectral resolution ≤4 nm for accurate LSPR peak resolution. Existing mobile terminal integrated spectral solutions, whether based on filter arrays or simple gratings, lack the spectral resolution required for accurate LSPR peak detection. In terms of data processing, existing mobile terminal detection systems largely rely on traditional image algorithms, which are sensitive to factors such as illumination and angle, exhibiting poor robustness. Although some studies have attempted to introduce deep learning, these are mostly limited to single-modal data processing, failing to fully utilize the complementarity of multimodal information and lacking quantitative evaluation of the reliability of detection results, thus limiting the system's credibility. Summary of the Invention
[0005] This invention provides a multimodal high-throughput detection system and method based on a mobile terminal, which addresses the shortcomings of existing portable detection devices that can only acquire single-modal signals, are sensitive to differences in ambient light and equipment, have limited parallel detection capabilities, and have difficulty in quantifying reliability. It enables the simultaneous acquisition of colorimetric, spectral, and fluorescence information in a single or time-division multiplexing acquisition, and completes rapid and accurate quantification and result reliability assessment of multiple samples or multiple indicators. It realizes high-throughput, multimodal, and high-resolution spectral detection of multiple target analytes in a single sample.
[0006] This invention provides a multimodal high-throughput detection system based on a mobile terminal, comprising a microfluidic detection chip, an optical illumination module, a spectral imaging module, and an intelligent control module. The intelligent control module controls the optical illumination module to provide a light source to the microfluidic detection chip and the spectral imaging module. The microfluidic detection chip carries the sample to be tested and, based on the light source provided by the intelligent control module, performs biometric identification on the carried sample and induces observable spectral characteristic changes in the local surface plasmon resonance peak position, generating a light signal. The spectral imaging module receives the light signal from the microfluidic detection chip and couples the light signal into spectral stripes to the imaging surface. The intelligent control module also acquires colorimetric images, fluorescence images, and spectral images based on the signals fed back from the microfluidic detection chip and the spectral imaging module, extracts optical features from the colorimetric, fluorescence, and spectral images, performs cross-modal fusion, predicts concentration and reliability indicators, and compares the reliability indicators with preset threshold values to determine the output concentration prediction result or trigger a re-inspection prompt.
[0007] According to the present invention, a multimodal high-throughput detection system based on a mobile terminal and a microfluidic detection chip are provided, comprising: a substrate layer; a flow channel layer disposed on the substrate layer, the flow channel layer being used to enclose and form a flow path channel with the substrate layer, the flow path channel including a waste liquid collection tank, a sample injection channel, and a buffer channel; wherein, the sample injection channel is used to deliver the sample to be tested to the detection area, and the buffer channel is used to introduce buffer solution to rinse the detection area; a detection site array disposed on the detection area of the substrate layer, the detection site array including detection sites arranged in a preset array, and each detection site including a gold nanoparticle disposed on the substrate layer and a functionalized layer disposed on the gold nanoparticle, the gold nanoparticle containing a trapping molecule for specifically binding to the target analyte in the sample to be tested.
[0008] According to the present invention, a multimodal high-throughput detection system based on a mobile terminal includes an intelligent control module comprising: a data acquisition unit, used to identify alignment marks or array corner points of microfluidic detection chips in the corresponding images based on input colorimetric images, fluorescence images, and spectral images, so as to perform array positioning, perspective correction, and lens distortion correction on the corresponding images, and identify the regions of interest corresponding to each detection site in the corrected images; and a calibration unit, used to perform pixel wavelength calibration on the spectral bands of the corrected spectral images, and to perform intensity normalization processing on the one-dimensional spectral vectors extracted from the spectral bands to obtain calibrated spectral data, and to perform calibration on the corrected colorimetric images. The image and fluorescence image are subjected to photometric normalization processing to obtain calibrated colorimetric and fluorescence data. The feature extraction unit is used to extract features from the calibrated spectral, colorimetric, and fluorescence data respectively. The cross-modal fusion unit is used to perform weighted fusion of the extracted spectral, colorimetric, and fluorescence features using a preset gated fusion mechanism, and predict the concentration prediction value and reliability index based on the fusion features. The reliability evaluation unit is used to compare the reliability index with a preset index threshold. If the reliability index does not reach the preset index threshold, a re-examination prompt is triggered. If the reliability index reaches the preset index threshold, the concentration prediction result is output.
[0009] According to the present invention, a multimodal high-throughput detection system based on a mobile terminal includes a cross-modal fusion unit, which further comprises: a modality-specific encoder for encoding extracted spectral features, colorimetric features, and fluorescence features into aligned features; a gating subunit for determining the fusion weight of the corresponding modality based on the aligned spectral features, colorimetric features, and fluorescence features; a fusion layer for weighted fusion of the aligned spectral features, colorimetric features, and fluorescence features according to the fusion weight of each modality to obtain fused features; and an output head for predicting concentration values based on the fused features and determining reliability indicators based on the predicted concentration values.
[0010] According to the multimodal high-throughput detection system based on a mobile terminal provided by the present invention, the cross-modal fusion unit is further configured to: utilize Monte Carlo random inactivation Dropout to perform a preset number of forward propagations on the fusion network adopted by the cross-modal fusion unit to obtain a corresponding number of prediction results, and estimate the model uncertainty based on multiple sets of prediction results; wherein, any set of prediction results includes the concentration prediction values output by the single-modal regression head corresponding to the spectrum, colorimetry, and fluorescence, respectively; determine the weighted dispersion based on the prediction results to obtain the modal consistency index; and obtain the reliability index based on the model uncertainty, the modal consistency index, and the signal-to-noise ratio in the prior fluorescence feature.
[0011] According to the multimodal high-throughput detection system based on a mobile terminal provided by the present invention, the feature extraction unit is further configured to: extract local surface plasmon resonance peak position features from calibrated spectral data using a spectral feature extraction network to obtain spectral features; extract absorbance from calibrated colorimetric data using a colorimetric feature extraction network to obtain colorimetric features; and extract fluorescence intensity and signal-to-noise ratio from calibrated fluorescence data using a fluorescence feature extraction network to obtain fluorescence features.
[0012] According to the present invention, a multimodal high-throughput detection system based on a mobile terminal includes an optical illumination module comprising a white light source, a fluorescent excitation source, an excitation-side filter assembly, an emission-side filter assembly, a first condenser lens, a second condenser lens, a diffuser, a reflector, and a dichroic mirror module. The white light source is used for colorimetric and spectral detection; the fluorescent excitation source is used for fluorescence detection; the excitation-side filter assembly is disposed on the fluorescent excitation light path and located between the second condenser lens and the fluorescent excitation source, the second condenser lens being used to converge the illumination beam emitted by the fluorescent excitation source; the emission-side filter assembly is disposed on the emission light path; the first condenser lens is used to converge the illumination beam emitted by the white light source; the diffuser is disposed between the white light source and the first condenser lens; the reflector is disposed in the white light illumination channel; and the dichroic mirror module is disposed between the emission-side filter assembly and the microfluidic detection chip for separating and switching the excitation and emission light. The white light source has a color temperature range of 4000–6500 K and covers the visible spectrum, and the fluorescent excitation source has a center wavelength range of 450–550 K. nm and, together with excitation-side and emission-side filter components, suppress stray light.
[0013] According to the present invention, a multimodal high-throughput detection system based on a mobile terminal is provided. The spectral imaging module includes an entrance slit, a diffraction grating, a focusing lens, and a spectral imaging region. The width of the entrance slit is 50 µm-150 µm, the groove density of the diffraction grating is greater than or equal to 1200 lines / mm, and the focusing lens includes a cemented doublet achromatic lens.
[0014] This invention also provides a multimodal high-throughput detection method based on a mobile terminal, applicable to any of the above-mentioned multimodal high-throughput detection systems based on a mobile terminal. The method includes: controlling an optical illumination module to provide a light source to a microfluidic detection chip and a spectral imaging module; acquiring colorimetric images, fluorescence images, and spectral images based on signals fed back from the microfluidic detection chip and the spectral imaging module; extracting optical features based on the colorimetric images, fluorescence images, and spectral images and performing cross-modal fusion to predict concentration and reliability indicators; comparing the reliability indicators with preset indicator thresholds to determine the output concentration prediction result or trigger a re-inspection prompt.
[0015] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the multimodal high-throughput detection method based on a mobile terminal as described above.
[0016] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the multimodal high-throughput detection method based on a mobile terminal as described above.
[0017] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the multimodal high-throughput detection method based on a mobile terminal as described above.
[0018] The multimodal high-throughput detection system and method based on mobile terminals provided by this invention uses an intelligent control module as a unified acquisition and computing platform. By integrating colorimetric, spectral, and fluorescence information on the same optical axis or equivalent optical path, it achieves cross-modal mutual verification and result confidence quantification, reducing the risk of false positives / false negatives caused by a single modality. Furthermore, by forming a multi-point detection array on the reaction layer substrate using a microfluidic chip, it enables parallel processing of multiple samples / indicators, significantly increasing throughput per unit time and reducing the amount of data used per sample. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0020] Figure 1 This is a schematic diagram of the structure of the multimodal high-throughput detection system based on a mobile terminal provided by the present invention; Figure 2 This is one of the structural schematic diagrams of the microfluidic detection chip provided by the present invention; Figure 3 This is the second schematic diagram of the microfluidic detection chip provided by the present invention; Figure 4 This is a schematic diagram of the structure of the optical illumination module provided by the present invention; Figure 5 This is a schematic diagram of the structure of the spectral imaging module provided by the present invention; Figure 6 This is a logical schematic diagram of the intelligent control module provided by the present invention; Figure 7 This is a schematic diagram of the network architecture of the intelligent control module provided by the present invention; Figure 8 This is a flowchart illustrating the multimodal high-throughput detection method based on a mobile terminal provided by the present invention. Figure 9 This is a schematic diagram of the structure of the electronic device provided by the present invention.
[0021] Figure label: 100: Intelligent control module; 110: Camera; 200: Microfluidic detection chip; 210: Substrate layer; 220: Flow channel layer; 230: Alignment marker; 240: Detection site array; 300: Detection site; 310: Gold nanoparticles; 320: Functionalized layer; 330: Capture molecule; 400: Optical illumination module; 410: White light source; 420: Fluorescent excitation source; 430: Excitation-side filter assembly; 440: Emission-side filter assembly; 450: First condenser lens / Second condenser lens; 460: Dichroic mirror module; 470: Diffuser; 480: Mirror; 500: Spectral imaging module; 510: Entrance slit; 520: Diffraction grating; 530: Focusing lens; 540: Spectral imaging region. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0023] Figure 1 This is a schematic diagram of the structure of the multimodal high-throughput detection system based on a mobile terminal provided by the present invention, as shown below. Figure 1 As shown, the system includes a microfluidic detection chip 200, an optical illumination module 400, a spectral imaging module 500, and an intelligent control module 100, wherein: The intelligent control module 100 is used to control the optical illumination module 400 to provide a light source to the microfluidic detection chip 200 and the spectral imaging module 500. The microfluidic detection chip 200 is used to carry the sample to be tested and, based on the light source provided by the intelligent control module 100, performs biometric identification on the sample to be tested and induces observable spectral feature changes in the local surface plasmon resonance peak position to generate an optical signal. The spectral imaging module 500 is used to receive the optical signal from the microfluidic detection chip 200 and couple the optical signal into spectral stripes to the imaging surface. The intelligent control module 100 is also used to acquire colorimetric images, fluorescence images and spectral images based on the signals fed back by the microfluidic detection chip 200 and the spectral imaging module 500, extract optical features based on the colorimetric images, fluorescence images and spectral images and perform cross-modal fusion, predict concentration and reliability indicators, and compare the reliability indicators with preset indicator thresholds to determine the output concentration prediction result or trigger a re-inspection prompt.
[0024] Specifically, refer to Figures 2-3 A microfluidic detection chip includes: a substrate layer 210; a flow channel layer 220 disposed on the substrate layer 210, the flow channel layer 220 being used to enclose the substrate layer 210 to form a flow path channel, the flow path channel including a waste liquid collection tank, a sample injection channel and a buffer channel; wherein, the sample injection channel is used to deliver the sample to be tested to the detection area, and the buffer channel is used to introduce buffer solution to rinse the detection area; a detection site array 240 disposed on the detection area on the substrate layer 210, the detection site array 210 including detection sites 300 arranged in a preset array, and each detection site including gold nanoparticles 310 disposed on the substrate layer 210 and a functionalized layer 320 disposed on the gold nanoparticles, the gold nanoparticles 320 containing trapping molecules 330 for specifically binding to target analytes in the sample to be tested.
[0025] Furthermore, the microfluidic detection chip also includes alignment marks set at preset positions on the substrate layer for array positioning and subsequent image correction.
[0026] In the actual processing of the test sample, the test sample or diluted test sample is introduced into the flow path channel through the sample dispensing port of the microfluidic detection chip. Under the distribution and guiding structure of the flow path channel, the test sample is distributed to multiple detection sites in the detection site array, and specific binding reactions occur with fixed capture molecules at each detection site.
[0027] Specifically, probe molecules labeled with fluorescent or enzyme reporter molecules and chromogenic substrate solutions are sequentially introduced into the flow path channel to generate three detection responses—spectral, colorimetric, and fluorescence—at the same detection site. Furthermore, the detection site on the substrate layer includes at least a localized surface plasmon nanostructure region composed of gold nanoparticles, a colorimetric reaction region for generating a colorimetric signal, and / or a fluorescently labeled reaction region for generating a fluorescence signal. The surface plasmon nanostructure region induces observable spectral characteristic changes upon biorecognition, the colorimetric reaction region generates color or density differences after the colorimetric reaction, and the fluorescently labeled reaction region generates a fluorescence emission signal under excitation light.
[0028] Preferably, the surface plasmon nanostructure region, the colorimetric reaction region, and the fluorescent labeling reaction region are arranged on the same detection site to achieve mutual verification of the three detection modes of spectroscopy, colorimetry, and fluorescence in the same sample and at the same spatial site.
[0029] Furthermore, the flow path is suitable for distributing and replacing samples and reagents inside the chip through pressure drive, negative pressure suction, capillary force, or a combination thereof.
[0030] In one alternative embodiment, the base layer is a quartz substrate with a thickness of 1.0 mm to 2.0 mm. The size range can be from 15 mm × 15 mm to 25 mm × 25 mm. In this embodiment, the thickness can be 1.5 mm, and the size range can be 20 mm × 20 mm.
[0031] Furthermore, the detection site array integrates an array of gold nanoparticles (AuNPs) on the substrate surface to form N×M detection sites. Each detection site can be 100 µm × 100 µm to 1000 µm × 1000 µm in size, with adjacent sites spaced 100 µm to 1000 µm apart. The average particle size of the gold nanoparticles is selectable within a wide range, from 20 nm to 80 nm, and the initial value of its localized surface plasmon resonance (LSPR) peak is in the 510 nm to 530 nm wavelength range. In this embodiment, the detection site array can employ a 15 × 15 format with a total of 225 independent detection sites. Each detection site is 500 µm × 500 µm in size, with an adjacent site spaced 500 µm apart and a center-to-center distance of 1000 µm. The average particle size of the gold nanoparticles is preferably 40 nm, and the initial value of its LSPR peak can be 520 nm.
[0032] In addition, after the immune response occurs, the target analyte binds to the AuNP surface, causing a local change in refractive index, which leads to a redshift of the LSPR peak position. The typical redshift range is 5–30 nm, and the amount of redshift increases monotonically with the concentration of the target analyte.
[0033] Furthermore, the flow channel layer is fabricated by bonding a flow channel cap, also made of quartz or glass, onto the substrate layer. This cap is fabricated using photolithography and wet etching techniques to create the flow channel layer, which has a thickness of 0.3 mm. Further, the sample inlet channel has a width of 2.0 mm and a height of 300 µm, arranged in a Y-shape with the detection area, supporting sample volumes of 50–200 µL; the buffer channel has a width of 1.5 mm and a height of 300 µm, used for post-reaction rinsing and signal stabilization; and the waste liquid collection tank, with a volume of approximately 100 µL, is located at the chip outlet to prevent backflow contamination.
[0034] The alignment marks are fabricated by placing a 500 µm diameter circular black mark at each of the four corners of the chip. These serve as key markers for GridDet network positioning and identification, facilitating the subsequent data acquisition module's identification of the center coordinates of the four alignment marks and the coordinates of the four corner points of the detection site array, totaling eight key points. These points are used to calculate the homography matrix H and achieve accurate perspective correction and ROI positioning. See below for details. The alignment marks can be formed by screen printing or laser etching of the black ink layer, with a positioning accuracy error ≤5 µm. Additionally, a cross-shaped auxiliary mark can be placed at the intersection of the diagonals of the detection area for manual alignment calibration.
[0035] Specifically, the optical illumination module is positioned around the camera of the intelligent control module to form an illumination and filtering optical path, as shown in the reference. Figure 4 The optical illumination module includes a white light source 410, a fluorescent excitation source 420, an excitation-side filter assembly 430, an emission-side filter assembly 440, a first condenser lens 450, a second condenser lens 450, a diffuser 470, a reflector 480, and a dichroic mirror module 460. The white light source 410 is used for colorimetric and spectral detection; the fluorescent excitation source 420 is used for fluorescence detection; and the excitation-side filter assembly 430 is disposed on the fluorescent excitation light path and located between the second condenser lens 450 and the fluorescent excitation source. Between 420, the second condenser lens 450 is used to converge the illumination beam emitted by the fluorescent excitation source; the emission-side filter assembly 440 is disposed in the emission light path; the first condenser lens 450 is used to converge the illumination beam emitted by the white light source; the diffuser 470 is disposed between the white light source 410 and the first condenser lens 450; the reflector 480 is disposed in the white light illumination channel; and the dichroic mirror module 460 is disposed between the emission-side filter assembly 440 and the microfluidic detection chip 200 to separate and switch the excitation light and the emission light.
[0036] It should be added that the excitation-side filter assembly can be an excitation-side bandpass filter, and the emission-side filter assembly can be an emission-side longpass filter.
[0037] Furthermore, the white light source preferably uses broadband white LEDs with a color temperature range of 4000 K to 6500 K and covering the visible spectrum. The output power of the white light source is adjustable, ranging from 50 to 500 mW, and the brightness can be controlled by current modulation or pulse width modulation.
[0038] In a preferred embodiment, the white light source is a white LED with a color temperature of about 5000 K, whose spectrum covers the visible light band of about 400 to 750 nm.
[0039] To improve illumination uniformity, a diffuser 470 is installed at the white light source emission end to homogenize the light beam. Then, the homogenized light beam is collimated and moderately focused by the first condenser lens 450. The white light illumination beam, after being shaped by the first condenser lens 450, is then refracted by a reflector 480 set at an inclined angle and incident on the detection area of the microfluidic detection chip 200, thereby achieving uniform illumination of the chip surface.
[0040] The reflector 480 is a non-filtering planar reflective element, which is only used to change the propagation direction of the white light illumination path. It does not perform wavelength selective filtering on the light passing through, and its position is staggered from the imaging light path of the intelligent control module camera 110 to avoid blocking or affecting the signal light reflected or transmitted back to the camera 110 by the microfluidic detection chip 200.
[0041] The fluorescence excitation source 420 can be designed with a single wavelength or a combination of multiple wavelengths. The center wavelength range of the fluorescence excitation source is 450–550 nm, and it is used in conjunction with excitation-side filter components and emission-side filter components to suppress stray light.
[0042] In one embodiment, to accommodate different types of fluorescent dyes, the fluorescent excitation light source 420 employs a switchable dual-wavelength structure. The first wavelength channel uses a blue LED with a center wavelength between 450 and 490 nm, and the second wavelength channel uses a green LED with a center wavelength between 510 and 540 nm. The output power of a single fluorescent excitation LED is adjustable, preferably in the range of 50–300 mW, to achieve approximately 0.5–3 mW / cm² on the chip surface. 2 The excitation optical density.
[0043] After the fluorescent excitation light is filtered by the excitation-side bandpass filter 430 to remove stray light and broadband background light, it is focused by the second condenser lens 450 and refracted by the dichroic mirror module 460 to the gold nanoparticle array region on the microfluidic detection chip 200, thereby achieving effective excitation of the sample to be tested. The fluorescent emission light emitted from the microfluidic detection chip 200 is transmitted through the dichroic mirror module 460 and further filtered by the emission-side long-pass filter 440 to remove residual excitation light and background light before entering the imaging system of the camera 110, thereby significantly improving the signal-to-noise ratio of the fluorescence signal.
[0044] Preferably, the illumination geometry adopts an annular dark field or oblique incidence arrangement, so that the angle between the multiple light sources and the normal of the microfluidic detection chip 200 is approximately 45°±5°, in order to reduce saturation and glare caused by strong specular reflections directly entering the camera 110. During multimodal (colorimetric / fluorescence / spectral) switching, the corresponding optical path can be selectively connected through the dichroic mirror module 460 or other optical path switching devices to achieve rapid switching of detection modes. The optical path switching time is preferably no higher than 50ms to meet the real-time imaging and analysis requirements in high-throughput detection processes.
[0045] In addition, refer to Figure 5 The spectral imaging module includes an entrance slit 510, a diffraction grating 520, a focusing lens 530, and a spectral imaging region 540. The width of the entrance slit is 50 µm-150 µm, the groove density of the diffraction grating is greater than or equal to 1200 lines / mm, and the focusing lens includes a cemented doublet achromatic lens.
[0046] Furthermore, the camera combination of the spectral imaging m-module and the intelligent control module is configured to achieve a spectral resolution of no more than 4 nm in the visible light band, in order to resolve the redshift of spectral peaks in the range of 5–30 nm caused by local surface plasmon resonance at the detection site.
[0047] It should be added that the entrance slit can be fabricated using laser etching technology to create a rectangular slit with a width of 100 µm and a height of 8 mm on a 0.1 mm thick black anodized aluminum sheet. The positioning accuracy of the slit center and the optical axis is adjusted by a precision displacement stage, with a deviation of ≤10 µm. The diffraction grating can be a transmission-type blazed grating with a typical groove density of 1200 lines / mm and a blaze wavelength of approximately 550 nm, operating in first-order diffraction. The grating is mounted using a precision rotating stage, and the incident angle can be finely adjusted within the range of -10° to +10° to optimize the alignment of the dispersion direction with the CMOS pixel arrangement of the camera. The focusing lens is a cemented doublet achromatic lens with a typical focal length of f=35 mm, a light-transmitting aperture ≥15 mm, and an F-number of approximately F / 2.8. The lens is aligned with the grating and the optical axis of the mobile terminal camera through a precision lens barrel to ensure spectral imaging quality.
[0048] In addition, the spectral resolution is calculated based on the grating equation and resolution formula, where the typical slit width is 100 µm, the collimation focal length is 35 mm, the grating constant is 833 nm, and the diffraction angle is about 20°. The theoretically calculated resolution is about 2.2 nm. The system can achieve a spectral resolution of ≤ 5 nm in the target working band, preferably no higher than about 4 nm, to meet the detection requirements of fine spectral features such as the peak position of localized surface plasmon resonance (LSPR).
[0049] Before the intelligent control module acquires images, the microfluidic detection chip (200) that has completed the reaction is placed in the dark box or bracket assembly, which precisely aligns the microfluidic detection chip with the camera of the optical illumination module, the spectral imaging module and the intelligent control module.
[0050] The signal recording and analysis application running on the intelligent control module is launched to automatically control the collaborative work of the optical illumination module and the camera according to a preset time-division multiplexing sequence. The specific acquisition process includes: controlling the white light source to turn on, providing an illumination source covering the visible light range for colorimetric detection and spectral analysis; after a stabilization time, preferably 100 ms to ensure stable light source brightness, the camera performs autofocus or uses a preset focal length to acquire the first frame of global image I. white Global Image I white The system fully records the colorimetric information of the detection site array on the microfluidic detection chip and the spectral band information formed in the spectral imaging region after dispersion by the spectral imaging module, obtaining colorimetric and spectral images. It controls the white light source to be turned off and immediately enters a preset anti-crosstalk delay stage, preferably lasting 5 ms to 10 ms. During this period, all light sources remain off to ensure that the afterglow of the white light source is completely extinguished and the camera has completed I / O. white The readout of frames prevents white light signals from interfering with subsequent fluorescence images.
[0051] Furthermore, after the anti-crosstalk delay ends, the control immediately turns on the fluorescence excitation source to provide excitation light of a specific wavelength for fluorescence detection; after a stabilization period, preferably 100 ms, the camera, in conjunction with the emission-side filter component, acquires the second frame of fluorescence image I. fluo Fluorescence diagram I fluo The image only contains signals emitted by fluorescent markers at the detection sites; after acquisition, all light sources are turned off to end the current detection cycle.
[0052] In this embodiment, to ensure that the camera's exposure window falls precisely within the time period when the light source is stably lit, the synchronization jitter between the application's light source control command and the camera's exposure trigger signal is strictly controlled.
[0053] Preferably, the synchronization jitter is no more than 5 ms, more preferably no more than 1 ms, in order to avoid acquiring unstable signals during the process of the light source being turned on or off.
[0054] Optionally, if the detection scheme includes multiple fluorescent dyes with different emission wavelengths, the application can control the sequential switching of different wavelengths of fluorescence excitation light sources, such as excitation light with a center wavelength of 470 nm or excitation light of 525 nm, and acquire corresponding multiple frames of fluorescence images to achieve multicolor fluorescence detection and analysis.
[0055] Optionally, to improve the signal-to-noise ratio, before acquiring spectral images and colorimetric or fluorescence images, the application can control the camera to acquire one or more frames of dark background images under completely dark conditions for subsequent background subtraction and intensity correction.
[0056] It is worth noting that the camera in the intelligent control module needs to meet the requirements of high-sensitivity multimodal image acquisition. The camera should have a resolution of at least 8 megapixels, a sensor size selectable from 1 / 2.0 to 1 / 3.0 inches, and a pixel size ranging from 0.8 to 2.0 µm. The lens aperture is preferably set in the range of F / 1.8 to F / 2.2 to ensure sufficient light intake under low-light conditions, especially during fluorescence detection. The lens focal length and field of view can be adapted according to the overall design of the optical module.
[0057] In one embodiment, the equivalent focal length can be 26-28 mm, the field of view can be no less than 75°, and the system preferably supports autofocus and high dynamic range imaging. More preferably, the system supports RAW format output to retain maximum information for subsequent algorithm processing.
[0058] Furthermore, the image acquisition parameters can be adaptively adjusted according to different detection modes, including colorimetric, fluorescence, and spectral modes. When acquiring colorimetric images, a preset white balance mode can be set, and an exposure time in the range of 1 / 100 to 1 / 30 s and an ISO sensitivity in the range of 100 to 400 can be selected to avoid overexposure. When acquiring fluorescence images, white balance can be turned off, and an exposure time in the range of 0.1 to 1 s and an ISO sensitivity in the range of 800 to 3200 can be selected according to the fluorescence signal intensity. A noise reduction algorithm can also be optionally enabled. In addition, spectral images can be acquired simultaneously with colorimetric images. The system extracts spectral bands from a specific imaging region 540, the wavelength range of which preferably covers visible light, and can be 405 to 785 nm. The system's spectral resolution is preferably 4 nm or better, and in some embodiments, it can be 3.5 to 4.0 nm.
[0059] In addition, the intelligent control module can use intelligent mobile terminals, such as smartphones, to achieve timing control with millisecond (ms) precision by utilizing the internal clock source of the mobile terminal. The hardware requirements for the smartphone terminal, including memory and storage space, only need to meet the requirements for the smooth operation of the application and deep learning model inference.
[0060] In this embodiment, reference Figure 6The intelligent control module includes: a data acquisition unit, used to identify alignment marks or array corner points of the microfluidic detection chip in the corresponding images based on the input colorimetric image, fluorescence image, and spectral image, so as to perform array positioning, perspective correction, and lens distortion correction on the corresponding images, and identify the region of interest corresponding to each detection site in the corrected image; and a calibration unit, used to calibrate the pixel wavelengths of the spectral bands of the corrected spectral image, and to perform intensity normalization processing on the one-dimensional spectral vector extracted from the spectral bands to obtain the calibrated spectral data, and to perform photometric analysis on the corrected colorimetric image and fluorescence image respectively. The system performs normalization to obtain calibrated colorimetric and fluorescence data; a feature extraction unit extracts features from the calibrated spectral, colorimetric, and fluorescence data respectively; a cross-modal fusion unit uses a preset gating fusion mechanism to perform weighted fusion of the extracted spectral, colorimetric, and fluorescence features, and predicts the concentration prediction value and reliability index based on the fusion features; a reliability assessment unit compares the reliability index with a preset threshold, triggers a re-inspection prompt when the reliability index does not reach the preset threshold, and outputs the concentration prediction result when the reliability index reaches the preset threshold.
[0061] It should be noted that the data acquisition unit can use a GridDet network for detection. When performing detection using the above embodiment, the original image from the camera is first input into the GridDet network. At this time, the GridDet network realizes array positioning and perspective correction of the original image by identifying alignment marks or array corners on the microfluidic detection chip, estimating geometric transformations such as homography matrix, and outputting the corrected image and a list of coordinates of the region of interest corresponding to each detection point.
[0062] In addition, the GridDet network analyzes spectral and colorimetric images, automatically estimates and calculates the homography matrix H by identifying alignment marks (230) or array corners on the microfluidic detection chip, and uses the homography matrix H to perform perspective correction on two frames of images, namely spectral, colorimetric and fluorescence images, to eliminate geometric deformation caused by the tilt of the shooting angle.
[0063] Optionally, if lens distortion correction is required, before perspective correction, the intrinsic parameters obtained from camera calibration are used. These intrinsic parameters include radial distortion coefficients k_1, k_2, k_3 and tangential distortion coefficients p_1, p_2. The original image is then distorted using the intrinsic parameters to eliminate radial and tangential lens distortion, and then homography transformation is performed to achieve perspective correction.
[0064] Specifically, refer to Figure 7The GridDet network is used for array localization and geometric correction. The input to the GridDet network is the original high-resolution image with a resolution of 1920×1080×3. The image is first preprocessed, proportionally scaled and padded to a uniform size of 640×640, and lens distortion correction is performed. Its backbone network adopts a lightweight inverse residual structure, which contains a series of inverse residual blocks. The feature maps extracted by the backbone network are then passed through global average pooling and fed into fully connected layers. These fully connected layers consist of a 1024-dimensional fully connected layer followed by a 16-dimensional output fully connected layer. The output of the GridDet network is the coordinates (x1, y1)...(x8, y8) of eight keypoints in the preprocessed image coordinate system. These eight keypoints include the four alignment marker centers of the chip and the four corner points of the 15x15 detection array.
[0065] Furthermore, by utilizing the coordinates of the eight predicted keypoints output by the GridDet network and their corresponding coordinates on the ideal template, the homography matrix H is calculated. This matrix H is then used to perform a perspective transformation on the original high-resolution image, resulting in a corrected "frontal view" image I. corrected On the corrected image, based on the known 15x15 array layout, the global image is precisely segmented into 225 independent regions of interest (ROIs) to facilitate parallel and independent quantitative analysis of the signal at each detection site. For each ROI, three-modal data are extracted: a 512×1 dimensional spectral vector is extracted from the spectral channel, a 64×64×3 dimensional colorimetric vector is extracted from the colorimetric channel, and a 64×64×1 dimensional fluorescence vector is extracted from the fluorescence channel.
[0066] To elaborate further, the training loss function of the GridDet network Using smoothed L1 loss, it can be expressed as: in, Indicates the first The true coordinates of each key point This represents the coordinates predicted by the GridDet network.
[0067] In addition, the calibration unit preferably uses a CalibNet network to receive the corrected image output by the GridDet network. The CalibNet network includes two sub-networks: CalibNet-λ and CalibNet-I. CalibNet-λ is used to perform pixel-to-wavelength mapping calibration on the spectral imaging region, while CalibNet-I is dedicated to intensity normalization correction of the spectral vector (1D). For the colorimetric and fluorescence channels, the system uses conventional image processing methods for photometric normalization and does not use the CalibNet-I network.
[0068] It should be added that CalibNet-λ's function is to implement pixel-wavelength mapping. Its structure uses a one-dimensional convolutional network, taking a 512-dimensional sequence of pixel positions as input. This input passes through three one-dimensional convolutional layers and two fully connected layers, ultimately outputting a 512-dimensional wavelength mapping vector. Its loss function... The mean squared error is expressed as: in, Indicates the sequence length. Indicates the first The true wavelength corresponding to each pixel position can be obtained through offline calibration using a standard light source. This represents the wavelength predicted by the CalibNet-λ network.
[0069] Furthermore, CalibNet-I is specifically designed for correcting the intensity of one-dimensional spectral vectors. Its structure employs a multilayer perceptron, taking a 512-dimensional spectral vector as input, passing it through three fully connected layers, and outputting a corrected 512-dimensional spectral vector. Its loss function... Including MSE and structural similarity loss SSIM, it is expressed as: in, Represents the target spectral vector. This represents the network output spectral vector. This represents the loss weighting coefficient. To compute structural similarity on a one-dimensional spectral vector, a sliding window approach is used. The CalibNet-I network uses paired raw / flat-field corrected spectral vectors as training data to achieve adaptive intensity correction across conditions.
[0070] Furthermore, CalibNet-λ uses a standard light source with known peak positions for pixel-to-wavelength calibration. The CalibNet-λ neural network takes a pixel index sequence (N=512) as input and directly regresses a wavelength mapping vector of the same length. The specific process is as follows: A mercury lamp spectrum is captured, with characteristic peak positions at 435.8 nm, 546.1 nm, 577.0 nm, and 579.1 nm, to extract the pixel coordinates of each peak position on the CMOS; a least-squares fit is performed using a cubic polynomial to obtain the wavelength label vector λ corresponding to each pixel position; during network training, this wavelength vector is directly regressed, rather than the polynomial coefficients. It should be noted that the standard light source is only used for offline generation of supervised labels; during runtime, the network directly inputs 512 pixel indices and outputs 512 wavelength values, eliminating the need for manual calibration each time.
[0071] In addition, CalibNet-I is only used for intensity normalization correction of spectral vectors. The input is a 512-dimensional spectral vector after wavelength calibration of a single ROI, and the output is a 512-dimensional spectral vector after intensity correction.
[0072] Furthermore, the photometric normalization of the colorimetric and fluorescence channels employs traditional image processing methods, including white balance correction, flat field correction, dark field correction, and gain correction, without relying on CalibNet-I prediction. Further, the colorimetric channel undergoes white balance, flat field correction, and gamma correction; the fluorescence channel undergoes dark current subtraction, flat field correction, and gain correction.
[0073] Specifically, the feature extraction unit is also used to: extract local surface plasmon resonance peak position features from calibrated spectral data using a spectral feature extraction network to obtain spectral features; extract absorbance from calibrated colorimetric data using a colorimetric feature extraction network to obtain colorimetric features; and extract fluorescence intensity and signal-to-noise ratio from calibrated fluorescence data using a fluorescence feature extraction network to obtain fluorescence features.
[0074] Furthermore, using a spectral feature extraction network, local surface plasmon resonance peak features are extracted from the calibrated spectral data to obtain spectral features, including: for each ROI, the corresponding spectral strips are extracted from the spectral imaging region (540) of the global image, and the dimensionality is reduced to 1D spectral vector by weighted integration along the slit direction; Gaussian weighted summation is used for integration along the vertical direction, and if the strip is tilted, the center line tracking algorithm is first used to determine the center line of the strip and rotate it for correction, and then the intensity values of 512 pixels are extracted along the horizontal direction; when the effective length of the strip in the horizontal direction is... When the pixel coordinate system is in use, it is first resampled at equal intervals using linear interpolation to N=512 points; for out-of-bounds areas, a boundary copying strategy is used.
[0075] Preferably, it is made by CalibNet- First, pixel-wavelength calibration is performed to map the pixel sequence to a wavelength sequence. Then, CalibNet-I performs intensity normalization correction on the spectral vector to eliminate flat-field differences and background drift. Subsequently, the SpecNet network processes the corrected spectral vector to extract spectral features related to LSPR. Preferably, the feature is the LSPR peak wavelength. .
[0076] A colorimetric feature extraction network is used to extract absorbance from calibrated colorimetric data to obtain colorimetric features, including: extracting absorbance from global images. The ROI data for colorimetric regions are normalized using traditional image processing methods such as white balance correction, flat field correction, and gamma correction. The normalized ROI image is then processed by the ColorNet network to extract colorimetric features related to color or density. Preferably, the feature is absorbance. .
[0077] A fluorescence feature extraction network is used to extract fluorescence intensity and signal-to-noise ratio from calibrated fluorescence data to obtain fluorescence features. This includes: normalizing the photometric intensity of the ROI data from the fluorescence image using methods such as dark current subtraction, flat-field correction, and gain correction; and processing the normalized ROI image using the FluoNet network to extract fluorescence emission features. Preferably, the features are the integrated fluorescence intensity I_fluo and the signal-to-noise ratio SNR.
[0078] In addition, for each of the 225 ROIs, the extracted three-modal features λ_peak, A, I_fluo, and SNR are combined into a multi-dimensional feature vector.
[0079] It should be added that the feature extraction unit includes a parallel spectral feature extraction network (SpecNet), a colorimetric feature extraction network (ColorNet), and a fluorescence feature extraction network (FluoNet). Specifically, the SpecNet network is used to regress the LSPR peak positions from the corrected 512-dimensional spectral curves. Its structure employs a 1D-CNN based on one-dimensional residual blocks. The network consists of four stacked residual blocks. Within each residual block, when the number of input and output channels does not match, a skip connection is used to match the dimensions via a 1x1 projective convolution. Finally, the output is regressed through global average pooling and fully connected layers. The ColorNet network is used to regress absorbance from a normalized 64×64×3 colorimetric ROI. Its structure employs a lightweight 2D-CNN, consisting of three stacked convolutional blocks for downsampling and feature extraction. Finally, it regresses one-dimensional absorbance through global average pooling and fully connected layers. The FluoNet network is used to regress fluorescence intensity I_fluo and signal-to-noise ratio (SNR) from a normalized 64×64×1 fluorescent ROI. Its structure is similar to ColorNet, employing a lightweight 2D-CNN, with the final fully connected layer outputting a two-dimensional vector.
[0080] Additionally, SpecNet's loss function Using the negative log-likelihood loss with uncertainty, it is expressed as: in, Indicates batch size. This represents the log-variance of the network's prediction for the j-th sample. Indicates the true peak position label. This indicates the peak position predicted by the SpecNet network.
[0081] Loss function of ColorNet network The mean square error (MSE) is used, and it is expressed as: in, Indicates batch size. This represents the true absorbance label of the j-th sample. This represents the absorbance predicted by the ColorNet network.
[0082] The loss function of the FluoNet network For multi-tasking MSE loss: in, and This represents the loss weighting coefficient. and These represent the true values of fluorescence intensity and signal-to-noise ratio, respectively. and This represents the corresponding network prediction value.
[0083] In an optional embodiment, before inputting the extracted spectral, colorimetric, and fluorescence features into the cross-modal fusion unit, the features from the three-modal feature extraction network are z-score normalized or Min-Max normalized to map each feature quantity to the same numerical range, thus obtaining the corresponding... .
[0084] Specifically, the z-score standardization formula is: ,in and Let SNR be the mean and standard deviation of the feature, respectively. The standardized SNR is denoted as SNR. ,and , , Maintaining consistency ensures that the features of different modalities are comparable, facilitating weighted fusion in subsequent fusion networks.
[0085] In addition, the cross-modal fusion unit also includes: a modality-specific encoder for encoding the extracted spectral features, colorimetric features, and fluorescence features into aligned features; a gating subunit for determining the fusion weights of the corresponding modes based on the aligned spectral features, colorimetric features, and fluorescence features; a fusion layer for weighted fusion of the aligned spectral features, colorimetric features, and fluorescence features based on the fusion weights of each mode to obtain fused features; and an output head for predicting concentration values based on the fused features and determining reliability indicators based on the predicted concentration values.
[0086] It is evident that by integrating a series of deep learning networks such as GridDet, CalibNet, SpecNet, ColorNet, FluoNet, and FusionNet, fully automated intelligent analysis from raw image acquisition to final quantitative results and reliability assessment has been achieved, providing technical support for high-throughput parallel detection, multimodal cross-validation, and high-reliability result output.
[0087] It should be added that the cross-modal fusion unit preferably adopts a fused FusionNet network, which receives the trimodal feature vectors output by SpecNet, ColorNet, and FluoNet networks; the FusionNet network uses a gated fusion mechanism to perform weighted fusion of the trimodal features and outputs the final concentration prediction value. .
[0088] In addition, the output head includes a concentration regression head for outputting the concentration C_pred and a reliability assessment head for outputting the reliability index RI. The concentration regression head includes single-mode regression heads for spectral, colorimetric and fluorescence channels to independently output the concentration prediction values for the corresponding channels.
[0089] Furthermore, fusion features , represented as: in, Aligned spectral, colorimetric, and fluorescence characteristics. This represents the fusion weights corresponding to spectral features, colorimetric features, and fluorescence features.
[0090] In addition, the FusionNet network is also used to calculate the reliability index (RI). Specifically, the RI is calculated based on the modal consistency index (δ'), model uncertainty (σ), and signal-to-noise ratio (SNR), and the RI is compared with a preset threshold to confirm the quantitative results, prompt for re-examination, or reject the output.
[0091] Specifically, the cross-modal fusion unit is also used to: utilize Monte Carlo random inactivation Dropout to perform a preset number of forward propagations on the fusion network used by the cross-modal fusion unit, obtain a corresponding number of prediction results, and estimate the model uncertainty based on multiple sets of prediction results; wherein any set of prediction results includes the concentration prediction values output by the single-modal regression head corresponding to the spectrum, colorimetry, and fluorescence, respectively; determine the weighted dispersion based on the prediction results to obtain the modal consistency index; and obtain the reliability index based on the model uncertainty, the modal consistency index, and the signal-to-noise ratio based on the prior fluorescence features.
[0092] It should be added that the formula for calculating the reliability index RI is as follows: in, Indicates the preset coefficient. Standardized fluorescence signal-to-noise ratio, Indicates model uncertainty. This represents the modal consistency index.
[0093] Furthermore, model uncertainty The Monte Carlo Dropout method can be used to estimate the Dropout activation in the evaluation head during inference, while maintaining the Dropout activation for the same ROI. After a random forward propagation, the prediction is obtained. , Defined as this The standard deviation of the second prediction ,in, express The mean of the predictions. This indicates the predicted concentration corresponding to time step t.
[0094] In addition, modal consistency index , represented as: in, This represents the weighted average. Indicates the weighted dispersion. This indicates the preset lower limit of quantification. To prevent the minimum value of division by zero, These are the concentration prediction values independently output by the single-modal regression heads of the spectral, colorimetric, and fluorescence channels, respectively.
[0095] Furthermore, the FusionNet network is trained using a multi-task loss function. , which is the weighted sum of concentration regression loss, consistency regularization term, and RI supervision loss, is expressed as: in, Indicates the concentration regression loss. This represents the consistency regularization loss. Indicates reliability monitoring loss, This represents the weighting coefficient for each type of loss.
[0096] It should be noted that the aforementioned networks can be trained by combining network prediction data and real samples with the corresponding loss functions to achieve high sensitivity and specificity. Building upon the above embodiments, a data generator is used to generate millions of precisely labeled synthetic data for pre-training, enabling the network to learn robust physical characteristics. Fine-tuning is then performed using a small number of rigorously labeled real samples. The loss function employs multi-task learning and consistency regularization to ensure that the model maintains the physical correlation between different modalities while predicting concentrations. This approach has good scalability; for example, the fusion network can be replaced with a more advanced cross-modal attention network; and the target analytes can be expanded from tumor markers to viral antigens, environmental pollutants, etc.
[0097] In one optional embodiment, the target analyte is selected from one or more of tumor markers (such as CEA, CA125, CYFRA21-1, GAGE7), viral antigens, environmental pollutants, or food safety markers; array localization can achieve perspective correction and site mapping by identifying alignment marks, edges, or feature points and estimating geometric transformations to support parallel detection and result traceability of multiple samples, multiple indicators, and control sites.
[0098] In addition, the network width, depth, activation function and regularization method of each of the above networks can be adjusted according to specific application scenarios and hardware conditions. As long as the functions of array positioning and geometric correction, spectral calibration and intensity correction, three-modal feature extraction and multimodal fusion and reliability assessment are achieved, no further limitations are made here.
[0099] In summary, the embodiments of the present invention use an intelligent control module as a unified acquisition and computing platform, integrates colorimetric, spectral, and fluorescence information on the same optical axis or equivalent optical path, realizes cross-modal mutual verification and result confidence quantification, reduces the risk of false positives / false negatives caused by a single modality, and forms a multi-point detection array on the reaction layer substrate through a microfluidic chip, making it possible to perform multiple samples / multi-indicators in parallel, significantly improving throughput per unit time and reducing the amount of data used per sample.
[0100] The following describes the multimodal high-throughput detection method based on mobile terminals provided by the present invention. The multimodal high-throughput detection method based on mobile terminals described below can be referred to in correspondence with the multimodal high-throughput detection system based on mobile terminals described above.
[0101] Figure 8 A flowchart illustrating a multimodal high-throughput detection method based on a mobile terminal is shown. Applied to any of the above-mentioned multimodal high-throughput detection systems based on a mobile terminal, the method includes: S81 controls the optical illumination module to provide a light source to the microfluidic detection chip and the spectral imaging module; S82 acquires colorimetric, fluorescence, and spectral images based on signals fed back from the microfluidic detection chip and the spectral imaging module. S83 extracts optical features from colorimetric, fluorescence, and spectral images and performs cross-modal fusion to predict concentration and reliability indicators; S84 compares the reliability index with the preset index threshold to determine the output concentration prediction result or trigger a re-inspection prompt.
[0102] In an optional embodiment, the method for fabricating the microfluidic detection chip used in the system of the present invention includes: providing a substrate layer; forming a gold nanoparticle array on the substrate layer; forming a flow channel layer and bonding the flow channel layer to the substrate layer on which the gold nanoparticle array is formed; inputting a preset solution into the flow channel of the flow channel layer to fix carboxyl groups on the surface of the gold nanoparticles and activate the fixed carboxyl groups; and inputting a preset specific capture molecule solution into the flow channel to form a capture antibody on the gold nanoparticles, thereby forming a microfluidic detection chip.
[0103] Specifically, forming the flow channel layer includes: providing and cleaning a quartz or glass plate; spin-coating photoresist onto the quartz or glass plate and drying it; exposing and developing it using a photomask and forming a microfluidic channel structure of the target depth through wet etching; stripping the photoresist by soaking in acetone; immersing and cleaning it in a mixed solution of hydrogen peroxide and concentrated sulfuric acid with a volume ratio of 1:3, rinsing it repeatedly with deionized water, drying it with nitrogen gas, and then treating it with oxygen plasma to obtain the flow channel layer.
[0104] It should be noted that the target depth can be determined based on the actual quartz or glass plate selected and the design requirements, such as 300 µm, etc., and no further limitation is made here.
[0105] In addition, a preset solution was introduced into the flow channel of the flow channel layer to fix carboxyl groups on the surface of gold nanoparticles. This included: using a syringe pump to push ethanol into the microfluidic channel and cleaning it by passing the solution through the microfluidic channel at a flow rate of 50 μL / min for 20 minutes; preparing a 10 mM 11-MUA ethanol solution and injecting it into the microfluidic channel at a flow rate of 0.83 μL / min using a syringe pump, and continuing the reaction for 12 to 14 hours to fix the 11-MUA with carboxyl groups at the end onto the surface of the gold nanoparticles (240) using gold-sulfur bonds; after the reaction, the ethanol cleaning step was performed again.
[0106] Activation of immobilized carboxyl groups involved preparing a 75 mM EDC ethanol solution and a 25 mM S-NHS PBS solution at a volume ratio of 3:1; passing the mixture through the channels at a flow rate of 25 μL / min for 30 minutes to activate the immobilized carboxyl groups in S830; and performing a rapid PBS wash after the reaction by manually passing 0.2 mL of PBS buffer into each channel.
[0107] In addition, a pre-defined specific capture molecule solution is introduced into the flow channel to allow capture antibodies to form on the gold nanoparticles. This includes: introducing the diluted specific capture molecule solution into the channel at a flow rate of 1.67 μL / min for 2 hours, allowing the capture antibodies to covalently bind to the chip surface via amide bonds; after the reaction, a rapid PBS wash is performed again.
[0108] In an optional embodiment, after introducing a pre-defined specific capture molecule solution into the flow path channel to allow capture antibodies to form on gold nanoparticles, the process includes: preparing a PBS solution of 1M ethanolamine and introducing it into the channel at a flow rate of 13.3 μL / min for 1 hour to block unreacted active sites when the carboxyl groups are activated and immobilized, thereby reducing non-specific adsorption; after the reaction, a rapid PBS wash is performed again.
[0109] In actual testing, the sample solution containing the target analyte is injected into the microfluidic channel, and then the chip is placed in a 37°C oven for 2 hours to allow the target antigen in the sample to specifically bind to the fixed capture antibody.
[0110] In an optional embodiment, to simultaneously acquire signals of three modalities at the same detection site (310), a multifunctional labeled secondary antibody is used. The secondary antibody is simultaneously labeled with a fluorescent group, specifically fluorescein isothiocyanate (FITC) and horseradish peroxidase (HRP). This multifunctional labeled secondary antibody solution is passed into a microfluidic channel at a flow rate of 1.67 μL / min for 1 hour to allow it to bind to the target analyte-capture molecule complex that has already specifically bound, forming a capture molecule-target analyte-labeled secondary antibody structure. After the reaction, rapid washing with PBS buffer is performed to remove unbound labeled secondary antibody.
[0111] Furthermore, the spectral signal originates from the localized surface plasmon resonance (LSPR) effect of the gold nanoparticle array (240). Specifically, when the aforementioned trap molecule-target analyte-labeled secondary antibody structure, including the trap molecule, target antigen, and secondary antibody labeled with FITC and HRP, accumulates on the surface of the gold nanoparticles at the detection site, it leads to a significant increase in the local refractive index of the AuNP surface. This refractive index change causes a redshift in the LSPR absorption peak, and the amount of redshift is positively correlated with the concentration of the target analyte bound at the site.
[0112] The fluorescence signal originates from the fluorescein isothiocyanate (FITC) fluorescent group coupled to the labeled secondary antibody. In subsequent data acquisition steps, when the detection site is illuminated by excitation light of a specific wavelength emitted from the fluorescence excitation source, the FITC group is excited and emits fluorescence through the emission-side filter assembly; the fluorescence intensity I captured by the camera is measured. fluo It is positively correlated with the number of bound labeled secondary antibodies, i.e., the concentration of the target analyte.
[0113] The colorimetric signal originates from the catalytic colorimetric reaction of horseradish peroxidase (HRP) coupled to the labeled secondary antibody. After secondary antibody binding and washing, a 3,3',5,5'-tetramethylbenzidine (TMB) chromogenic substrate solution is injected into the microfluidic channel. The HRP enzyme catalyzes the reaction of the TMB substrate, producing an insoluble blue precipitate in situ on the surface of the detection site. The reaction is terminated after 5 minutes and washed off. This blue precipitate absorbs light of a specific wavelength under white LED (410) illumination, causing changes in the color and optical density of the detection site. The absorbance A collected by the camera (110) is positively correlated with the amount of precipitate produced, i.e., the concentration of the target analyte.
[0114] It should be noted that the process involves immobilizing the captured molecules on the detection sites of the substrate layer and completing the sample introduction and reaction. Sample processing may include introducing the sample to be tested into the microfluidic channel via the sample application layer, followed by sealing, cleaning, and labeling steps. Colorimetric / spectral images and fluorescence images are acquired sequentially according to time-division multiplexing or equivalent sequences. White light illumination is used to form a global image containing the colorimetric region and spectral bands, while fluorescence excitation illumination is used to form a fluorescence emission image. The images are then array-localized and geometrically corrected, and pixel-wavelength mapping, intensity flat field, or dark background correction are performed. Feature quantities related to peak position or spectral shape are extracted from the spectral channel, feature quantities related to color or density are extracted from the colorimetric channel, and feature quantities related to emission intensity or signal-to-noise ratio are extracted from the fluorescence channel. Cross-modal fusion calculations are then performed to obtain the concentration or level information of the target being tested, and a reliability index is calculated. During multi-modal fusion, weights can be assigned based on channel quality or consistency. The results are judged based on the reliability index and a preset threshold. If the threshold is not reached, a prompt for retesting or repeated acquisition is issued. For details, please refer to the system embodiment described above, which will not be repeated here. The target analytes applicable to this method include, but are not limited to, tumor markers, viral antigens, environmental pollutants, or food safety markers.
[0115] In summary, the embodiments of the present invention use an intelligent control module as a unified acquisition and computing platform, integrates colorimetric, spectral, and fluorescence information on the same optical axis or equivalent optical path, realizes cross-modal mutual verification and result confidence quantification, reduces the risk of false positives / false negatives caused by a single modality, and forms a multi-point detection array on the reaction layer substrate through a microfluidic chip, making it possible to perform multiple samples / multi-indicators in parallel, significantly improving throughput per unit time and reducing the amount of data used per sample.
[0116] Figure 9 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 9 As shown, the electronic device may include a processor 910, a communication interface 920, a memory 930, and a communication bus 940. The processor 910, communication interface 920, and memory 930 communicate with each other via the communication bus 940. The processor 910 can call logical instructions in the memory 930 to execute a multimodal high-throughput detection method based on a mobile terminal. This method includes: controlling an optical illumination module to provide a light source to a microfluidic detection chip and a spectral imaging module; acquiring colorimetric images, fluorescence images, and spectral images based on signals fed back from the microfluidic detection chip and the spectral imaging module; extracting optical features from the colorimetric images, fluorescence images, and spectral images and performing cross-modal fusion to predict concentration and reliability indicators; comparing the reliability indicators with preset indicator thresholds to determine the output concentration prediction result or trigger a re-inspection prompt.
[0117] Furthermore, the logical instructions in the aforementioned memory 930 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0118] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the multimodal high-throughput detection method based on a mobile terminal provided by the above methods. The method includes: controlling an optical illumination module to provide a light source to a microfluidic detection chip and a spectral imaging module; acquiring colorimetric images, fluorescence images, and spectral images based on signals fed back from the microfluidic detection chip and the spectral imaging module; extracting optical features based on the colorimetric images, fluorescence images, and spectral images and performing cross-modal fusion to predict concentration and reliability indicators; comparing the reliability indicators with a preset indicator threshold to determine the output concentration prediction result or trigger a re-inspection prompt.
[0119] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the multimodal high-throughput detection method based on a mobile terminal provided by the methods described above. The method includes: controlling an optical illumination module to provide a light source to a microfluidic detection chip and a spectral imaging module; acquiring colorimetric images, fluorescence images, and spectral images based on signals fed back from the microfluidic detection chip and the spectral imaging module; extracting optical features based on the colorimetric images, fluorescence images, and spectral images and performing cross-modal fusion to predict concentration and reliability indicators; comparing the reliability indicators with preset indicator thresholds to determine the output concentration prediction result or trigger a re-inspection prompt.
[0120] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0121] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0122] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multimodal high-throughput detection system based on a mobile terminal, characterized in that, It includes a microfluidic detection chip, an optical illumination module, a spectral imaging module, and an intelligent control module, among which: The intelligent control module is used to control the optical illumination module to provide a light source to the microfluidic detection chip and the spectral imaging module; The microfluidic detection chip is used to carry the sample to be tested, and based on the light source provided by the intelligent control module, it performs biometric identification on the sample and induces observable spectral feature changes in the local surface plasmon resonance peak position to generate an optical signal. The spectral imaging module is used to receive the optical signal from the microfluidic detection chip and couple the optical signal into spectral stripes to the imaging surface; The intelligent control module is also used to acquire colorimetric images, fluorescence images and spectral images based on the signals fed back by the microfluidic detection chip and the spectral imaging module, extract optical features based on the colorimetric images, fluorescence images and spectral images and perform cross-modal fusion, predict concentration and reliability indicators, compare the reliability indicators with preset indicator thresholds, and determine the output concentration prediction result or trigger a re-inspection prompt.
2. The multimodal high-throughput detection system based on a mobile terminal according to claim 1, characterized in that, The microfluidic detection chip includes: basal layer; A flow channel layer is disposed on the substrate layer, and the flow channel layer is used to enclose the substrate layer to form a flow path channel. The flow path channel includes a waste liquid collection tank, a sample injection channel, and a buffer solution channel. The sample injection channel is used to transport the sample to be tested to the detection area, and the buffer solution channel is used to introduce buffer solution to rinse the detection area. The detection site array is disposed in the detection region on the substrate layer. The detection site array includes detection sites arranged in a preset array, and each detection site includes a gold nanoparticle disposed on the substrate layer and a functionalized layer disposed on the gold nanoparticle. The gold nanoparticle contains a trapping molecule for specifically binding to the target analyte in the sample to be tested.
3. The multimodal high-throughput detection system based on a mobile terminal according to claim 2, characterized in that, The intelligent control module includes: The data acquisition unit is used to identify alignment marks or array corner points of the microfluidic detection chip in the corresponding image based on the input colorimetric image, fluorescence image and spectral image, so as to perform array positioning, perspective correction and lens distortion correction on the corresponding image, and identify the region of interest corresponding to each detection site in the corrected image; The calibration unit is used to calibrate the pixel wavelength of the spectral bands of the corrected spectral image, and to perform intensity normalization processing on the one-dimensional spectral vector extracted from the spectral bands to obtain calibrated spectral data. It also performs photometric normalization processing on the corrected colorimetric image and fluorescence image to obtain calibrated colorimetric data and fluorescence data, respectively. The feature extraction unit is used to extract features from the calibrated spectral data, colorimetric data, and fluorescence data, respectively. The cross-modal fusion unit is used to perform weighted fusion of extracted spectral features, colorimetric features and fluorescence features using a preset gated fusion mechanism, and predict concentration prediction values and reliability indicators based on the fusion features; The reliability assessment unit is used to compare the reliability index with a preset index threshold. When the reliability index fails to reach the preset index threshold, a re-inspection prompt is triggered. When the reliability index reaches the preset index threshold, the concentration prediction result is output.
4. The multimodal high-throughput detection system based on a mobile terminal according to claim 3, characterized in that, The cross-modal fusion unit also includes: A modality-specific encoder is used to encode extracted spectral, colorimetric, and fluorescence features into aligned features; The gated subunit determines the fusion weights of the corresponding modes based on the aligned spectral, colorimetric, and fluorescence characteristics. The fusion layer performs weighted fusion of aligned spectral features, colorimetric features, and fluorescence features according to the fusion weights of each mode to obtain fused features; The output head predicts the concentration prediction value based on the fusion characteristics, and determines the reliability index based on the predicted concentration prediction value.
5. The multimodal high-throughput detection system based on a mobile terminal according to claim 4, characterized in that, The cross-modal fusion unit is also used for: Monte Carlo random inactivation Dropout is used to perform a preset number of forward propagations on the fusion network used by the cross-modal fusion unit to obtain a corresponding number of prediction results, and the model uncertainty is estimated based on multiple sets of prediction results; wherein, any set of prediction results includes the concentration prediction values output by the single-modal regression head corresponding to the spectrum, colorimetry and fluorescence, respectively. Based on the prediction results, the weighted dispersion is determined to obtain the modal consistency index; The reliability index is obtained based on the model uncertainty, the modal consistency index, and the signal-to-noise ratio based on the fluorescence features.
6. The multimodal high-throughput detection system based on a mobile terminal according to claim 3, characterized in that, The feature extraction unit is also used for: A spectral feature extraction network is used to extract the local surface plasmon resonance peak position features from the calibrated spectral data to obtain spectral features; A colorimetric feature extraction network is used to extract absorbance from calibrated colorimetric data to obtain colorimetric features; A fluorescence feature extraction network is used to extract fluorescence intensity and signal-to-noise ratio from calibrated fluorescence data to obtain fluorescence features.
7. The multimodal high-throughput detection system based on a mobile terminal according to claim 1, characterized in that, The optical illumination module includes a white light source, a fluorescent excitation source, an excitation-side filter assembly, an emission-side filter assembly, a first condenser lens, a second condenser lens, a diffuser, a reflector, and a dichroic mirror module, wherein: The white light source is used for colorimetric and spectral detection. The fluorescence excitation source is used for fluorescence detection; The excitation-side filter assembly is disposed on the fluorescence excitation light path and located between the second condenser lens and the fluorescence excitation light source. The second condenser lens is used to converge the illumination beam emitted by the fluorescence excitation light source. The emission-side filter assembly is disposed on the emission optical path; The first focusing lens is used to converge the illumination beam emitted by the white light source, and the diffuser is disposed between the white light source and the first focusing lens; The reflector is disposed in the white light illumination channel, and the dichroic mirror module is disposed between the emission-side filter component and the microfluidic detection chip, for separating and switching the excitation light and the emission light; The white light source has a color temperature range of 4000–6500 K and covers the visible spectrum. The fluorescent excitation light source has a center wavelength range of 450–550 nm and works with excitation-side and emission-side filter components to suppress stray light.
8. The multimodal high-throughput detection system based on a mobile terminal according to claim 1, characterized in that, The spectral imaging module includes an entrance slit, a diffraction grating, a focusing lens, and a spectral imaging region. The width of the entrance slit is 50 µm-150 µm, the groove density of the diffraction grating is greater than or equal to 1200 lines / mm, and the focusing lens includes a cemented doublet achromatic lens.
9. A multimodal high-throughput detection method based on a mobile terminal, characterized in that, The method, applied in any one of the mobile terminal-based multimodal high-throughput detection systems as described in claims 1-8, comprises: The optical illumination module controls the light source to provide light to the microfluidic detection chip and the spectral imaging module; Based on the signals fed back from the microfluidic detection chip and the spectral imaging module, colorimetric images, fluorescence images, and spectral images are acquired. Optical features are extracted from the colorimetric image, the fluorescence image, and the spectral image, and cross-modal fusion is performed to predict concentration and reliability indicators. The reliability index is compared with the preset index threshold to determine the output concentration prediction result or trigger a re-inspection prompt.