Method and system for predicting classification of pathological images based on multi-wavelength diffraction neural network

By using a three-channel, multi-wavelength diffraction neural network with separate inputs, combined with nano-silicon pillars and a two-dimensional tunable transparent film, the speed and energy consumption problems of traditional electronic computing models are solved, improving the speed and accuracy of pathological image classification and reducing equipment costs.

CN121616901BActive Publication Date: 2026-04-07NANCHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-02-02
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Traditional electronic computing models suffer from problems such as slow inference speed, high energy consumption, and channel crosstalk in pathological image classification. Furthermore, optical computing methods do not fully utilize the multi-channel information of visible light, making it difficult to meet high-precision requirements.

Method used

A three-channel, separate-input, multi-wavelength diffraction neural network is employed, which achieves optical parallel computing through nano-silicon pillars and a two-dimensional tunable transparent film. Combined with dynamic weight fusion technology, the computing speed and accuracy are improved.

Benefits of technology

It achieves efficient and low-energy-consumption pathological image classification, significantly improving inference speed and accuracy, and reducing equipment deployment costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121616901B_ABST
    Figure CN121616901B_ABST
Patent Text Reader

Abstract

The application discloses a pathological image classification method and system based on a multi-wavelength diffraction neural network, and relates to the technical field of optical and artificial intelligence cross. The method comprises the following steps: separating an RGB three-channel digital image into three-channel gray images; inputting the three-channel gray images into a spatial light modulator respectively to generate laser with corresponding wavelengths and carrying pixel-level phase information; inputting red channel laser, green channel laser and blue channel laser into independent corresponding multi-wavelength diffraction neural networks respectively for calculation; and finally outputting a classification result by optical fusion of three-channel information. The application avoids the channel crosstalk problem of the traditional coaxial beam combination mode, fully utilizes optical information, improves the classification task processing speed based on optical characteristics, reduces the inference delay, is suitable for intraoperative rapid diagnosis, high-throughput pathological screening, automatic driving high-speed inference, production rapid sorting and the like, and has the advantages of strong parallelism, high precision, good interpretability, zero electronic power consumption inference and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of interdisciplinary technology of optics and artificial intelligence, and in particular to a classification method and system for predicting pathological images based on multi-wavelength diffraction neural networks. Background Technology

[0002] With the widespread application of artificial intelligence and combinatorial optimization in industries such as manufacturing, transportation, and healthcare, the consumption of computing resources and energy is growing exponentially. Deep learning models based on electronic computing architectures, such as Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs), face the "Von Neumann bottleneck," where frequent data transfer between computing and storage leads to high energy consumption and latency, thus limiting computing speed. The serial computing mode of electronic chips results in high inference latency. Simultaneously, this design continuously increases the energy consumption of computing; large-scale parameter calculations require a continuous power supply, which is unfavorable for deployment in portable devices. Currently, artificial intelligence relies heavily on proprietary equipment; complex models require high-performance GPUs or CPUs, resulting in high costs and insufficient flexibility.

[0003] Optical computing, with its high parallelism, high-speed light propagation, and low power consumption, offers a new direction for solving these problems. Diffraction neural networks, as the core architecture of optical computing, achieve weight calculation and feature extraction through light field modulation. However, current optical computing methods mostly focus on single-channel tasks, failing to fully utilize the multi-channel information of visible light. Furthermore, their lens structure design is not optimized, resulting in insufficient phase modulation accuracy and limited feature representation capabilities, making it difficult to meet the high-precision requirements of pathological image classification. Summary of the Invention

[0004] This invention aims to provide a classification method and system for predicting pathological images based on a multi-wavelength diffraction neural network. It solves the problems of slow inference speed, high energy consumption, and channel crosstalk in traditional electronic computing models by using three-channel separate input, multi-wavelength modulation, and dynamic weight fusion technology. At the same time, it fills the gap in optical computing that cannot perform multi-channel calculations, further improving the speed and accuracy of image classification tasks in optical computing, and reducing equipment deployment costs and energy consumption.

[0005] In a first aspect, the present invention provides a classification method for predicting pathological images based on a multi-wavelength diffraction neural network, comprising the following steps:

[0006] Step S1: Obtain an RGB three-channel digital image, preprocess the RGB three-channel digital image to obtain the preprocessed red channel grayscale image, green channel grayscale image and blue channel grayscale image;

[0007] Step S2: Input the preprocessed red channel grayscale image, green channel grayscale image and blue channel grayscale image into the spatial light modulator to generate red channel laser, green channel laser and blue channel laser respectively;

[0008] Step S3: The red channel laser, green channel laser, and blue channel laser are respectively incident on the corresponding multi-wavelength diffraction neural network for parallel optical calculation to obtain the output light field;

[0009] The multi-wavelength diffraction neural network is composed of a nano-silicon pillar, a lens, and a two-dimensional tunable transparent film. The nano-silicon pillar is used to achieve phase modulation of the incident light wave, and the two-dimensional tunable transparent film is used to achieve optical nonlinear activation.

[0010] Step S4: At the exit surface of the multi-wavelength diffraction neural network, the output light field is received by the image sensor and converted into the corresponding electrical signal.

[0011] Step S5: Convert the spatial feature maps represented by the electrical signals of the red, green, and blue channels into one-dimensional feature vectors respectively; use a dynamic weight balancing strategy to perform weighted fusion of the three one-dimensional feature vectors of the red, green, and blue channels, and finally output the classification result.

[0012] Secondly, the present invention provides a classification system for predicting pathological images based on a multi-wavelength diffraction neural network, the system comprising:

[0013] Image input and preprocessing module: used to collect RGB three-channel digital images and perform specification adjustment, channel separation and grayscale correction;

[0014] Optical signal modulation output module: including a spatial light modulator, used to convert a three-channel grayscale image into a phase-modulated laser of the corresponding wavelength;

[0015] Multiwavelength diffraction neural network module: contains three independent two-level multiwavelength diffraction neural networks for parallel computation of incident laser light;

[0016] Optical signal receiving and electrical signal conversion module: including an image sensor, used to convert optical field signals into electrical signals;

[0017] Data processing terminal module: used for feature vector processing, dynamic weighted fusion and classification result output of electrical signals.

[0018] The beneficial effects of this invention are:

[0019] This invention organically combines a nano-silicon pillar, a lens, and a two-dimensional tunable transparent film. The nano-silicon pillar is etched onto the front of the lens, while the two-dimensional tunable transparent film is adsorbed onto the back of the lens. The nano-silicon pillar achieves continuous phase modulation from 0 to 2π through diameter adjustment, the two-dimensional tunable transparent film provides nonlinear activation, and combined with the lens's high transmittance and precise positioning, ensures the accuracy and efficiency of optical calculations.

[0020] The optical part of this invention achieves high-speed parallel feature extraction, while the electrical part fuses three-channel features through a dynamic weight balancing strategy, balancing the high efficiency of optical computing with the flexibility of electrical fusion, and achieving high speed, high precision, and low power consumption. Attached Figure Description

[0021] Figure 1 This is a flowchart of the present invention.

[0022] Figure 2 This is a schematic diagram of the structure of the multi-wavelength diffraction neural network in this embodiment. Detailed Implementation

[0023] To better understand the above-mentioned objects, features, and advantages of the present invention, the present invention will be further described in detail below with reference to embodiments. The specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0024] like Figure 1 As shown, this invention provides a classification method for predicting pathological images based on a multi-wavelength diffraction neural network, comprising the following steps:

[0025] Step S1: Obtain an RGB three-channel digital image, preprocess the RGB three-channel digital image to obtain the preprocessed red channel grayscale image, green channel grayscale image and blue channel grayscale image;

[0026] Step S2: Input the preprocessed red channel grayscale image, green channel grayscale image and blue channel grayscale image into the spatial light modulator to generate red channel laser, green channel laser and blue channel laser respectively;

[0027] Step S3: The red channel laser, green channel laser, and blue channel laser are respectively incident on the corresponding multi-wavelength diffraction neural network for parallel optical calculation to obtain the output light field;

[0028] The multi-wavelength diffraction neural network is composed of a nano-silicon pillar, a lens, and a two-dimensional tunable transparent film. The nano-silicon pillar is used to achieve phase modulation of the incident light wave, and the two-dimensional tunable transparent film is used to achieve optical nonlinear activation.

[0029] Step S4: At the exit surface of the multi-wavelength diffraction neural network, the output light field is received by the image sensor and converted into the corresponding electrical signal.

[0030] Step S5: Convert the spatial feature maps represented by the electrical signals of the red, green, and blue channels into one-dimensional feature vectors respectively; use a dynamic weight balancing strategy to perform weighted fusion of the three one-dimensional feature vectors of the red, green, and blue channels, and finally output the classification result.

[0031] Furthermore, the nano-silicon pillars are arranged at equal intervals around the center on the lens. Each nano-silicon pillar receives the laser output of one pixel in the spatial light modulator. By adjusting the diameter of the nano-silicon pillars, the incident light wave achieves a continuous transmission phase shift of 0-2π, as shown in the following formula:

[0032] ;

[0033] in, Let i be the transmission phase of the nano-silicon pillar with x-axis i and y-axis j. The incident light wavelength, Because of the difference in refractive index between silicon and air, For the fixed height of the nano-silicon pillars, The mapping function between the diameter and phase of the nano-silicon pillar;

[0034] Nano-silicon pillars on incident light field The modulation follows the complex amplitude transmission rule, yielding the output light field of the nano-silicon pillar, as shown in the following formula:

[0035] ;

[0036] in, The output light field of the nano-silicon pillar. denoted as the complex transmission coefficient of the nano-silicon pillar. , The amplitude transmittance of the nano-silicon pillar. It is the imaginary unit.

[0037] Furthermore, based on its saturable absorption characteristics, the output light field of the nano-silicon pillars is nonlinearly modulated to obtain the light intensity of the two-dimensional tunable transparent film, as shown in the following formula:

[0038] ;

[0039] in, The light intensity of a two-dimensional tunable light-transmitting film;

[0040] The transmittance of a two-dimensional tunable light-transmitting film varies with light intensity, as shown in the following formula:

[0041] ;

[0042] in, For the transmittance of two-dimensional tunable light-transmitting films, Unsaturated transmittance For saturation transmittance, The saturation light intensity;

[0043] The output light field of the laser after passing through a nano-silicon pillar-lens-two-dimensional tunable transparent film is expressed by the following formula:

[0044] ;

[0045] in, The output light field of the laser after passing through a nano-silicon pillar-lens-two-dimensional tunable transparent film.

[0046] In this embodiment, the publicly available Camelyon16 dataset is used as both the training and validation sets. The Camelyon16 dataset consists of 171 whole-slide images (WSI) of breast cancer without lymph node metastasis and 219 whole-slide images of breast cancer with lymph node metastasis, along with corresponding annotations of the metastatic lesions by professional pathologists. This invention divides the whole-slide images into equidistant 256-pixel × 256-pixel × 3-channel color blocks, and performs quality control on the segmented color blocks to obtain a color block set. Based on the tumor region annotations in the Camelyon16 dataset, the color blocks are divided into positive-label color blocks and negative-label color blocks, resulting in a total of 5,943,990 color blocks, of which 816,571 are positive-label color blocks and 5,127,419 are negative-label color blocks. This invention randomly selects 100,000 images from each of the positive and negative-label color blocks to form the training set, and then randomly selects 50,000 images to form the validation set.

[0047] The color patches in the training and validation sets are preprocessed as follows: all color patches are uniformly scaled to 256 pixels × 256 pixels × 3 channels, the scaled color patches are separated into three-channel grayscale images, namely red channel grayscale image, green channel grayscale image and blue channel grayscale image, the grayscale values ​​of the three channels are linearly normalized to the [0,1] interval, and gamma correction is performed on the linearly normalized grayscale values, with a gamma correction coefficient of γ=0.8.

[0048] (I) Training Phase: Optimization of Digital Simulation Model Parameters

[0049] A digital simulation model of a multi-wavelength diffraction neural network is constructed in a computer (such as one equipped with an Intel Core i7 processor and 16GB of memory). Using the aforementioned training set, it is trained using a deep learning framework (such as PyTorch). During training, the loss between the probability distribution predicted by the digital simulation model of the multi-wavelength diffraction neural network and the actual labels is calculated. In the digital simulation model, the mapping relationship between the diameter of the silicon nanopillar and the optical phase modulation it generates is defined by a mapping function. Using a gradient backpropagation algorithm, the optimal network parameters are determined by iteratively minimizing the loss function by adjusting the diameter of the silicon nanopillar and the thickness of the titanium tricarbide film corresponding to each pixel in the digital simulation model of the multi-wavelength diffraction neural network. The mapping function, i.e., the mapping function between the diameter of the silicon nanopillar and the phase modulation, is defined as follows: , It is an exponential function. The attenuation coefficient is determined by electromagnetic simulation fitting using the finite element method. Let i be the diameter of the nano-silicon column with x-axis and y-axis.

[0050] After training, the optimized parameters are mapped to the structural dimensions of the physical device. Nanoimprint lithography with wet etching is used to fabricate nano-silicon pillars for each pixel according to the mapped structural dimensions. A titanium tricarbide thin film is then adsorbed onto the back of the silicon dioxide lens using atomic layer deposition (ALD) assisted growth technology.

[0051] (ii) Inference stage: Classification based on a trained multi-wavelength diffraction neural network.

[0052] The preprocessed three-channel grayscale image is converted into an electrical signal and input to a spatial light modulator (SLM). The SLM uses a liquid crystal display with a resolution of 1920×1080, a pixel size of 8 micrometers, and a response time of less than 5 milliseconds. The SLM linearly maps grayscale values ​​to driving voltages using a pixel lookup table. The driving voltages adjust the tilt angle of the liquid crystal molecules, converting the grayscale values ​​into pixel-level phase. The SLM outputs red, green, and blue channel lasers. The wavelength of the red channel laser is... =700 nanometers, wavelength of green channel laser =546.1 nanometers, the wavelength of the blue channel laser =435.8 nm, emits a continuous laser with a power of 10 mW, has a beam quality M² < 1.2, the wavelength of the laser at each pixel remains constant, and its phase is dynamically adjusted according to the gray value at the corresponding position in the input grayscale image.

[0053] The output light field is obtained by separately incidenting red, green, and blue channel lasers onto their respective corresponding multi-wavelength diffraction neural networks. The multi-wavelength diffraction neural network consists of nano-silicon pillars, silicon dioxide lenses, and titanium tricarbide thin films. The structure of the multi-wavelength diffraction neural network is as follows: Figure 2 As shown.

[0054] At the exit surface of the multi-wavelength diffraction neural network, the output light field is received by an image sensor and converted into a corresponding electrical signal. The image sensor is a complementary metal-oxide-semiconductor (CMOS) sensor with a resolution of 12 bits, a frame rate of 30 frames per second, and a pixel size of 5.5 micrometers. The electrical signals of the red, green, and blue channels are represented as spatial feature maps, which are 256-pixel × 256-pixel single-channel light intensity distributions. Flattening the spatial feature map according to pixel coordinates yields a 65536-dimensional one-dimensional feature vector. The mapping relationship between the index p of the one-dimensional feature vector and the pixel coordinates is as follows: Where x is the x-coordinate of the pixel coordinate and y is the y-coordinate of the pixel coordinate. The one-dimensional vector is normalized using the L2 norm to obtain a normalized vector, which is then input into a fully connected layer to reduce the dimension to 2048. The fully connected layer performs a linear transformation through the weight matrix and bias term, and introduces non-linearity through the ReLU activation function. During the training phase, the reduced feature vector is randomly deactivated. A dynamic weight balancing strategy is adopted, where the 2048-dimensional one-dimensional feature vectors of the red, green, and blue channels are weighted and summed with their corresponding dynamic weight coefficients to obtain a fused feature vector. The dynamic weight coefficients are obtained through a normalized exponential function based on the channel importance score, and the sum of the weight coefficients of all channels is 1. The channel importance score is a learnable parameter, initialized to 1, and obtained through backpropagation optimization. The fused feature is input into a classifier to obtain the original scores for each category. Finally, the original scores of each category are normalized to obtain the probability distribution of each category, and the category with the highest probability is output as the final classification result.

[0055] This invention was compared with traditional models on the Camelyon16 dataset for a tumor cell identification task. The comparison metrics included ROC value, single-image inference speed, and total energy consumption during the inference phase. The comparison results are shown in Table 1. This invention achieves a single-image inference speed of up to 50 microseconds, significantly improved compared to the 2000 microseconds of the traditional ResNET-50 model and the 3200 microseconds of ViT-B / 16. During the inference phase, the total energy consumption of this invention is only 40 watts, far lower than the 380 watts of ResNET-50 and the 420 watts of ViT-B / 16, highlighting its low energy consumption characteristics. The ROC value of this invention reaches 95.9%, comparable to ResNET-50's 95.2%, and although slightly lower than ViT-B / 16's 97.0%, it still maintains extremely high classification accuracy. In summary, this invention achieves significant improvements in inference speed and energy consumption while maintaining high classification accuracy.

[0056] Table 1 Performance comparison between the present invention and the traditional model

[0057]

[0058] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the implementation of the present invention. Those skilled in the art can make other variations or modifications based on the above description. It is neither necessary nor possible to exhaustively describe all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the claims of the present invention.

Claims

1. A classification method for predicting pathological images based on multi-wavelength diffraction neural networks, characterized in that, Includes the following steps: Step S1: Obtain an RGB three-channel digital image, preprocess the RGB three-channel digital image to obtain the preprocessed red channel grayscale image, green channel grayscale image and blue channel grayscale image; The preprocessing includes: scaling the RGB three-channel digital image to a preset size; separating the scaled image into a red channel grayscale image, a green channel grayscale image, and a blue channel grayscale image; linearly normalizing the grayscale values ​​of the red channel grayscale image, the green channel grayscale image, and the blue channel grayscale image to a preset range; and performing gamma correction. Step S2: Input the preprocessed red channel grayscale image, green channel grayscale image and blue channel grayscale image into the spatial light modulator to generate red channel laser, green channel laser and blue channel laser respectively; Step S3: The red channel laser, green channel laser, and blue channel laser are respectively incident on the corresponding multi-wavelength diffraction neural network for parallel optical calculation to obtain the output light field; The multi-wavelength diffraction neural network is composed of a nano-silicon pillar, a lens, and a two-dimensional tunable transparent film. The nano-silicon pillar is used to achieve phase modulation of the incident light wave, and the two-dimensional tunable transparent film is used to achieve optical nonlinear activation. The nano-silicon pillars are arranged at equal intervals around the center of the lens. Each nano-silicon pillar receives the laser output of one pixel in the spatial light modulator. By adjusting the diameter of the nano-silicon pillars, the incident light wave achieves a continuous transmission phase shift of 0-2π, as shown in the following formula: ; in, Let i be the transmission phase of the nano-silicon pillar with x-axis i and y-axis j. The incident light wavelength, Because of the difference in refractive index between silicon and air, For the fixed height of the nano-silicon pillars, The mapping function between the diameter and phase of the nano-silicon pillar; Nano-silicon pillars on incident light field The modulation follows the complex amplitude transmission rule, yielding the output light field of the nano-silicon pillar, as shown in the following formula: ; in, The output light field of the nano-silicon pillar. denoted as the complex transmission coefficient of the nano-silicon pillar. , The amplitude transmittance of the nano-silicon pillar. The imaginary unit; The two-dimensional tunable transparent film, based on its saturable absorption characteristics, nonlinearly modulates the output light field of the nano-silicon pillars to obtain the light intensity of the two-dimensional tunable transparent film, as shown in the following formula: ; in, The light intensity of a two-dimensional tunable light-transmitting film; The transmittance of a two-dimensional tunable light-transmitting film varies with light intensity, as shown in the following formula: ; in, For the transmittance of two-dimensional tunable light-transmitting films, Unsaturated transmittance For saturation transmittance, The saturation light intensity; The output light field of the laser after passing through a nano-silicon pillar-lens-two-dimensional tunable transparent film is expressed by the following formula: ; in, The output light field of the laser after passing through a nano-silicon pillar-lens-two-dimensional tunable transparent film; Step S4: At the exit surface of the multi-wavelength diffraction neural network, the output light field is received by the image sensor and converted into the corresponding electrical signal. Step S5: Convert the spatial feature maps represented by the electrical signals of the red, green, and blue channels into one-dimensional feature vectors respectively; use a dynamic weight balancing strategy to perform weighted fusion of the three one-dimensional feature vectors of the red, green, and blue channels, and finally output the classification result; the specific steps are as follows: A dynamic weight coefficient is calculated for the feature vectors of the red, green, and blue channels respectively; the dynamic weight coefficient is obtained by a normalized exponential function of the learnable channel importance score, and the sum of the dynamic weight coefficients of all channels is 1; The feature vectors of the red, green, and blue channels are weighted and summed with their corresponding dynamic weight coefficients to obtain the fused features. The fused features are input into a classifier to obtain the raw scores for each category; The original scores of each category are normalized to obtain the probability distribution of each category, and the category with the highest probability is output as the final classification result.

2. The classification method for predicting pathological images based on a multi-wavelength diffraction neural network according to claim 1, characterized in that, In step S2, the spatial light modulator modulates the grayscale information to generate red channel laser, green channel laser and blue channel laser with corresponding wavelengths and carrying pixel-level phase information; wherein the wavelength of the red channel laser is 700nm, the wavelength of the green channel laser is 546.1nm and the wavelength of the blue channel laser is 435.8nm.

3. A classification system for predicting pathological images based on a multi-wavelength diffraction neural network, used to implement the classification method for predicting pathological images based on a multi-wavelength diffraction neural network as described in any one of claims 1 to 2, characterized in that, Classification systems include: Image input and preprocessing module: used to collect RGB three-channel digital images and perform specification adjustment, channel separation and grayscale correction; Optical signal modulation output module: including a spatial light modulator, used to convert a three-channel grayscale image into a phase-modulated laser of the corresponding wavelength; Multiwavelength diffraction neural network module: contains three independent two-level multiwavelength diffraction neural networks for parallel computation of incident laser light; Optical signal receiving and electrical signal conversion module: including an image sensor, used to convert optical field signals into electrical signals; Data processing terminal module: used for feature vector processing, dynamic weighted fusion and classification result output of electrical signals.

4. The classification system for predicting pathological images based on a multi-wavelength diffraction neural network according to claim 3, characterized in that, The spatial light modulator in the optical signal modulation output module linearly maps grayscale values ​​to driving voltages. The driving voltages regulate the tilt angle of liquid crystal molecules, thereby converting grayscale values ​​into pixel-level phase.

5. The classification system for predicting pathological images based on a multi-wavelength diffraction neural network according to claim 3, characterized in that, The image sensor is a complementary metal-oxide-semiconductor sensor or a charge-coupled device sensor.

Citation Information

Patent Citations

  • Scattering medium-penetrating perception and display integrated augmented reality device and method

    CN118409435A

  • Online detection method and system for laser-induced damage of optical lens

    CN120801365A