Eye white image acquisition and disease auxiliary diagnosis method and device
By automatically positioning the eye orbit and adaptively collecting the white area map, combining the target detection and medical segmentation model, an abnormal white heat map is generated, which solves the problems of low eye diagnosis efficiency and subjective omission in traditional Chinese medicine, and achieves efficient auxiliary diagnosis of diseases.
Patent Information
- Application Number
- CN202510546874.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-07-04
AI Technical Summary
Traditional Chinese medicine's eye-catching diagnosis relies on manual observation, is inefficient and prone to subjective omissions, and lacks intuitive auxiliary means of assessment of diseases.
Automatic positioning of the eye sockets is adopted, and the five-directional white area maps are adaptively collected. The eye sockets are accurately positioned through object detection and PID control algorithms, and the abnormal parts are automatically marked with the medical segmentation model to calculate and display the abnormal heat map of the eye sockets.
It realizes batch acquisition of patient eye white images, improves diagnostic efficiency, provides intuitive auxiliary means for the evaluation of diseases, and reduces manual intervention.
Smart Images

Figure CN120240986A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical diagnosis, and particularly to a method and device for collecting sclera images and assisting in disease diagnosis. Background Art
[0002] In traditional Chinese medicine theory, the eyes are closely related to the five internal organs and six hollow organs, and the health status of the internal organs can be diagnosed by observing the eyes. Clinically, observing the eyes mainly involves traditional Chinese medicine practitioners observing the vein morphology, color distribution, and abnormal plaque positions in the sclera area of patients to judge visceral diseases. There are two problems with traditional eye observation diagnosis: First, the collection of sclera images relies on the naked-eye observation of traditional Chinese medicine practitioners, and it is necessary to repeatedly observe the spirit, color, shape, and state of the sclera area, which is rather cumbersome and inefficient; Second, there may be subjective omissions in the identification of abnormalities by traditional Chinese medicine, and there is a lack of auxiliary means for intuitive evaluation of disease observation. Therefore, a method and device for collecting sclera images and assisting in disease diagnosis are designed to achieve full-automatic and full-region collection of sclera images, and assist doctors to quickly and conveniently complete disease observation and evaluation to meet the auxiliary technical support for batch and centralized traditional Chinese medicine eye observation diagnosis. Summary of the Invention
[0003] The present invention provides a method and device for collecting sclera images and assisting in disease diagnosis, which automatically locates the eye sockets, adaptively collects sclera area images in five directions, and realizes hierarchical display of sclera abnormal heat maps to meet the batch collection of patients' sclera images, improve the diagnosis efficiency, and assist doctors in completing disease diagnosis.
[0004] To achieve the above object, the present invention is implemented by adopting the following technical solutions:
[0005] A method for collecting sclera images and assisting in disease diagnosis includes the following steps:
[0006] S1. Obtain the video frames collected by the camera in real time, and obtain the information of the eye socket detection frame in the video frames through the target detection algorithm;
[0007] S2. Based on the information of the eye socket detection frame, adjust the displacements of the X, Y, and Z axes of the motor through the PID control algorithm;
[0008] S3. Calculate the input error values of the X, Y, and Z axes in the PID control algorithm;
[0009] S4. Repeat steps S1 - S3 until the input error value is less than the preset threshold, indicating that the eye socket has been located, and stop the motor drive;
[0010] S5. Obtain the video frames collected by the camera in real time, and obtain the information of the pupil detection frame in the video frames through the target detection algorithm;
[0011] S6. Based on the pupil detection box information, command the subject to look up, left, middle, right, and down in five directions through voice;
[0012] S7. When it is detected that the pupil is facing the specified five directions, save the current video frame as the sclera area map corresponding to the direction;
[0013] S8. Repeat steps S5 - S7 until the image acquisition of the sclera area maps in five directions is completed;
[0014] S9. Extract features from the acquired images, and automatically label abnormal parts based on the offline trained medical segmentation model;
[0015] S10. According to the labeled images, establish a dynamic coordinate system and a mapping area with the pupil center as the pole, and calculate the polar coordinate parameters, the polar radius r and the polar angle θ:
[0016]
[0017] Among them, (x0, y0) is the pupil center coordinate, (x i , y i ) is the central point coordinate of the abnormal area in the marked two - dimensional coordinate system, and i is each abnormal serial number; the mapping area is the standard eye shape area of a person, which is used as a position reference for mapping the abnormal part under the standard eye shape. The mapping area is divided into eight regions A1 - A8 by a straight line according to the pupil position; d max and d min are the differences between the maximum and minimum effective radii of the mapping area, representing the range size of the mapping area, S max is the maximum visible radius of the sclera, is the Euclidean distance from the pupil center coordinate to the central point coordinate of the abnormal area, is the scaling coefficient between the mapping area and the sclera image area, and the coordinate points in the polar coordinate system are calculated in the form of (r, θ);
[0018] S11. According to the coordinate system information, calculate the spatial weight value w j of the mapping abnormality through the distance attenuation function:
[0019]
[0020] Among them, p i is the central coordinate of the abnormal position in the polar coordinate system, represented by (r i , θ i ), c j is the central coordinate of the mapping area numbers A1 - A8, where j is the mapping area number; σ is the standard radius of the mapping area, n is the number of labeled abnormal parts, and the final result w j is the spatial weight value of the j - th area of the mapping area number;
[0021] S12. Calculate the comprehensive gain value G of the mapping area j ;
[0022]
[0023] where α is the spatial weight coefficient, β is the abnormal point density weight coefficient, represents the number of the j-th abnormality in the total abnormalities, indicating the abnormal point density, γ is the clinical prior knowledge weight coefficient, and C j is the number of historically diagnosed cases in this mapping area;
[0024] S13. Normalize the comprehensive gain value G j to generate a sclera abnormality heat map and display it in grades.
[0025] Furthermore, the normalization process of the gain value is as follows:
[0026]
[0027] where G j is the comprehensive gain value, G min is the minimum value of the comprehensive gain degrees of all abnormalities, G max is the maximum value of the comprehensive gain degrees of all abnormalities, and finally the heat value G' j is obtained. This operation maps the value to the interval [0, 1].
[0028] Furthermore, it also includes generating a sclera abnormality graded heat map of the original sclera image. The original sclera image is calculated based on a two-dimensional coordinate system, and Gaussian kernel density estimation is introduced to perform spatial smoothing on the discrete gain values:
[0029]
[0030] where x is the x coordinate in the heat map, y is the y coordinate in the heat map, x j is the x coordinate of the center of the abnormal area, and y j is the y coordinate of the center of the abnormal area; ρ is a parameter controlling the bandwidth of the kernel function, used to adjust the smoothness of the color switching in the heat map, H(x, y) is the gain value of this pixel point, and through normalization processing and color mapping, a sclera abnormality graded heat map of the original sclera image is obtained.
[0031] An apparatus for sclera image acquisition and disease-assisted diagnosis includes:
[0032] A camera module: used for visual positioning and acquiring high-resolution sclera images;
[0033] A mechanical module: a mechanical framework supporting all other modules;
[0034] Motion module: Composed of 4 motors, applied to three-axis motion control, with 1 motor for the X-axis, 2 motors for the Y-axis, and 1 motor for the Z-axis;
[0035] Central module: The central brain of the entire device, used to drive the motors and process information;
[0036] Voice module: Used to broadcast the current information collection status;
[0037] Graphic display module: A software module built into the computer, used for the visual display of functions such as inputting information, displaying image information, and annotating information;
[0038] Computer module: Used to automatically annotate information and display all information;
[0039] The camera module, motion module, and voice module are all connected to the central module. The central module is connected to the computer module, and the graphic display module is built into the computer module.
[0040] Further, the mechanical frame includes a U-shaped bracket, an X-axis translation structure, a connecting rod, a Y1-axis translation structure, a Y2-axis translation structure, and a Z-axis lifting structure. The Y1-axis translation structure and the Y2-axis translation structure are symmetrically arranged. One end of the Y1-axis translation structure and the Y2-axis translation structure is connected by a connecting rod. The X-axis translation structure is vertically arranged and connected to the Y1-axis translation structure and the Y2-axis translation structure. The Z-axis lifting structure is arranged on the X-axis translation structure and is perpendicular to the X-axis translation structure, the Y1-axis translation structure, and the Y2-axis translation structure. The camera module is arranged on the Z-axis lifting structure. The U-shaped bracket is fixedly vertically arranged on the Y1-axis translation structure and the Y2-axis translation structure, and the height of the U-shaped bracket matches the position of the lens of the camera module. One motor drives the X-axis translation structure to move along the X-axis, two motors respectively drive the Y1-axis translation structure and the Y2-axis translation structure to move synchronously along the Y-axis, and another motor drives the Z-axis translation structure. The voice module and the control circuit module are arranged on the connecting rod.
[0041] Further, the X-axis translation structure, the connecting rod, the Y1-axis translation structure, the Y2-axis translation structure, and the Z-axis lifting structure all include a coupling, a smooth rod, a lead screw, and a motor bracket. A lead screw is arranged on the smooth rod. The lead screw is connected to the motor output shaft through a coupling, and the motor is arranged on the motor bracket.
[0042] Further, the current information collection status includes the collection quantity, whether the collection is successful, and the collected parts.
[0043] Compared with the prior art, the beneficial effects of the present invention are:
[0044] 1) A method and device for collecting sclera images and assisting in disease diagnosis according to the present invention can fully automatically locate the patient's eye sockets and adaptively collect sclera area images in five directions of the patient, and use the graded sclera abnormality heat map as an auxiliary means for diagnosis, realizing batch collection of the patient's sclera images and improving the diagnosis efficiency;
[0045] 2) Through the collaborative mechanism of the target detection algorithm and the PID control algorithm, the present invention realizes the automatic and precise positioning of the human eye socket, eliminating the need for doctors to manually adjust the camera and improving the collection efficiency;
[0046] 3) Aiming at the problems that the whole sclera is less exposed and blocked in normal people when opening their eyes, the present invention has established a collection mechanism for sclera area images in five directions, namely the upper sclera, the left sclera, the middle sclera, the right sclera, and the lower sclera, and realizes the adaptive collection of sclera area images in five directions through the target detection algorithm for the pupil, so as to meet the batch collection of the patient's sclera images;
[0047] 4) By combining the automatically marked abnormal part information, the present invention calculates the comprehensive gain value of the abnormality in the mapping area, generates a graded sclera abnormality heat map, highlights the high-incidence areas of abnormalities, and uses it as an auxiliary means for diagnosis to improve the diagnosis efficiency. Description of the Drawings
[0048] Figure 1 It is a schematic diagram of the module structure of the sclera image collection and disease diagnosis assistance device according to the embodiment of the present invention.
[0049] Figure 2 It is an electrical connection diagram of the sclera image collection and disease diagnosis assistance according to the embodiment of the present invention.
[0050] Figure 3 It is a mechanical connection diagram of the mechanical module and the motion module according to the embodiment of the present invention.
[0051] Figure 4 It is a schematic diagram of the mechanical structure of the image collection device according to the embodiment of the present invention.
[0052] Figure 5 It is a side view of the mechanical mechanism of the image collection device according to the embodiment of the present invention.
[0053] Figure 6 It is a regional division diagram of the standard eye shape diagram according to the embodiment of the present invention.
[0054] Figure 7 It is a flowchart of the method according to the embodiment of the present invention.
[0055] Figure 8 It is a flowchart of obtaining the orbital detection box information in the real-time video frame based on the target detection algorithm according to the embodiment of the present invention.
[0056] Figure 9 It is the regional distribution diagram of the five regions for sclera acquisition in the embodiments of the present invention.
[0057] Figure 10 It is the example diagram of the regional distribution of the five regions for sclera acquisition in the embodiments of the present invention.
[0058] In the figure: 1. U-shaped bracket; 2. Voice module; 3. Central module; 4. X-axis translation structure; 5. X-axis stepper motor; 6. Y1-axis stepper motor; 7. Y1-axis translation structure; 8. Y2-axis stepper motor; 9. Y2-axis translation structure; 10. Z-axis stepper motor; 11. Z-axis lifting structure; 12. Camera. Detailed implementation manners
[0059] The following further explains the detailed implementation manners of the present invention with reference to the accompanying drawings:
[0060] See Figure 1-2 , which is the overall structure diagram and electrical connection relationship diagram of the present invention. A method and device for sclera image acquisition and disease auxiliary diagnosis of the present invention, the image acquisition device includes:
[0061] Camera module: Composed of the Huiboshi 13-megapixel autofocus camera 12, used for visual positioning and acquiring high-resolution sclera images;
[0062] Mechanical module: The mechanical framework that supports all other modules;
[0063] Motion module: Composed of the X-axis stepper motor 5, Y1-axis stepper motor 6, Y2-axis stepper motor 8, and Z-axis stepper motor 10, controlling the movement of the camera 12 in the X, Y, and Z axes. The forward and reverse rotations of the axis motors respectively correspond to the forward and reverse directions of the axis; the X-axis stepper motor 5 controls the movement of the camera 12 in the X axis; the Y1-axis stepper motor 6 and the Y2-axis stepper motor 8 control the movement of the camera 12 in the Y axis, and the rotation speeds and directions of the two motors are exactly the same; the Z-axis stepper motor 10 controls the movement of the camera 12 in the Z axis;
[0064] Central module 3: Composed of the RDK-X5 development board, control circuit board, and battery. The control circuit board is responsible for receiving serial port information and controlling the stepper motor through the PID algorithm to make the camera 12 approach and align with the eye; the RDK-X5 development board is responsible for image processing, communication, and voice broadcast functions, including:
[0065] Driving the Huiboshi 13-megapixel autofocus camera 12;
[0066] Running the YOLOv11 model target detection algorithm to obtain the human eye target detection frame;
[0067] Setting the communication baud rate to 115200 and communicating with the STM32 core board through the serial port TX-RX;
[0068] Set a fixed IP to perform SSH remote communication with the computer module in the same local area network;
[0069] Control the voice module 2 to broadcast information through the I2S interface;
[0070] Mark the abnormal part information according to the position clicked by the doctor;
[0071] Establish a mapping relationship based on the marked abnormal parts and the standard eye shape diagram, and calculate the comprehensive gain value of each mapped partition area;
[0072] Generate a heat map display according to the mapping relationship and the comprehensive gain value.
[0073] Voice module 2: Consists of the Newman BT55 mini audio voice module 2, which is used to receive the audio information of the RDK-X5, broadcast the current acquisition status, including acquisition quantity, acquisition success or failure, acquired parts, etc. acquisition information, and voice command the subject to turn the eyes in different directions to acquire different areas of the sclera;
[0074] Graphic display module: Composed of four modules: user information display, image display, doctor's marked abnormal information, and function button display. Among them, the user information display module includes the display of personal information such as doctor's login account and subject, etc. The image display module includes the five-region map of the acquired sclera and the generated heat map. The doctor's marked abnormal information module includes all abnormal area information marked by the doctor on the five-region map of the sclera. The function button module is used to execute various functions, and the button names include: start acquisition, stop acquisition, end acquisition, start marking, end marking, analysis display, save result. The graphic display module is a software module that runs in the computer module and executes the visual interaction of all the above functions.
[0075] Computer: Includes basic hardware devices and a graphic display module, and performs visual interaction functions through the embedded graphic display module. The interaction functions include:
[0076] Display all information in the graphic display module;
[0077] Click on the user information display module to input the doctor's account information and the subject's personal information;
[0078] Control the acquisition image process through the "start acquisition", "stop acquisition", and "end acquisition" buttons;
[0079] Mark the abnormal content of the sclera through the "start marking" and "end marking" buttons;
[0080] Analyze the abnormal information through the "analysis display" button to generate a heat map;
[0081] Save the user information, captured images, marked abnormal information, and heat maps by pressing the "Save Results" button.
[0082] The camera module is connected to the central module 3 via a USB cable. The motion module is connected to the control circuit board of the central module 3. The voice module 2 is connected to the RDK-X5 development board of the central module 3 via an I2S audio cable. The central module 3 is connected to the computer module via the remote SSH protocol.
[0083] Specifically describe the electrical connection relationships between the modules:
[0084] Power supply connection: The battery uses a standard 11.1V RC model battery and directly inputs power to the control circuit board. The control circuit board has a built-in buck circuit that outputs 5V to power the STM32 core board and the RDK-X5 development board. The RDK-X5 development board also powers the camera module and the voice module 2 at the same time;
[0085] Wireless signal transmission: Between the computer and the RDK-X5 development board, when the two devices are in the same local area network, the user uses the MobaXterm software, uses a fixed IP address, and remotely logs in to the RDK-X5 development board via the SSH protocol, starts the X11 remote graphical display. The RDK-X5 development board runs the graphical display module and displays the graphical interface on the computer via the X11 protocol;
[0086] Wired signal transmission: The RDK-X5 development board communicates with the STM32 core board via a serial port. The RDK-X5 development board sends control commands to the STM32 core board via the UART serial port. The STM32 core board receives the commands and controls the motor through the motor driver. Set the serial port baud rate to 115200, and use / n as the data frame header and frame tail for verification; The RDK-X5 development board sends audio data to the voice module 2 via the I2S interface. The sampling rate of I2S is selected as 44.1kHz, the bit depth is selected as 16 bits, the number of channels is 2, and the frame synchronization method is selected as the I2S standard mode. The voice module 2 receives the audio data and plays it; The camera module transmits the captured image data to the RDK-X5 development board via the USB interface. The RDK-X5 development board receives the image data and processes it. The USB transmission uses the USB2.0 transmission protocol;
[0087] Motor drive connection: In the control circuit board, the STM32 core board outputs control signals to the X-axis motor drive, Y1-axis motor drive, Y2-axis motor drive, and Z-axis motor drive. The X-axis motor drive is connected to the X-axis stepper motor 5 for driving. The Y1-axis motor drive is connected to the Y1-axis stepper motor 6 for driving. The Y2-axis motor drive is connected to the Y2-axis stepper motor 8 for driving. The Z-axis motor drive is connected to the Z-axis stepper motor 10 for driving.
[0088] SeeFigures 3-5 The mechanical frame includes a U-shaped bracket 1, an X-axis translation structure 4, a connecting rod, a Y1-axis translation structure 7, a Y2-axis translation structure 9, and a Z-axis lifting structure 11. The Y1-axis translation structure 7 and the Y2-axis translation structure 9 are symmetrically arranged. One end of the Y1-axis translation structure 7 and the Y2-axis translation structure 9 is connected by a connecting rod. The X-axis translation structure 4 is vertically arranged and connected to the Y1-axis translation structure 7 and the Y2-axis translation structure 9. The Z-axis lifting structure 11 is arranged on the X-axis translation structure 4 and is perpendicular to the X-axis translation structure 4, the Y1-axis translation structure 7, and the Y2-axis translation structure 9. The camera module is arranged on the Z-axis lifting structure 11. The U-shaped bracket 1 is fixedly vertically arranged on the Y1-axis translation structure 7 and the Y2-axis translation structure 9, and the height of the U-shaped bracket 1 matches the position of the camera 12. One motor drives the X-axis translation structure 4 to move along the X-axis, two motors respectively drive the Y1-axis translation structure 7 and the Y2-axis translation structure 9 to move synchronously along the Y-axis, and another motor drives the Z-axis translation structure. The voice module 2 and the control circuit module 3 are arranged on the connecting rod. Among them, the Y-axis moving axis is at the bottom of the whole device and has a large load capacity. To ensure balanced force, a double-axis translation structure composed of the Y1-axis translation structure 7 and the Y2-axis translation structure 9 is adopted. X-axis translation structure 4: The X-axis stepper motor 5 controls the movement of the camera on the X-axis moving axis. Double-axis translation structure: The Y1 stepper motor and the Y2 stepper motor jointly control the movement of the camera 12 on the Y-axis moving axis, and the rotational speeds and directions of the two motors are exactly the same. Z-axis lifting structure 11: The Z-axis stepper motor 10 controls the lifting movement of the camera 12 on the Z-axis moving axis.
[0089] The X-axis translation structure 4, the connecting rod, the Y1-axis translation structure 7, the Y2-axis translation structure 9, and the Z-axis lifting structure 11 all include a coupling, a smooth rod, a lead screw, and a motor bracket. A lead screw is arranged on the smooth rod, the lead screw is connected to the motor output shaft through a coupling, and the motor is arranged on the motor bracket.
[0090] When the device is in use, the head of the person to be collected is placed on the U-shaped bracket 1, and the eyes are facing the camera module.
[0091] See Figure 6 It is a regional division diagram of the standard eye shape diagram. According to the pupil position, a dividing line is made to divide the regions A1 - A8, which is used for the doctor's disease assessment.
[0092] See Figure 7 It is the method flow chart of the embodiment of the present invention. An eye white image acquisition and disease auxiliary diagnosis method includes the following steps:
[0093] S1: Obtain the video frames collected by the camera in real time, and obtain the orbital detection frame information in the video frames through the target detection algorithm. See Figure 8 Specifically, it includes the following steps:
[0094] S1.1: System initialization;
[0095] Start the computer and the image acquisition device, configure the WIFI of the computer and the RDK-X5 development board to be in the same local area network, complete the remote connection between the computer and the device through the SSH protocol, and complete the information input and parameter configuration required for system operation;
[0096] S1.2: Run the function program, open the graphic display module, and input user information;
[0097] Run the compiled code file, the system automatically opens the graphic display module, and input the doctor's account information and the personal information of the person being collected in the user information module of the graphic display module;
[0098] S1.3: Select the eye to be collected, click the "Start Collection" button, and drive the camera;
[0099] Set the sclera image to be collected as the left eye or the right eye, click the "Start Collection" button of the graphic display module, the RDK-X5 development board drives the camera 12 through the USB interface, set the collection rate to 30 frames per second, and prepare to start collecting video frames to complete the initialization configuration of image collection;
[0100] S1.4: Run the object detection algorithm for the eye socket;
[0101] According to the left eye or right eye selected in step S1.3, it is used for information statistics and to determine whether the data of the corresponding detection box sent in the subsequent detection steps is the left detection box or the right detection box.
[0102] The object detection algorithm uses the object detection model of YOLOv11 version, the input format is a standard RGB image of 640*640, and the output data format is the detection box information of [box center point coordinate X, box center point coordinate Y, box length, box width];
[0103] Each video frame is unified into a 640*640 RGB image through the image preprocessing step, and then the image is input into the object detection model. The model outputs the detection box information. After the post-processing operations of non-maximum suppression and confidence threshold filtering, the detection box information is finally obtained; if the number of detection boxes is 2, the left and right eye detection boxes are classified according to the size of the box center point coordinate X. The detection box with the smaller box center point coordinate X is the left eye object detection box, and the detection box with the larger box center point coordinate X is the right eye object detection box; finally, the RDK-X5 development board inputs the detection box information into the STM32 core board through the serial port;
[0104] The offline training process of the YOLOv11 object detection algorithm is as follows:
[0105] The training dataset uses a public face dataset, and the labels of human eyes are marked through Darklabel software, and the originalFigure 1 Use it as the training dataset. The pre-trained model uses yolov11n.pt, the initial learning rate is 0.01, the target image size is 640, the number of iterations is 100, the number of label types is 1, and the label name is: eye;
[0106] Input the training set images into the YOLOv11 training model for iterative training, and perform cross-validation through the validation set to obtain the model weight file best_eye.pt of the optimal YOLOv11 network. Import this weight file into the YOLOv11 object detection model for prediction function;
[0107] S2: Based on the orbital detection box information, adjust the displacements of the X, Y, and Z axes of the motor through the PID control algorithm;
[0108] First, call the PID algorithm to control the X-axis motor and the Z-axis motor;
[0109] The motor uses a stepper motor with model number 42. The PID algorithm runs on the STM32 core board. The PID algorithm uses the incremental PID. The specific formula of PID is as follows:
[0110] U(t) = K p ·(e(t) - e(t - 1)) + K i ·e(t) + K d ·(e(t) - 2e(t - 1) + e(t - 2)) (1)
[0111] Among them, U(t) is the PWM value of the output stepper motor, K p is the proportional gain value, K i is the integral gain value, K d is the derivative gain value, e(t) is the current input error value, e(t - 1) is the error value collected last time, and e(t - 2) is the error collected the time before last;
[0112] The calculation formula of the input error is as follows:
[0113] e(t) x = X cur - X tar (2)
[0114] Among them, e(t) x is the PWM value for controlling the X-axis stepper motor 5, X cur is the X coordinate of the center point of the orbital detection box obtained by the STM32 core board, X tar is the X coordinate of the target position;
[0115] e(t) xInput the PID formula (1), and the PWM value for controlling the X-axis stepper motor 5 is output. The STM32 core board generates the corresponding control signal and inputs it into the X-axis motor driver to complete the control of the X-axis motor;
[0116] e(t) z = Y cur - Y tar (3)
[0117] Among them, e(t) z is the PWM value for controlling the Z-axis stepper motor 10, Y cur is the Y coordinate of the center point of the orbital detection frame obtained by the STM32 core board, and Y tar is the Y coordinate of the target position;
[0118] e(t) z Input the PID formula (1), and the PWM value for controlling the Z-axis stepper motor 10 is output. The STM32 core board generates the corresponding control signal and inputs it into the Z-axis motor driver to complete the control of the Z-axis motor;
[0119] Then, call the PID algorithm to control the Y-axis stepper motor;
[0120] The PID algorithm uses formula (1);
[0121] The formula for the input error is as follows:
[0122] e(t) y = W * H - S1 (4)
[0123] Among them, W is the width of the orbital detection frame obtained by the STM32 core board, H is the height of the orbital detection frame, and S1 is the set target area (number of pixel points); W * H represents the area of the obtained orbital detection frame, and e(t) y represents the area difference, and e(t) y is input into the PID formula (1), and the PWM value of the Y-axis stepper motor is output. The STM32 core board generates the corresponding control signal and inputs it into the Y1-axis motor driver and the Y2-axis motor driver to complete the control of the Y-axis motor;
[0124] S3: Calculate the input error values of the X, Y, and Z axes in the PID control algorithm;
[0125] Detect whether the lens of the camera 12 is close to the eye, that is, calculate the input error values of the X-axis, Y-axis, and Z-axis according to formulas (2), (3), and (4);
[0126] S4: Repeat steps S1 - S3 in this embodiment until the input error value is less than the preset threshold, indicating that the camera has been aligned with the orbit, and stop the motor drive;
[0127] Determine whether the input error value is less than the lowest threshold set by the program. If it is less than the threshold, it means the lens is close enough, and the motor drive is stopped; during the calibration process, if an abnormality is encountered, the "Stop Acquisition" button can be clicked to stop the motor drive; the "End Acquisition" button can also be clicked to return the motor to its original position; among them, steps S1.1 - S1.3 are only executed once and do not participate in the repeated steps. After calibration is completed, the RDK-X5 development board transmits the audio signal to the voice module 2 through the I2S audio transmission protocol, and the voice module 2 broadcasts "Position determined".
[0128] S5: Obtain the pupil detection box information in the video frame through the object detection algorithm;
[0129] The object detection algorithm uses the object detection model of YOLOv11 version. The input format is a standard RGB image of 640*640, and the output data format is the detection box information of [coordinate X of the center point of the box, coordinate Y of the center point of the box, box length, box width];
[0130] Each video frame is unified into a 640*640 RGB image through image preprocessing, and then the image is input into the object detection model to obtain two types of output detection box information: white, black;
[0131] The offline training process of the YOLOv11 object detection algorithm is as follows:
[0132] The training dataset uses a public face dataset. The labels of the eyes and the pupils inside the eyes are marked through the Darklabel software. The eyes are marked as white, and the pupils are marked as black, which are used as the training dataset together with the original Figure 1 The pre-trained model uses yolov11n.pt, the initial learning rate is 0.01, the target image size is 640, the number of iterations is 100 times, and the number of label types is 2;
[0133] Input the training set images into the YOLOv11 training model for iterative training, and perform cross-validation through the validation set to obtain the model weight file best_black.pt of the optimal YOLOv11 network;
[0134] After obtaining all the detection boxes, perform post-processing operations of non-maximum suppression and confidence threshold filtering, and then screen the situation where only one black target box is included in the white target box, that is, the inclusion degree is greater than 90%. The inclusion degree formula is calculated as follows:
[0135]
[0136] Among them, S b represents the area of the overlapping part of the black target box and the white target box, S blackLet \(S\) represent the area of the black target box, and \(C\) represent the calculated overlap degree. In the same white target box, only the black target box with the largest area is retained, and finally the remaining black and white boxes are taken as valid boxes.
[0137] S6: Based on the pupil detection box information, the subject being sampled is commanded by voice to look in five directions: up, left, middle, right, and down.
[0138] See Figure 9 , and the sclera determination area specifically includes the following:
[0139] The entire ellipse represents an eye image. There are five types of images to be collected: upper sclera, left sclera, middle sclera, right sclera, and lower sclera. The type of the currently collected image is judged according to the area block in the eye corresponding to the center coordinates of the collected black box.
[0140] See Figure 10 , and the five-region sclera collection example is as follows:
[0141] The solid ellipse frame represents the eye socket of the whole eye, the black solid ellipse represents the pupil in the eye, the smaller dashed box represents the black box, and the larger dashed box represents the white box. The black box is included in the white box. The image type is delimited according to the determination area rule of Figure 5 .
[0142] According to the five sclera regions to be collected, the voice module 2 is commanded to broadcast "Please look in the XX direction".
[0143] S7: When it is detected that the pupil is facing the specified five directions, the video frame is saved as the sclera region map in the corresponding direction.
[0144] If the type to be collected is consistent with the target type, the voice module 2 broadcasts "Collection successful", and the current video frame is saved, named as "subject number_collection area.jpg", and saved to the subject's folder; otherwise, it broadcasts: "Collection failed, target area error, please look in the XX direction".
[0145] S8: Repeat steps S5 - S7 in this embodiment until the image collection of the sclera region maps in the five directions is completed.
[0146] Count all the pictures in the subject's folder. If the five-region requirement is not met, repeat steps S5 - S7. If it has been met, broadcast "XX region has been collected", the collection ends, and the motor returns to its original position.
[0147] S9: Feature extraction is performed on the collected images, and based on the offline-trained medical segmentation model, the information of abnormal parts is automatically labeled.
[0148] The medical segmentation model uses SAM-Med2D. The input format is a high-precision sclera image with a resolution of 1024*1024. The output image is a mask image with different colors, where different colors represent different categories of abnormalities, including: blood streaks, black spots, blood patches, and yellow spots. It also outputs relevant evaluation metrics for the segmentation results, such as the F1 score and Dice coefficient, which are used to measure the accuracy and quality of the segmentation.
[0149] The offline training process of the SAM-Med2D segmentation model is as follows:
[0150] The training dataset uses a public sclera dataset. Through the QuaPath medical segmentation tool, the abnormal labels in the sclera image are marked as the label set and used together with the original Figure 1 as the training dataset. The model is trained based on the MONAI medical model framework. The training set images are input into the SAM-Med2D segmentation model for iterative training, and cross-validation is performed through the validation set to obtain the optimal model weight file.
[0151] S10: According to the annotated image, establish a dynamic coordinate system and mapping area with the pupil center as the pole, and calculate the polar coordinate parameters, the radial distance r and the polar angle θ;
[0152]
[0153] According to formula (6), where (x0, y0) is the pupil center coordinate, which serves as the pole of the polar coordinate system and is the reference point for calculating the relative positions of other points. (x i , y i ) is the center point coordinate of the abnormal area in the marked two-dimensional coordinate system, and i is each abnormal serial number; the mapping area is the standard eye shape area of a person, which is used as a position reference for mapping the abnormal part under the standard eye shape. The purpose is to adapt to different people's eye shapes through normalization and eliminate individual eye shape differences. The mapping area is divided into eight regions A1 - A8 by straight lines according to the pupil position for doctors' disease assessment; d max and d min are the difference between the maximum and minimum effective radii of the mapping area, representing the range size of the mapping area. S max is the maximum visible radius of the sclera, is the Euclidean distance from the pupil center coordinate to the center point coordinate of the abnormal area, is the scaling coefficient between the mapping area and the sclera image area;
[0154] The doctor clicks the "Analysis Display" button, and the system starts to analyze the image and the marked content. By establishing a mathematical model, the abnormal areas marked are standardized. Considering that the eyes of different people are different, if the traditional static two-dimensional coordinate system depends on external fixed reference points, it is easily affected by actions such as slight head movement or blinking displacement. Therefore, the method of using a polar coordinate system is adopted. With the pupil center as the pole, a polar coordinate system is established, and a dynamic coordinate system mapping model is constructed, which has stronger robustness. By calculating the Euclidean distance between the feature point (x i , y i ) and the pupil center (x0, y0), the physical coordinates in the image are mapped to the polar radius r. At the same time, the polar angle θ relative to the pupil center coordinates is calculated. In subsequent coordinate calculations, the coordinate points in the polar coordinate system are calculated in the form of (r, θ);
[0155] S11: According to the coordinate system information, calculate the spatial weight value w of the mapped anomaly through the distance attenuation function j ;
[0156]
[0157] where p i is the central coordinate of the abnormal position in the polar coordinate system, represented by (r i , θ i ), where i is the abnormal serial number; c j is the central coordinate of the mapping area numbered A1 - A8, where j is the mapping area number; σ is the standard radius of the mapping area, n is the number of abnormal parts marked by the doctor, and the final result w j is the spatial weight value of the j-th area of the mapping area number;
[0158] w j quantifies the spatial aggregation degree of the feature points related to each abnormal center j. In practical applications, if the number of abnormal areas marked in a certain area is large, that is, when the abnormal points cluster at the core of the reflection area, the calculated reflection area weight w j increases significantly;
[0159] In actual parameter debugging, by adjusting σ, the probability attenuation rate can be controlled, and the distance correlation strength can be controlled. In practical applications, if the actual eyes of the patient in the image are larger than the average area of human eyes, the value of σ needs to be appropriately increased to expand the effective action range; generally, σ takes 1 - 2 times the standard radius of the reflection area;
[0160] S12. Calculate the comprehensive gain value G j of the mapping area according to parameters such as the spatial weight value w j ;
[0161]
[0162] where α is the spatial weight coefficient, w j is the spatial weight value calculated in formula (7), β is the outlier density weight coefficient, represents the number of the j-th outlier in the total outliers, indicating the outlier density, γ is the clinical prior knowledge weight coefficient, and C j is the number of historically confirmed cases in this mapped area; the log logarithmic transformation is added to reduce the magnitude difference of the cases. This formula balances the three conditions of abnormal spatial distribution, abnormal density, and confirmed cases, and adjusts the priority of the corresponding conditions through the weight coefficient;
[0163] According to formula (8), the proportion of the three weight coefficients of abnormal spatial distribution, abnormal density, and clinical prior knowledge weight coefficient reflects the importance of the corresponding class conditions. Under the default parameters, the spatial weight coefficient is taken as 0.6, the outlier density coefficient is taken as 0.3, and the clinical prior knowledge coefficient is taken as 0.1;
[0164] S13. Normalize the comprehensive gain value G j to generate a sclera abnormality heat map and display it in grades to assist doctors for reference.
[0165] Furthermore, the normalization process of the gain value is as follows:
[0166]
[0167] where G j is the comprehensive gain value, G min is the minimum value of the comprehensive gain degree of all abnormalities, G max is the maximum value of the comprehensive gain degree of all abnormalities, and finally the heat value G' j is obtained. This operation maps the numerical value to the interval [0, 1]. Using the linear mapping method and the hot color mapping scheme in the Matplotlib library of Python, corresponding colors are filled in different regions on the mapping graph to obtain the heat map of the mapped area, which is displayed in the image display module of the computer to assist doctors for reference.
[0168] Furthermore, it is also necessary to generate a sclera abnormality grading heat map of the original sclera image. The original sclera image is calculated based on a two-dimensional coordinate system, and Gaussian kernel density estimation is introduced to perform spatial smoothing on the discrete gain values:
[0169]
[0170] where x is the x coordinate in the heat map, y is the y coordinate in the heat map, x j is the x coordinate of the center of the abnormal area, and y jy is the y - coordinate of the center of the abnormal area, n is the number of abnormal parts marked by the doctor, and i is the abnormal serial number; ρ is a parameter for controlling the bandwidth of the kernel function, which is used to adjust the smoothness of color switching in the heat map, and H(x, y) is the gain value of this pixel point.
[0171] According to formula (10), Gaussian kernel density estimation is introduced to spatially smooth the discrete gain values, obtain the heat value of each pixel point in the heat map, generate the heat map, and display it in the image display module. In practical applications, ρ is used to adjust the smoothness of color switching in the heat map, and doctors can adjust the value of ρ according to their own observation needs; in the probability distribution of the mapping area of the heat map, red - yellow is used to simply represent the lesion risk level to assist doctors in identifying areas with high incidence of abnormalities, meeting the intuitive requirements of clinical diagnosis.
[0172] Doctors can click the "Save Results" button of the graphic display module, and the system automatically generates a report based on the automatically marked content and images to save the results.
[0173] The above embodiments are implemented on the premise of the technical solution of the present invention, and detailed implementation methods and specific operation processes are given. However, the protection scope of the present invention is not limited to the above embodiments. The methods used in the above embodiments are conventional methods unless otherwise specified.
Claims
1. An eye white image acquisition and disease auxiliary diagnosis method, characterized in that, It includes the following steps: S1. Obtain the video frames captured by the camera in real time, and obtain the orbital detection box information in the video frames through the target detection algorithm; S2. Based on the orbital detection box information, adjust the displacements of the X, Y, and Z axes of the motor through the PID control algorithm; S3. Calculate the input error values of the X, Y, and Z axes in the PID control algorithm; S4. Repeat steps S1 - S3 until the input error value is less than the preset threshold, indicating that the orbit has been located, and stop the motor drive; S5. Obtain the video frames captured by the camera in real time, and obtain the pupil detection box information in the video frames through the target detection algorithm; S6. Based on the pupil detection box information, command the subject to look in the five directions of up, left, middle, right, and down through voice; S7. When the pupil is detected to be facing the specified five directions, save the current video frame as the sclera area map in the corresponding direction; S8. Repeat steps S5 - S7 until the image acquisition of the sclera area maps in the five directions is completed; S9. Extract the features of the acquired images, and automatically label the abnormal parts based on the offline trained medical segmentation model; S10. According to the labeled images, establish a dynamic coordinate system and a mapping area with the pupil center as the pole, and calculate the polar coordinate parameters, the polar radius r and the polar angle θ: Among them, (x0, y0) is the pupil center coordinate, and (x i , y i ) is the center point coordinate of the abnormal area in the marked two-dimensional coordinate system, and i is each abnormal serial number; the mapping area is the standard eye shape area of a person, which is used as a position reference for mapping the abnormal part under the standard eye shape. The mapping area is divided into eight regions A1 - A8 by straight lines according to the pupil position; d max and d min are the differences between the maximum and minimum effective radii of the mapping area, representing the range size of the mapping area, S max is the maximum visible radius of the sclera, is the Euclidean distance from the pupil center coordinate to the center point coordinate of the abnormal area, is the scaling factor between the mapping area and the sclera image area. The coordinate points in the polar coordinate system are calculated in the form of (r, θ); S11. Calculate the spatial weight value w of the mapping anomaly through a distance attenuation function according to the coordinate system information j : Among them, p i is the central coordinate of the abnormal position in the polar coordinate system, represented by (r i , θ i ), c j is the central coordinate of the mapping area numbers A1 - A8, where j is the mapping area number; σ is the standard radius of the mapping area, n is the number of marked abnormal parts, and the final result w j is the spatial weight value of the area with the mapping area number j; S12. Calculate the comprehensive gain value G of the mapping area j ; Among them, α is the spatial weight coefficient, β is the outlier density weight coefficient, represents the number of the j-th outlier in the total outliers, indicating the outlier density, γ is the clinical prior knowledge weight coefficient, and C j is the number of historically confirmed cases in this mapped area; S13. Normalize the comprehensive gain value G j to generate a sclera abnormality heat map and display it in a graded manner.
2. The method for collecting sclera images and assisting in disease diagnosis according to claim 1, wherein The gain value normalization process is as follows: Among them, G j is the comprehensive gain value, G min is the minimum value of the comprehensive gain degrees of all anomalies, G max is the maximum value of the comprehensive gain degrees of all anomalies, and finally the heat value G' j is obtained. This operation maps the value to the interval [0, 1].
3. The method for collecting sclera images and assisting in disease diagnosis according to claim 2, characterized in that, It also includes generating a sclera abnormal grading heat map of the original sclera image. The original sclera image is calculated based on the two-dimensional coordinate system, and the Gaussian kernel density estimation is introduced to perform spatial smoothing on the discrete gain values: where x is the x - coordinate in the heat map, y is the y - coordinate in the heat map, x j is the x - coordinate of the center of the abnormal area, y j is the y - coordinate of the center of the abnormal area; ρ is a parameter for controlling the kernel bandwidth, which is used to adjust the smoothness of color switching in the heat map, H(x, y) is the gain value of this pixel point. Through normalization processing and color mapping, the heat map of the sclera abnormal grading of the original sclera image is obtained.
4. An eye white image acquisition and disease auxiliary diagnosis device, characterized in that, It includes: Camera module: Used for visual positioning and capturing high-resolution sclera images; Mechanical module: The mechanical framework that supports all other modules; Motion module: Composed of 4 motors, applied to three-axis motion control, 1 motor for the X axis, 2 motors for the Y axis, and 1 motor for the Z axis; Central module: The central brain of the entire device, used to drive the motors and process information; Voice module: Used to broadcast the current information acquisition status; Graphic display module: A software module built into the computer, used for the visual display of functions such as inputting information, displaying image information, and annotating information; Computer module: Used to automatically annotate information and display all information; The camera module, the motion module, and the voice module are all connected to the central module. The central module is connected to the computer module, and the graphic display module is built into the computer module.
5. The oculus sclera image acquisition and disease assisted diagnosis device according to claim 4, wherein, The mechanical frame includes a U-shaped bracket, an X-axis translation structure, a connecting rod, a Y1-axis translation structure, a Y2-axis translation structure, and a Z-axis lifting structure. The Y1-axis translation structure and the Y2-axis translation structure are symmetrically arranged. One end of the Y1-axis translation structure and the Y2-axis translation structure is connected by a connecting rod. The X-axis translation structure is vertically arranged and connected to the Y1-axis translation structure and the Y2-axis translation structure. The Z-axis lifting structure is arranged on the X-axis translation structure and is perpendicular to the X-axis translation structure, the Y1-axis translation structure, and the Y2-axis translation structure. The camera module is arranged on the Z-axis lifting structure. The U-shaped bracket is fixedly arranged vertically on the Y1-axis translation structure and the Y2-axis translation structure, and the height of the U-shaped bracket matches the position of the lens of the camera module. One motor drives the X-axis translation structure to move along the X-axis, two motors respectively drive the Y1-axis translation structure and the Y2-axis translation structure to move synchronously along the Y-axis, and another motor drives the Z-axis lifting structure. The voice module and the control circuit module are arranged on the connecting rod.
6. The device for collecting sclera images and assisting in disease diagnosis according to claim 5, characterized in that, The X-axis translation structure, the connecting rod, the Y1-axis translation structure, the Y2-axis translation structure, and the Z-axis lifting structure all include a coupling, a smooth rod, a lead screw, and a motor bracket. A lead screw is arranged on the smooth rod. The lead screw is connected to the motor output shaft through a coupling. The motor is arranged on the motor bracket.
7. An eye white image acquisition and disease auxiliary diagnosis device according to claim 4, characterized in that, The current information acquisition status includes the acquisition quantity, whether the acquisition is successful, and the acquired part.
Citation Information
Cited By
Method for measuring percentage of visible eye white of dairy cow
CN120827335A
Multispectral dynamic visual function training system and method
CN120860500A