A dual-lens capsule endoscope intelligent film reading system and film reading method

Through the dual-lens capsule endoscopic intelligent film reading system and the intelligent analysis algorithm combined with deep learning methods, the problem of the capsule endoscopic system lacking efficient image processing and intelligent analysis in the existing technology is solved, and more efficient and accurate diagnosis of digestive tract diseases is achieved.

CN118969207BActive Publication Date: 2025-05-13WUXI FUSHENG SMART MEDICAL TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411086824.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-08
Publication Date
2025-05-13
Estimated Expiration
2044-08-08

AI Technical Summary

Technical Problem

The existing capsule endoscopy system lacks efficient image processing and intelligent analysis methods, which leads to a lot of time and energy required by doctors during the film reading process, and is susceptible to subjective factors, affecting the accuracy of the diagnosis. At the same time, single-lens capsules have problems such as limited field of vision, many blind spots, and inaccurate diagnosis.

Method used

It adopts a dual-lens capsule endoscope intelligent film reading system, integrating image acquisition and processing module, wireless transmission module, receiver module and intelligent film reading module. The system uses a high-definition, low-distortion dual-lens design, combined with deep learning methods of lesion abnormality detection and polyp detection and quantitative analysis algorithms to achieve comprehensive monitoring and intelligent analysis of the digestive tract.

Benefits of technology

It achieves more efficient and accurate diagnosis of digestive tract diseases, reduces the time and burden of doctors for reading the film, improves the accuracy and efficiency of diagnosis, and reduces the probability of misdiagnosis or missed diagnosis due to subjective factors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118969207B_ABST
    Figure CN118969207B_ABST
Patent Text Reader

Abstract

The present invention discloses a dual-lens capsule endoscope intelligent film reading system and film reading method, the film reading system includes a dual-lens capsule endoscope body, a receiver module and an intelligent film reading module; the intelligent film reading module includes: a permission management module, a case management module, a fast film reading module and a report editing and output module, the fast film reading module: for performing deduplication, abnormal lesion detection and polyp detection and quantitative analysis on the image data stored in the receiver module, and performing dual-lens film reading by means of different frame rate film reading, thumbnail fast preview, thumbnail positioning, forward and backward film reading, frame-by-frame forward and backward film reading and / or lens mode switching. The present invention combines the dual-lens design with the intelligent film reading algorithm, realizes comprehensive monitoring and intelligent analysis of the front and back directions of the body; introduces the deduplication algorithm, abnormal lesion detection algorithm and polyp detection and quantitative analysis algorithm, and significantly improves the efficiency and accuracy of diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to an intelligent film reading system and a film reading method for a dual-lens capsule endoscope, belonging to the technical field of capsule endoscopes. Background Art

[0002] Traditional endoscopic examination methods often require patients to endure great discomfort and are complicated to operate, which limits their widespread clinical application. With the advancement of medical technology, capsule endoscopes have gradually emerged as a non-invasive and convenient examination tool. Capsule endoscopes are tiny wireless devices that use the natural peristalsis of the digestive tract to capture images of the inside of the digestive tract and transmit them wirelessly to an external receiving device. However, most existing capsule endoscope systems only have basic image acquisition and transmission functions, lack efficient image processing and intelligent analysis methods, which requires doctors to spend a lot of time and energy in the process of reading films, and are easily affected by subjective factors, affecting the accuracy of diagnosis. At the same time, most traditional capsules are single-lens, with problems such as limited field of view, many blind spots, and inaccurate diagnosis. Summary of the invention

[0003] The present invention provides a dual-lens capsule endoscope intelligent film reading system and film reading method, which integrates advanced image acquisition and processing technology, wireless transmission technology and intelligent film reading technology, and realizes more efficient and accurate diagnosis of digestive tract diseases.

[0004] In order to solve the above technical problems, the technical solutions adopted by the present invention are as follows:

[0005] A dual-lens capsule endoscope intelligent film reading system, comprising: a dual-lens capsule endoscope body, a receiver module and an intelligent film reading module;

[0006] The main body of the dual-lens capsule endoscope includes an image acquisition and processing module and a wireless transmission module. The image acquisition and processing module is used to acquire and process images, and the wireless transmission module is used to transmit the images processed by the image acquisition and processing module to the receiver module.

[0007] The receiver module is used for receiving, processing, storing and displaying images;

[0008] The intelligent film reading module is used to quickly read, analyze and generate reports on the image data stored in the receiver module;

[0009] The intelligent film reading module includes:

[0010] Permission management module: used for permission control, only authorized doctors can access and operate the system;

[0011] Case management module: used to add, delete, modify and query case information, so that doctors can quickly obtain and maintain case data;

[0012] Fast film reading module: used for deduplication, lesion abnormality detection and polyp detection and quantitative analysis of the image data stored in the receiver module, and dual-lens film reading by reading at different frame rates, quick thumbnail preview, thumbnail positioning, forward and backward film reading, frame-by-frame forward and backward film reading and / or lens mode switching; wherein, deduplication is to identify and remove image frames with a similarity greater than 98% in consecutive frames (i.e., repeated or highly similar image frames), thereby reducing the burden and time of doctors reading films; lesion abnormality detection is to train a lesion abnormality detection model using a deep learning method, and use the obtained lesion abnormality detection model to analyze feature information such as texture, color and shape in the image, identify possible lesion areas (such as inflammation, ulcers, tumors, etc.), and mark them in the form of highlights or frames; polyp detection and quantitative analysis is to identify polyp areas and draw precise boundary boxes, and use image processing methods to measure the size and shape of polyps, and grade and risk assess polyps according to preset evaluation criteria;

[0013] Report editing and output module: used to automatically capture lesion images after film reading is completed, and allow doctors to add explanatory text to generate a detailed PDF report containing patient information, image data, diagnosis results and suggestions.

[0014] The dual lenses of the dual-lens capsule endoscope body can realize comprehensive monitoring and image acquisition in both the front and back directions inside the body.

[0015] The image acquisition and processing module uses high-definition, low-distortion dual lenses, built-in optical system and CMOS sensor, as well as image processing chip to achieve preliminary image processing.

[0016] The wireless transmission module uses low-power, high-efficiency wireless communication technology to transmit the processed image data to the receiver module outside the body in real time.

[0017] The receiver module is designed with a compact shape to reduce the size and weight of the receiver; it adopts a touch screen mode to reduce the number of physical buttons, making the device light and easy to carry; it has a built-in 32G large-capacity storage module that can store 3 million images at a time; it is also equipped with a 480x800 high-resolution touch screen display that can receive and display the dual-lens image data from the capsule endoscope in real time. It uses a long-lasting battery to ensure the patient's freedom of movement during the examination.

[0018] In order to solve the problem of duplicate images that may be generated during the movement of capsule endoscopes, this application has built-in efficient deduplication function. Deduplication can automatically identify and remove duplicate or highly similar image frames in consecutive frames, thereby reducing the burden and time cost of doctors reading images.

[0019] The lesion anomaly detection model uses Vision Transformer (ViT), which has global context capture capabilities, self-attention mechanism, and multi-layer encoder architecture, enabling the model to have a high balance between sensitivity and specificity, ensuring the accuracy of the detection results.

[0020] This application provides a special polyp detection and quantitative analysis function for polyps, a common type of digestive tract lesions.

[0021] The lesion abnormality detection algorithm and the polyp detection and quantitative analysis algorithm have been trained and optimized with a large amount of data, and have a high recognition rate and a low false alarm rate.

[0022] After the film reading is completed, this application automatically captures the lesion image and allows the doctor to add explanatory text, and finally generates a detailed PDF report containing patient information, image data, diagnosis results and suggestions. It greatly simplifies the doctor's workflow and improves the efficiency and quality of report generation. Select the specified save location to save.

[0023] The present invention fully considers the comfort of patients and the convenience of use of doctors, and has good clinical application prospects and promotion value.

[0024] The above-mentioned fast film reading module: uses XAML to define the user interface, including dual-lens image display in the main area, lens mode switching buttons, frame rate control components (acceleration and deceleration buttons and frame rate progress bar), thumbnail progress bar, forward and backward buttons, frame-by-frame forward and backward buttons, and lesion capture area; uses the .NET data binding mechanism to bind the image data and case information received by the receiver to the XAML interface elements to achieve dynamic display and update of data; implements a frame rate manager to dynamically adjust the image update rate within a range of 1-40 frames per second by monitoring the events of the frame rate control component; draws thumbnails on the thumbnail progress bar, monitors mouse events to display the thumbnail of the current position, and processes mouse click events to achieve rapid positioning; monitors the mouse double-click event in the main film reading area, determines the double-click position and captures the corresponding lesion image data, and saves it to the local file system. This module displays images from two lenses simultaneously, making it easier for doctors to make comparative observations and improving the accuracy and efficiency of diagnosis. It allows doctors to adjust the reading rate as needed, from slow motion to fast browsing, to meet the reading needs in different scenarios. Through the thumbnail preview function, doctors can quickly locate the lesion and reduce unnecessary browsing time.

[0025] This application simultaneously uses deduplication, abnormal lesion detection, polyp detection and quantitative analysis for intelligent film reading, which reduces the workload of doctors, speeds up film reading and increases the accuracy of film reading.

[0026] In order to improve the efficiency and accuracy of deduplication, a deduplication algorithm is used to extract the features of the image using a visual model, and the vector list is divided into small batches for calculation; the parallelism of matrix operations is used to improve the computational efficiency; matrix multiplication or batch calculation of cosine similarity is used, and finally images with a similarity greater than 98% are filtered out, greatly reducing the burden and time cost of doctors reading films; compared with traditional image deduplication methods, such as deduplication based on hash algorithms, this deduplication algorithm extracts image features through visual models and uses the parallelism of matrix operations to significantly improve computational efficiency.

[0027] In order to improve the accuracy of lesion abnormality recognition, the Vision Transformer (VIT) algorithm is used for lesion abnormality detection. The input two-dimensional image is divided into multiple fixed-size image blocks (for example, 16x16 pixels), and each image block is flattened into a one-dimensional vector. The one-dimensional vector of each image block is embedded into a fixed-dimensional feature vector space through linear transformation to form an embedded representation of the image block. Since the image block sequence is disordered, Vision The Transformer (VIT) algorithm retains the position information of the image block by adding position encoding to ensure that the model can capture the spatial structure of the image; collects and annotates a large number of medical image data sets, uses the cross entropy loss function and Adam optimizer to optimize the parameters, and continuously trains the model; then inputs the embedded image block sequence into the Transformer encoder, which contains 12 layers of self-attention layers (Self-Attention) and feedforward neural network layers. Through the self-attention mechanism, the model can dynamically focus on important areas in the image and capture global and local features; by adding a classification head to the last feature vector, the classification results of the image block are output, and these classification results are used to determine whether there are lesions in each area of ​​the image; the annotated image results are displayed in real time in the main area to help doctors quickly and accurately identify and diagnose lesions; compared with some algorithms that only focus on image features without considering position information, VIT retains the position information of the image block by adding position encoding. This helps the model better understand the spatial structure of the image when identifying lesions and improves the accuracy of detection; VIT's self-attention mechanism also enables the model to dynamically focus on important areas in the image, further improving the accuracy of detection.

[0028] In this application, multiple refers to more than two.

[0029] In order to improve the efficiency and accuracy of the analysis, polyp detection and quantitative analysis are as follows: during the reading process, the detection algorithm YoloV9 is called, the input image is adjusted to a fixed size, and the features of the image are extracted using a multi-layer convolutional neural network to generate a feature map; the feature map is divided into an SxS grid, and each grid unit is responsible for detecting targets in a part of the image; each grid unit predicts three bounding boxes, each of which contains the position of the target (center coordinates, width, and height) and a confidence score (indicating the probability that the bounding box contains the target), and each bounding box also predicts the category probability distribution of the target to determine whether the target belongs to a polyp; YoloV9 uses a joint loss function, including position loss (bounding box regression error), confidence loss (confidence error of predicting that the bounding box contains the target), and classification loss (target category prediction error) to optimize the model parameters; through data enhancement methods (such as random cropping, rotation In the process of reasoning, the non-maximum suppression algorithm is used to remove redundant bounding boxes with low confidence and high overlap, and the bounding box with the highest confidence is retained as the final detection result, which helps doctors to diagnose polyps quickly and accurately, and then save the lesion image through the lesion capture function; some traditional two-stage detection algorithms (such as the R-CNN series) need to generate candidate regions first and then perform classification and regression, and the detection speed is slow, while YoloV9 adopts a one-stage detection strategy, which directly predicts the category and position of the target in the output layer. This strategy gives YoloV9 a significant advantage in detection speed and can achieve real-time detection; YoloV9 provides a simple and easy-to-use API interface and pre-trained models to facilitate users to develop and apply target detection tasks, which is particularly important for scenarios where rapid diagnosis is required in clinical applications.

[0030] The above-mentioned receiver module receives the radio frequency signal from the wireless transmission module and converts it into a usable electrical signal, then uses an analog-to-digital converter to convert the electrical signal into a digital signal, and after error detection and correction and data compression and decompression, the digital signal is restored to the original data, and then the original data is reorganized in JPEG format to form complete image data for storage and display.

[0031] The above-mentioned permission management module has a database containing user information (storing user names, encrypted passwords, roles, etc.). When a user logs in, the user password and private key are mixed and encrypted and stored using the MD5 hash algorithm, and compared with the encrypted password stored in the database, and the verification result is returned.

[0032] The above encryption method using user password plus private key can doubly guarantee the security of user information.

[0033] The above case management module is as follows: after entering the system, doctors with operation authority can query the corresponding case according to the patient's name, age and medical record number in the case, and check the patient's report. According to the actual situation, they can add case information to facilitate subsequent medical management; if the case information is incorrect, it can be modified or deleted; case data is stored in SQLite, which takes up very little resources and can be combined with many programming languages, such as C#, PHP, Java, etc., and has an ODBC interface. Compared with other databases, it has a faster processing speed.

[0034] A method for reading images with a capsule endoscope that integrates an intelligent reading function, wherein a dual-lens capsule endoscope transmits collected and processed images to a receiver wirelessly, performs deduplication, lesion abnormality detection, polyp detection and quantitative analysis on the image data in the receiver, and performs dual-lens reading by reading images at different frame rates, thumbnail quick preview, thumbnail positioning, forward and backward reading, frame-by-frame forward and backward reading and / or lens mode switching; wherein deduplication is to identify and remove image frames with a similarity greater than 98% in consecutive frames (i.e., repeated or highly similar image frames), thereby reducing the burden and time of doctors reading images; lesion abnormality detection is to use a deep learning method to detect and remove the image frames with a similarity greater than 98% in consecutive frames, thereby reducing the burden and time of doctors reading images; and The lesion abnormality detection model is trained and used to analyze the feature information such as texture, color and shape in the image, identify possible lesion areas (such as inflammation, ulcers, tumors, etc.), and mark them in the form of highlights or frames; polyp detection and quantitative analysis are to identify polyp areas and draw precise bounding boxes, and use image processing methods to measure the size and shape of polyps, and grade and risk assess polyps according to preset evaluation criteria; after the film reading is completed, the lesion image is automatically captured, and the doctor is allowed to add explanatory text to generate a detailed PDF report containing patient information, image data, diagnosis results and suggestions.

[0035] The above deduplication adopts a deduplication algorithm, uses a visual model to extract the features of the image, divides the vector list into small batches for calculation; uses the parallelism of matrix operations to improve the calculation efficiency; uses matrix multiplication or batch calculation of cosine similarity, and finally filters out images with a similarity greater than 98%;

[0036] The above-mentioned lesion abnormality detection uses the Vision Transformer (VIT) algorithm to divide the input two-dimensional image into multiple fixed-size image blocks (for example, 16x16 pixels), and each image block is flattened into a one-dimensional vector; the one-dimensional vector of each image block is embedded into a fixed-dimensional feature vector space through linear transformation to form an embedded representation of the image block; since the image block sequence is disordered, Vision The Transformer (VIT) algorithm retains the position information of the image blocks by adding position encoding to ensure that the model can capture the spatial structure of the image; collects and annotates a large number of medical imaging data sets, uses the cross entropy loss function and Adam optimizer to optimize the parameters, and continuously trains the model; then inputs the embedded image block sequence into the Transformer encoder, which contains 12 layers of self-attention layers (Self-Attention) and feedforward neural network layers. Through the self-attention mechanism, the model can dynamically focus on important areas in the image and capture global and local features; by adding a classification head to the final feature vector, the classification results of the image blocks are output, and these classification results are used to determine whether there are lesions in each area of ​​the image; the annotated image results are displayed in real time in the main area to help doctors quickly and accurately identify and diagnose lesions;

[0037] The above polyp detection and quantitative analysis is as follows: during the reading process, the detection algorithm YoloV9 is called, the input image is adjusted to a fixed size, and the image features are extracted using a multi-layer convolutional neural network to generate a feature map; the feature map is divided into an SxS grid, and each grid unit is responsible for detecting targets in a certain area of ​​the image; each grid unit predicts three bounding boxes, each of which contains the location of the target (center coordinates, width, and height) and a confidence score (indicating the probability that the bounding box contains the target), and each bounding box also predicts the category probability distribution of the target to determine whether the target belongs to a polyp; YoloV9 A joint loss function, including position loss (bounding box regression error), confidence loss (confidence error that the predicted bounding box contains the target) and classification loss (target category prediction error), is used to optimize model parameters. Data enhancement methods (such as random cropping, rotation, scaling, etc.) are used to increase the diversity of training data and improve the generalization ability of the model. During the inference process, redundant bounding boxes with low confidence and high overlap are removed through the non-maximum suppression algorithm, and the bounding box with the highest confidence is retained as the final detection result, helping doctors to quickly and accurately diagnose polyps, and then save the lesion image through the lesion capture function.

[0038] This application is suitable for non-invasive examination of internal environments such as the digestive tract, and achieves comprehensive optimization from image acquisition, wireless transmission, real-time display to intelligent analysis, greatly improving the efficiency and accuracy of medical diagnosis.

[0039] The technologies not mentioned in the present invention are all referred to the prior art.

[0040] The dual-lens capsule endoscope intelligent film reading system of the present invention combines the dual-lens design with the intelligent film reading algorithm to realize comprehensive monitoring and intelligent analysis of the front and back directions of the body; the receiver module supports real-time display of dual-lens images, which greatly increases the timeliness of detection; the introduction of deduplication algorithm, lesion abnormality detection algorithm and polyp detection and quantitative analysis algorithm significantly improves the efficiency and accuracy of diagnosis; the lesion abnormality detection algorithm and polyp detection and quantitative analysis algorithm in the intelligent film reading software have been trained and optimized with a large amount of data, and have a high recognition rate and a low false alarm rate; the overall design of the system fully considers the comfort of patients and the convenience of use of doctors, and has good clinical application prospects and promotion value. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 It is a schematic diagram of the overall structure of a dual-lens capsule endoscope intelligent film reading system. DETAILED DESCRIPTION

[0042] In order to better understand the present invention, the content of the present invention is further explained below in conjunction with the embodiments, but the content of the present invention is not limited to the following embodiments.

[0043] Example 1

[0044] like Figure 1 As shown, the dual-lens capsule endoscope intelligent film reading system includes a dual-lens capsule endoscope body, a receiver module and an intelligent film reading system.

[0045] Dual-lens capsule endoscope body: The dual-lens capsule endoscope body includes an image acquisition and processing module and a wireless transmission module. The image acquisition and processing module is used to acquire and process images, and the wireless transmission module is used to transmit the images processed by the image acquisition and processing module to the receiver module. The image acquisition and processing module adopts a high-definition, low-distortion dual-lens design. It uses a medical lens of model HR0502A, which can simultaneously capture images in both the front and back directions of the digestive tract, providing a wider field of view and more comprehensive image coverage, which is helpful for later image analysis. The use of the HM0360 CMOS sensor ensures the high resolution and color reproduction of the image. Through the RF chip in the wireless transmission module, the processed image data is stably transmitted to the receiver module outside the body in a low-latency manner. At the same time, in order to ensure the security of data transmission, custom header data is added to the image data, and then encrypted using MD5 encryption technology to prevent data leakage. In this case, the dual-lens capsule endoscope body uses a 130mAh long-life battery of model CR9885 to ensure the patient's freedom of movement during the examination.

[0046] Receiver module: The receiver module is used to receive image data from the capsule endoscope body in real time, and process, store and display the image data; it is equipped with a 480x800 high-resolution touch screen display, which can simultaneously display the images captured by the dual lenses in the front and back directions. The receiver module adopts wireless receiving technology to receive radio frequency signals and convert them into usable electrical signals, and then uses an analog-to-digital converter to convert the electrical signals into digital signals. After error detection and correction and data compression and decompression, the digital signals are restored to the original data; then the original data is reorganized according to a specific JPEG format to form complete image data and transmitted to the display and storage module. The receiver module adds a timestamp to each frame of the image to ensure synchronization during the image processing process; at the same time, the decoding and rendering algorithms are optimized to ensure the real-time display of the dual-lens images, which helps to observe the internal structure of the digestive tract more comprehensively. The receiver module adopts a compact design to reduce the size and weight of the receiver; the touch screen mode is adopted to reduce the number of physical buttons, making it easier for patients to carry it with them. The shell of the receiver module is made of wear-resistant, non-slip, and high-temperature resistant materials ABS+PC, ensuring stability and comfort during daily activities such as walking and sitting, greatly reducing the patient's discomfort. In order to meet the patient's needs for long-term wearing, the receiver module has a built-in 5000mAh large-capacity lithium battery to support long-term continuous use. Using the RK3566 core board, it has a 32GB large-capacity storage module, which can effectively ensure that all image data sent by the endoscope body are successfully saved and then transmitted to the intelligent film reading system.

[0047] Intelligent film reading module: used for rapid film reading, intelligent analysis and report generation of image data stored in the receiver module. Including:

[0048] Permission management module: Design a database containing a user information table (storing user names, encrypted passwords, roles, etc.). This module receives the user name and password entered by the user, performs encryption comparison operations, and returns the verification results. The MD5 hash algorithm is used to encrypt and store the user password and private key. When the user logs in, the system encrypts the password and private key entered by the user with MD5 and compares them with the encrypted password stored in the database. The encryption method of user password + private key can double guarantee the security of user information.

[0049] Case management module: supports adding, deleting, modifying and querying case information, making it convenient for doctors to quickly obtain and maintain case data. After entering the system, doctors with operation permissions can query the corresponding cases according to the patient's name, age, medical record number and other conditions in the case, and check the patient's report. Doctors can add case information based on the actual situation of the patient to facilitate subsequent treatment management. If the case information is incorrect, it can be modified and deleted. Case data is stored in SQLite, which takes up very little resources and can be combined with many programming languages, such as C#, PHP, Java, etc., and has an ODBC interface. Compared with other databases, it has a faster processing speed.

[0050] Fast film reading module: used to perform deduplication, abnormal lesion detection and polyp detection and quantitative analysis on the image data stored in the receiver module, and perform dual-lens visual intelligent film reading through different frame rate reading, thumbnail quick preview, thumbnail positioning, forward and backward film reading, frame-by-frame forward and backward film reading and / or lens mode switching. Use XAML to define the user interface, including dual-lens image display in the main area, lens mode switching button, frame rate control component (acceleration and deceleration buttons and frame rate progress bar), thumbnail progress bar, forward and backward buttons, frame-by-frame forward and backward buttons, and lesion capture area. Use the .NET data binding mechanism to bind the image data and case information received by the receiver to the XAML interface elements to realize dynamic display and update of data. Implement a frame rate manager to dynamically adjust the image update rate by listening to the events of the frame rate control component, with the rate range of 1-40 frames / second. Draw thumbnails on the thumbnail progress bar, listen to mouse events to display the thumbnail of the current position, and handle mouse click events to achieve fast positioning. Monitor the mouse double-click event in the main reading area, determine the double-click position and capture the corresponding lesion image data, and save it to the local file system. This module displays images from two lenses at the same time, which is convenient for doctors to compare and observe, improving the accuracy and efficiency of diagnosis; allowing doctors to adjust the reading speed as needed, from slow motion to fast browsing, to meet the reading needs in different scenarios; through the thumbnail preview function, doctors can quickly locate the lesion position, reducing unnecessary browsing time.

[0051] The rapid film reading module simultaneously performs deduplication, abnormal lesion detection, polyp detection and quantitative analysis, reducing the workload of doctors, speeding up film reading and increasing the accuracy of film reading.

[0052] Deduplication is to identify and remove image frames with a similarity greater than 98% (i.e. repeated or highly similar image frames) in consecutive frames, reducing the burden and time of doctors reading images. After the doctor clicks the deduplication button on the page, the visual model is used to extract the features of the image, and the vector list is divided into small batches for calculation; the parallelism of matrix operations is used to improve the calculation efficiency; matrix multiplication or batch calculation of cosine similarity is used, and finally the images with high similarity are filtered out, which greatly reduces the burden and time cost of doctors reading images; compared with traditional image deduplication methods, such as deduplication based on hash algorithms, this deduplication algorithm extracts image features through visual models and uses the parallelism of matrix operations to significantly improve the calculation efficiency.

[0053] Lesion anomaly detection: A lesion anomaly detection model is trained using deep learning methods. The obtained lesion anomaly detection model is used to analyze the texture, color, shape and other feature information in the image, identify possible lesion areas (such as inflammation, ulcers, tumors, etc.), and mark them in the form of highlights or frames. The Vision Transformer (VIT) algorithm model is used to divide the input two-dimensional image into multiple fixed-size image blocks (16x16 pixels), and each image block is flattened into a one-dimensional vector; the one-dimensional vector of each image block is embedded into a fixed-dimensional feature vector space through linear transformation to form an embedded representation of the image block; since the image block sequence is disordered, the Vision Transformer (VIT) algorithm model retains the position information of the image block by adding position encoding to ensure that the model can capture the spatial structure of the image. Publicly available medical imaging datasets are collected, and image data actually used in clinical practice is obtained through cooperative hospitals and medical institutions. They are merged into a training dataset of more than 50,000 images, and the cross entropy loss function and Adam optimizer are used to optimize the parameters and continuously train the model; the embedded image block sequence is input into the Transformer encoder. The encoder contains 12 self-attention layers and feedforward neural network layers. Through the self-attention mechanism, the model can dynamically focus on important areas in the image and capture global and local features. By adding a classification head to the final feature vector, the classification results of the image blocks are output. These classification results are used to determine whether there are lesions in each area of ​​the image. The annotated image results are displayed in real time in the main area to help doctors quickly and accurately identify and diagnose lesions. Compared with some algorithms that only focus on image features without considering position information, VIT retains the position information of image blocks by adding position encoding. This helps the model better understand the spatial structure of the image when identifying lesions and improves the accuracy of detection. VIT's self-attention mechanism also enables the model to dynamically focus on important areas in the image, further improving the accuracy of detection.

[0054] Polyp detection and quantitative analysis, in order to identify the polyp area and draw an accurate bounding box, the image processing method is used to measure the size and shape of the polyp, and the polyp is graded and risk assessed according to the preset evaluation criteria. In the process of reading the film, the detection algorithm YoloV9 is called to adjust the input image to a fixed size, and the image features are extracted using a multi-layer convolutional neural network to generate a feature map; the feature map is divided into an SxS grid, and each grid unit is responsible for detecting the target in a part of the image area; each grid unit predicts 3 bounding boxes, each bounding box contains the position of the target (center coordinates, width and height) and the confidence score (indicating the probability that the bounding box contains the target), and each bounding box also predicts the class probability distribution of the target to determine whether the target belongs to a polyp; YoloV9 uses a joint loss function, including position loss (bounding box regression error), confidence loss (confidence error of predicting that the bounding box contains the target) and classification loss (target category prediction error) to optimize the model parameters; through data enhancement techniques (such as random cropping, rotation, scaling, etc.), the number of training data is increased. In the process of reasoning, the redundant bounding boxes with low confidence and high overlap are removed through the non-maximum suppression algorithm, and the bounding boxes with the highest confidence are retained as the final detection results, which helps doctors diagnose polyps quickly and accurately, and then save the lesion images through the lesion capture function; some traditional two-stage detection algorithms (such as the R-CNN series) need to generate candidate regions first and then perform classification and regression, and the detection speed is slow, while YoloV9 adopts a one-stage detection strategy, which directly predicts the category and position of the target in the output layer. This strategy gives YoloV9 a significant advantage in detection speed and can achieve real-time detection; YoloV9 provides a simple and easy-to-use API interface and pre-trained models to facilitate users to develop and apply target detection tasks, which is especially important for scenarios where rapid diagnosis is required in clinical applications.

[0055] Report editing and output module: After the doctor completes the intelligent film reading, it automatically captures the lesion image and allows the doctor to add explanatory text to generate a detailed PDF report containing patient information, image data, diagnosis results and suggestions. On the film reading page, the doctor can add explanatory text to the captured lesion image and explain the endoscopic findings and diagnosis. Click the Export button to automatically generate a PDF version of the detailed report containing patient information, image data, diagnosis results and suggestions, select the specified save location, and save it. This function automatically obtains patient-related lesion images, image descriptions, and diagnostic opinions, and outputs reports in template format, which greatly facilitates the doctor's work.

[0056] The above is applicable to non-invasive examinations of internal environments such as the digestive tract, and realizes comprehensive optimization from image acquisition, wireless transmission, real-time display to intelligent analysis. After testing, the accuracy of deduplication reached more than 98%, and the accuracy of abnormal lesion detection and polyp detection and quantitative analysis reached more than 95%. Doctors can obtain and analyze images faster and shorten the total time required for diagnosis. The diagnostic process that previously took 2 hours now only takes a few minutes, greatly improving the efficiency of medical diagnosis. With the support of the specific intelligent algorithm of this application, the probability of misdiagnosis or missed diagnosis due to different technical levels of doctors is reduced, thereby improving the reliability and effectiveness of diagnosis.

[0057] Example 2

[0058] A method for reading images with a capsule endoscope that integrates an intelligent reading function, wherein a dual-lens capsule endoscope transmits collected and processed images to a receiver wirelessly, performs deduplication, lesion abnormality detection, polyp detection and quantitative analysis on the image data in the receiver, and performs dual-lens reading by reading images at different frame rates, thumbnail quick preview, thumbnail positioning, forward and backward reading, frame-by-frame forward and backward reading and / or lens mode switching; wherein deduplication is to identify and remove image frames with a similarity greater than 98% in consecutive frames (i.e., repeated or highly similar image frames), thereby reducing the burden and time of doctors reading images; lesion abnormality detection is to use a deep learning method to detect and remove the image frames with a similarity greater than 98% in consecutive frames, thereby reducing the burden and time of doctors reading images; and The lesion abnormality detection model is trained and used to analyze the feature information such as texture, color and shape in the image, identify possible lesion areas (such as inflammation, ulcers, tumors, etc.), and mark them in the form of highlights or frames; polyp detection and quantitative analysis are to identify polyp areas and draw precise bounding boxes, and use image processing methods to measure the size and shape of polyps, and grade and risk assess polyps according to preset evaluation criteria; after the film reading is completed, the lesion image is automatically captured, and the doctor is allowed to add explanatory text to generate a detailed PDF report containing patient information, image data, diagnosis results and suggestions.

[0059] Deduplication: Use a deduplication algorithm to extract the features of the image using a visual model, divide the vector list into small batches for calculation; use the parallelism of matrix operations to improve calculation efficiency; use matrix multiplication or batch calculation of cosine similarity, and finally filter out images with a similarity greater than 98%;

[0060] For lesion abnormality detection, the Vision Transformer (VIT) algorithm model is used to divide the input two-dimensional image into multiple fixed-size image blocks (16x16 pixels), and each image block is flattened into a one-dimensional vector. The one-dimensional vector of each image block is embedded into a fixed-dimensional feature vector space through linear transformation to form an embedded representation of the image block. Since the image block sequence is disordered, Vision The Transformer (VIT) algorithm model retains the position information of the image blocks by adding position encoding to ensure that the model can capture the spatial structure of the image; collect publicly available medical image data sets, and obtain image data actually used in clinical practice through cooperating hospitals and medical institutions, merge them into a training data set of more than 50,000 images, use the cross entropy loss function and Adam optimizer to optimize the parameters, and continuously train the model; then input the embedded image block sequence into the Transformer encoder, which contains 12 layers of self-attention layers (Self-Attention) and feedforward neural network layers. Through the self-attention mechanism, the model can dynamically focus on important areas in the image and capture global and local features; by adding a classification head to the final feature vector, the classification results of the image blocks are output, and these classification results are used to determine whether there are lesions in each area of ​​the image; the annotated image results are displayed in real time in the main area to help doctors quickly and accurately identify and diagnose lesions;

[0061] Polyp detection and quantitative analysis involves calling the detection algorithm YoloV9 during the film reading process, adjusting the input image to a fixed size, using a multi-layer convolutional neural network to extract the features of the image, and generating a feature map; dividing the feature map into an SxS grid, with each grid unit responsible for detecting targets in a certain area of ​​the image; each grid unit predicts three bounding boxes, each of which contains the location of the target (center coordinates, width, and height) and a confidence score (indicating the probability that the bounding box contains the target); each bounding box also predicts the target's category probability distribution to determine whether the target is a polyp; YoloV9 uses A joint loss function is used, including position loss (bounding box regression error), confidence loss (confidence error of the predicted bounding box containing the target) and classification loss (target category prediction error) to optimize model parameters; data enhancement methods (such as random cropping, rotation, scaling, etc.) are used to increase the diversity of training data and improve the generalization ability of the model; in the reasoning process, redundant bounding boxes with low confidence and high overlap are removed by non-maximum suppression algorithm, and the bounding box with the highest confidence is retained as the final detection result, which helps doctors to diagnose polyps quickly and accurately, and then save the lesion image by grabbing the lesion function. The rest are all referred to Example 1. After testing, the accuracy of deduplication reached more than 98%, and the accuracy of lesion abnormality detection and polyp detection and quantitative analysis reached more than 95%. Doctors can obtain and analyze images faster, shorten the total time required for diagnosis, and the diagnosis process that used to take 2 hours now only takes a few minutes, which greatly improves the efficiency of medical diagnosis. With the support of the specific intelligent algorithm of this application, the probability of misdiagnosis or missed diagnosis due to different technical levels of doctors is reduced, thereby improving the reliability and effectiveness of diagnosis.

[0062] The specific implementation of the present invention has been described in detail above. It should be pointed out that the related technical documents can form similar technical solutions through the concept of the present invention through different descriptions, and should be included in the protection scope of the patent of the present invention.

Claims

1. A dual-lens capsule endoscope intelligent film reading system, comprising: Dual-lens capsule endoscope body, receiver module and intelligent film reading module; The main body of the dual-lens capsule endoscope includes an image acquisition and processing module and a wireless transmission module. The image acquisition and processing module is used to acquire and process images, and the wireless transmission module is used to transmit the images processed by the image acquisition and processing module to the receiver module. The receiver module is used for receiving, processing, storing and displaying images; The intelligent film reading module is used to quickly read, analyze and generate reports on the image data stored in the receiver module; Features: The intelligent film reading module includes: Permission management module: used for permission control, only authorized doctors can access and operate the system; Case management module: used to add, delete, modify and query case information, so that doctors can quickly obtain and maintain case data; Fast film reading module: used for deduplication, lesion abnormality detection, polyp detection and quantitative analysis of the image data stored in the receiver module, and dual-lens film reading by reading at different frame rates, quick thumbnail preview, thumbnail positioning, forward and backward film reading, frame-by-frame forward and backward film reading and / or lens mode switching; wherein, deduplication is to identify and remove image frames with a similarity greater than 98% in consecutive frames, thereby reducing the burden and time of film reading for doctors; lesion abnormality detection is to train a lesion abnormality detection model using a deep learning method, and use the obtained lesion abnormality detection model to analyze the texture, color and shape in the image, identify possible lesion areas, and mark them in the form of highlights or frames; polyp detection and quantitative analysis is to identify polyp areas and draw precise boundary boxes, and use image processing methods to measure the size and shape of polyps, and grade and risk assess polyps according to preset evaluation criteria; Report editing and output module: used to automatically capture lesion images after the film reading is completed, and allow doctors to add explanatory text to generate a detailed PDF report containing patient information, image data, diagnosis results and suggestions; Fast film reading module: Use XAML to define the user interface, including dual-lens image display in the main area, lens mode switching button, frame rate control component, thumbnail progress bar, forward and backward buttons, frame-by-frame forward and backward buttons, and lesion capture area; use .NET's data binding mechanism to bind the image data and case information received by the receiver to the XAML interface elements to achieve dynamic display and update of data; implement a frame rate manager to dynamically adjust the image update rate within the range of 1-40 frames per second by listening to events of the frame rate control component; draw thumbnails on the thumbnail progress bar, listen to mouse events to display the thumbnail of the current position, and process mouse click events to achieve fast positioning; listen to the mouse double-click event in the main film reading area, determine the double-click position and capture the corresponding lesion image data, and save it to the local file system; Deduplication: Use a deduplication algorithm to extract the features of the image using a visual model, divide the vector list into small batches for calculation; use the parallelism of matrix operations to improve calculation efficiency; use matrix multiplication or batch calculation of cosine similarity, and finally filter out images with a similarity greater than 98%; For abnormal lesion detection, the Vision Transformer algorithm model is used to divide the input two-dimensional image into multiple fixed-size image blocks, and each image block is flattened into a one-dimensional vector. The one-dimensional vector of each image block is embedded into a fixed-dimensional feature vector space through linear transformation to form an embedded representation of the image block. Since the image block sequence is disordered, the Vision Transformer algorithm model retains the position information of the image block by adding position encoding to ensure that the model can capture the spatial structure of the image. Medical imaging data sets are collected and annotated, and the cross entropy loss function and Adam optimizer are used to optimize the parameters and continuously train the Transformer algorithm model. The embedded image block sequence is then input into the Transformer encoder, which contains 12 self-attention layers and feedforward neural network layers. Through the self-attention mechanism, the Transformer algorithm model dynamically pays attention to important areas in the image and captures global and local features. A classification head is added to the final feature vector to output the classification results of the image block. These classification results are used to determine whether there are lesions in each area of ​​the image. The annotated image results are displayed in real time in the main area to help doctors quickly and accurately identify and diagnose lesions. Polyp detection and quantitative analysis: During the film reading process, the detection algorithm YoloV9 is called to adjust the input image to a fixed size, and a multi-layer convolutional neural network is used to extract the features of the image and generate a feature map. The feature map is divided into an SxS grid, and each grid unit is responsible for detecting targets in a part of the image. Each grid unit predicts three bounding boxes, each of which contains the location and confidence score of the target. Each bounding box also predicts the category probability distribution of the target to determine whether the target is a polyp. YoloV9 uses a joint loss function, including position loss, confidence loss, and classification loss, to optimize model parameters. Through data enhancement methods, the diversity of training data is increased and the generalization ability of the model is improved. During the reasoning process, the non-maximum suppression algorithm is used to remove redundant bounding boxes with low confidence and high overlap, and the bounding box with the highest confidence is retained as the final detection result, helping doctors to quickly and accurately diagnose polyps, and then save the lesion image through the lesion capture function.

2. The dual-lens capsule endoscope intelligent film reading system according to claim 1, characterized in that: The receiver module receives the radio frequency signal from the wireless transmission module and converts it into a usable electrical signal. It then uses an analog-to-digital converter to convert the electrical signal into a digital signal. After error detection and correction and data compression and decompression, the digital signal is restored to the original data. The original data is then reorganized in JPEG format to form complete image data for storage and display.

3. The dual-lens capsule endoscope intelligent film reading system according to claim 1 or 2, characterized in that: Permission management module: The permission management module has a database containing user information. When a user logs in, the MD5 hash algorithm is used to encrypt and store the user's password and private key, and then compare it with the encrypted password stored in the database, and return the verification result.

4. The dual-lens capsule endoscope intelligent film reading system according to claim 1 or 2, characterized in that: Case management module: After entering the system, doctors with operation permissions can query the corresponding case according to the patient's name, age and medical record number in the case, and check the patient's report. According to the actual situation, they can add case information to facilitate subsequent medical management; if the case information is incorrect, it will be modified and deleted; case data is stored using SQLite.

5. A method for reading images using the dual-lens capsule endoscope intelligent image reading system according to any one of claims 1 to 4, characterized in that: The dual-lens capsule endoscope transmits the collected and processed images to the receiver wirelessly, performs deduplication, lesion abnormality detection, polyp detection and quantitative analysis on the image data in the receiver, and performs dual-lens reading through different frame rates, thumbnail quick preview, thumbnail positioning, forward and backward reading, frame-by-frame forward and backward reading and / or lens mode switching; among them, deduplication is to identify and remove image frames with a similarity greater than 98% in consecutive frames, reducing the burden and time of doctors reading films; lesion abnormality detection is to obtain lesion abnormality detection trained by deep learning methods Model, using the obtained lesion abnormality detection model to analyze the texture, color and shape in the image, identify possible lesion areas, and mark them in the form of highlights or frames; polyp detection and quantitative analysis to identify polyp areas and draw precise bounding boxes, while using image processing methods to measure the size and shape of polyps, and grade and assess the risk of polyps according to preset evaluation criteria; after the film reading is completed, the lesion image is automatically captured, and the doctor is allowed to add explanatory text to generate a detailed PDF report containing patient information, image data, diagnosis results and suggestions; Deduplication: Use a deduplication algorithm to extract the features of the image using a visual model, divide the vector list into small batches for calculation; use the parallelism of matrix operations to improve calculation efficiency; use matrix multiplication or batch calculation of cosine similarity, and finally filter out images with a similarity greater than 98%; For abnormal lesion detection, the Vision Transformer algorithm model is used to divide the input two-dimensional image into multiple fixed-size image blocks, and each image block is flattened into a one-dimensional vector. The one-dimensional vector of each image block is embedded into a fixed-dimensional feature vector space through linear transformation to form an embedded representation of the image block. Since the image block sequence is disordered, the Vision Transformer algorithm model retains the position information of the image block by adding position encoding to ensure that the model can capture the spatial structure of the image. Medical imaging data sets are collected and annotated, and the cross entropy loss function and Adam optimizer are used to optimize the parameters, and the Vision Transformer algorithm model is continuously trained. The embedded image block sequence is then input into the Transformer encoder, which contains 12 self-attention layers and feedforward neural network layers. Through the self-attention mechanism, the Transformer algorithm model dynamically pays attention to important areas in the image and captures global and local features. A classification head is added to the last feature vector to output the classification results of the image block. These classification results are used to determine whether there are lesions in each area of ​​the image. The annotated image results are displayed in real time in the main area to help doctors quickly and accurately identify and diagnose lesions. Polyp detection and quantitative analysis: During the film reading process, the detection algorithm YoloV9 is called to adjust the input image to a fixed size, and a multi-layer convolutional neural network is used to extract the features of the image and generate a feature map. The feature map is divided into an SxS grid, and each grid unit is responsible for detecting targets in a part of the image. Each grid unit predicts three bounding boxes, each of which contains the location and confidence score of the target. Each bounding box also predicts the category probability distribution of the target to determine whether the target is a polyp. YoloV9 uses a joint loss function, including position loss, confidence loss, and classification loss, to optimize model parameters. Through data enhancement methods, the diversity of training data is increased and the generalization ability of the model is improved. During the reasoning process, the non-maximum suppression algorithm is used to remove redundant bounding boxes with low confidence and high overlap, and the bounding box with the highest confidence is retained as the final detection result, helping doctors to quickly and accurately diagnose polyps, and then save the lesion image through the lesion capture function.

Citation Information

Patent Citations

  • Capsule endoscopy system

    CN103494595A

  • Contour extraction and detection method and system of thoracic lesion image

    CN115131386A

  • Capsule endoscope image processing method, device and equipment

    CN115937649A

  • Colorectal polyp detection method, device and equipment and storage medium

    CN118333942A