Intelligent control method and system of ear-nose-throat display and storage medium
By using dynamic feature identification models in the ENT endoscopy system for real-time image analysis and update, the problem of low image acquisition and recognition accuracy of traditional endoscopy is solved, and more efficient and accurate diagnostic support is achieved.
Patent Information
- Application Number
- CN202510466163.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional ENT endoscopy has low image acquisition and recognition accuracy during the detection process, which affects the accuracy of diagnosis.
Through the dynamic feature identification model, the endopic image is analyzed in real time and identified with interest, and the image data is updated in a timely manner and transmitted to multiple display devices to improve the accuracy and efficiency of image acquisition and identification.
It improves the accuracy and efficiency of ENT endoscopic image acquisition and recognition, and enhances the accuracy of diagnosis.
Smart Images

Figure CN119993404A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical appliances, and in particular to an intelligent control method, system and storage medium of an ear, nose and throat display. Background Art
[0002] Otolaryngology is one of the important disciplines in the medical field, involving the examination and diagnosis of the ears, nasal cavity, throat and other parts. With the continuous advancement of medical technology, traditional otolaryngology examination methods are gradually developing in the direction of digitalization and intelligence. Existing otolaryngology examination equipment mostly uses endoscopes to observe and diagnose tissues or organs. However, the existing endoscopic technology still has some shortcomings during use. In traditional otolaryngology endoscopy, doctors manually adjust the focus and brightness of the endoscope to obtain clear images, but these operations often rely on the doctor's experience and skills, and are easily affected by factors such as differences in human body structure and changes in the operating environment, resulting in unstable image quality and affecting the accuracy of diagnosis. Summary of the invention
[0003] The present application provides an intelligent control method, system and storage medium for an ENT display, which solves the technical problem of low image acquisition and recognition accuracy during the detection process of traditional endoscopes. Through a dynamic feature identification model, real-time analysis and interest identification of endoscopic images are performed, and image data is updated in a timely manner and transmitted to multiple display devices, achieving the technical effect of improving the accuracy and efficiency of image acquisition and identification.
[0004] In view of the above problems, the present application provides an intelligent control method, system and storage medium for an ENT display.
[0005] According to a first aspect of the present application, there is provided an intelligent control method for an ear, nose and throat display, the method comprising: a multifunctional endoscope transmits a sequence of acquired endoscopic images back to a main display, wherein the multifunctional endoscope integrates ear detection function, nasal cavity detection function and throat detection function; the main display runs a feature identification unit to load a dynamic feature identification model according to the real-time function type of the multifunctional endoscope, wherein the CPU of the main display has an endoscope control unit, a transmission control unit, a display unit and the feature identification unit built in; the main display performs real-time key feature identification on the endoscopic image sequence through the dynamic feature identification model, and when the dynamic feature identification model outputs a real-time risk identification image, the endoscope control unit controls and updates the multifunctional endoscope according to the real-time risk identification image to acquire an updated image sequence; after receiving the updated image sequence, the feature identification unit runs the dynamic feature identification model to perform image feature identification and outputs an updated identification sequence; after interactively obtaining multiple display accuracy information of multiple secondary displays, the transmission control unit transmits the updated identification sequence to the multiple secondary displays according to the multiple display accuracy information.
[0006] According to a second aspect of the present application, an intelligent control system for an ear, nose and throat display is provided, the system comprising: an image return module: a multifunctional endoscope returns a sequence of acquired endoscopic images to a main display, wherein the multifunctional endoscope integrates ear detection function, nasal cavity detection function and throat detection function; a model loading module: the main display runs a feature identification unit to load a dynamic feature identification model according to the real-time function type of the multifunctional endoscope, wherein the CPU of the main display has an endoscope control unit, a transmission control unit, a display unit and the feature identification unit built in; a control update module: the main display updates the dynamic feature identification model through the dynamic feature identification model The dynamic feature identification model performs real-time key feature identification on the endoscopic image sequence. When the dynamic feature identification model outputs a real-time risk identification image, the endoscope control unit controls and updates the multifunctional endoscope according to the real-time risk identification image to acquire an updated image sequence; an image feature identification module: after the feature identification unit receives the updated image sequence, it runs the dynamic feature identification model to perform image feature identification and outputs an updated identification sequence; a sequence transmission module: after the transmission control unit interactively obtains multiple display accuracy information of multiple sub-displays, it transmits the updated identification sequence to the multiple sub-displays according to the multiple display accuracy information.
[0007] The third aspect of the present application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements an intelligent control method for an ENT display provided in the present application.
[0008] One or more technical solutions provided in this application have at least the following technical effects or advantages: First, the multifunctional endoscope transmits the acquired endoscopic image sequence back to the main display, wherein the multifunctional endoscope integrates ear detection function, nasal cavity detection function and laryngeal detection function; then, the main display runs the feature identification unit to load the dynamic feature identification model according to the real-time function type of the multifunctional endoscope, wherein the CPU of the main display is built-in with an endoscope control unit, a transmission control unit, a display unit and the feature identification unit; further, the main display performs real-time key feature identification on the endoscopic image sequence through the dynamic feature identification model, and when the dynamic feature identification model outputs a real-time risk identification image, the endoscope control unit controls and updates the multifunctional endoscope according to the real-time risk identification image to acquire an updated image sequence; thereafter, after receiving the updated image sequence, the feature identification unit runs the dynamic feature identification model to perform image feature identification and outputs an updated identification sequence; finally, after interactively obtaining multiple display accuracy information of multiple secondary displays, the transmission control unit transmits the updated identification sequence to the multiple secondary displays according to the multiple display accuracy information. It solves the technical problem of low image acquisition and recognition accuracy in the detection process of traditional endoscopes. Through the dynamic feature identification model, the endoscopic image is analyzed and the objects of interest are identified in real time, the image data is updated in time and transmitted to multiple display devices, achieving the technical effect of improving the accuracy and efficiency of image acquisition and identification. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0010] Figure 1 A schematic flow chart of an intelligent control method for an ENT display provided in an embodiment of the present application.
[0011] Figure 2 A schematic diagram of a multifunctional endoscope connection for an intelligent control method of an ENT display provided in an embodiment of the present application.
[0012] Figure 3 A schematic diagram of the structure of an intelligent control system for an ENT display provided in an embodiment of the present application.
[0013] Explanation of reference numerals: image return module 11 , model loading module 12 , control update module 13 , image feature identification module 14 , sequence transmission module 15 . DETAILED DESCRIPTION
[0014] The present application solves the technical problem of low image acquisition and recognition accuracy of traditional endoscopes during the detection process by providing an intelligent control method, system and storage medium for an ENT display. It performs real-time analysis and identification of interest on endoscopic images through a dynamic feature identification model, and updates image data in a timely manner and transmits it to multiple display devices, thereby achieving the technical effect of improving the accuracy and efficiency of image acquisition and identification.
[0015] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0016] It should be noted that the terms "including" and "having" are intended to cover non-exclusive inclusions. For example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or modules that are not explicitly listed or are inherent to these processes, methods, products or devices.
[0017] Embodiment 1, as Figure 1 , Figure 2 As shown, the present application provides an intelligent control method for an ENT display, wherein the method comprises: The multifunctional endoscope transmits the acquired endoscopic image sequence back to the main display, wherein the multifunctional endoscope integrates ear detection function, nasal cavity detection function and throat detection function.
[0018] In one embodiment, the multifunctional endoscope captures a sequence of images of the ear, nasal cavity and larynx through a camera or sensor, and transmits the image data back to the main display for display and analysis in real time. The multifunctional endoscope is equipped with an integrated display, which can be quickly connected and detached from the base of the multifunctional endoscope by magnetic attraction. It has a built-in battery and wireless module, and can communicate with the secondary display wirelessly or wired, so as to support the synchronous display of data. The multifunctional endoscope has three main functions, namely ear detection, nasal detection and throat detection. The ear detection function can clearly display the internal structure of the ear and possible problems. The nasal detection function is used to examine the inside of the nasal cavity, including the nasal mucosa, sinuses and their abnormalities, while the throat detection function focuses on the observation of the throat area, including the vocal cords, pharynx and larynx. These functions are integrated in one device, allowing medical staff to flexibly switch between different examination sites to provide comprehensive ENT examination services.
[0019] The main display runs a feature identification unit to load a dynamic feature identification model according to the real-time function type of the multifunctional endoscope, wherein the CPU of the main display has built-in endoscope control unit, transmission control unit, display unit and the feature identification unit.
[0020] In one embodiment, the main display automatically activates the corresponding processing program according to the real-time function type of the multifunctional endoscope (such as ear detection, nasal detection or laryngeal detection), runs the feature identification unit and loads a dynamic feature identification model adapted to the current function. Specifically, the central processing unit (CPU) of the main display integrates multiple core modules, including an endoscope control unit, a transmission control unit, a display unit and a feature identification unit. The endoscope control unit is responsible for real-time control of the endoscope and adjusting parameters such as the focus and brightness of the endoscope to ensure the clarity and stability of the image. The transmission control unit handles data transmission to ensure efficient communication between the main display and the multifunctional endoscope and other remote devices. The display unit is responsible for displaying real-time image data to provide medical staff with clear and intuitive endoscopic images. The feature identification unit loads a dynamic feature identification model to perform real-time analysis of the endoscopic image and identify key features, such as abnormal areas, thereby providing medical staff with auxiliary diagnosis information.
[0021] Furthermore, the main display runs a feature identification unit to load a dynamic feature identification model according to the real-time function type of the multifunctional endoscope, and the method includes: The ear detection function is used as a retrieval condition to collect image data and obtain an ear endoscopy image set; an ear ROI identification rule is predefined; the ear ROI identification rule is used to perform image identification on the ear endoscopy image set to obtain an ear identification image set, wherein the ear identification image set includes a plurality of binary mask images; a model architecture of an ear feature identification model is constructed based on a Mask R-CNN model, and the ear endoscopy image set and the ear identification image set are used as training data to perform model parameter optimization to complete the construction of the ear feature identification model; and by analogy, an ear feature identification model, a nasal feature identification model and a laryngeal feature identification model are constructed; the ear feature identification model, the nasal feature identification model and the laryngeal feature identification model are stored in the feature identification unit; the main display runs the feature identification unit according to the real-time function type of the multifunctional endoscope to filter and load the dynamic feature identification model from the ear feature identification model, the nasal feature identification model and the laryngeal feature identification model.
[0022] Preferably, the system terminal uses the ear detection function of the multifunctional endoscope as a retrieval condition, collects image data from historical detection records, obtains multiple ear endoscopic images (with abnormalities), and adds them to the ear endoscopic image set; then, according to the anatomical structure and diagnostic requirements of the ear, a set of ear ROI (Region of Interest) identification rules are predefined, which are used to guide the automatic identification and marking of key areas in the image, such as the ear canal, eardrum and other characteristic areas; then, the collected ear endoscopic image set is processed using the defined ear ROI identification rules to identify the key areas, and this process is completed through an identification feature extraction model, and a set of ear identification image sets is output, which contain multiple binary mask images for highlighting abnormal parts of the ear; then, based on the Mask R-CNN model (a deep learning model for image segmentation and target detection), an ear feature identification model is constructed, which can accurately identify key areas in ear endoscopic images, such as the inner wall of the ear canal, the eardrum and possible abnormal areas thereof. Specifically, the system terminal uses Mask R-CNN to extract the key areas of the ear endoscopic image. The R-CNN model is used as the basis to build the ear feature identification model, including the input layer, feature extraction layer, region proposal network (RPN), RoI Align layer, segmentation branch and output layer, etc., and then the ear endoscopy image set and ear identification image set are input into the model as training data for forward propagation. The ear endoscopy image is transmitted layer by layer through the input layer, feature extraction layer, RPN, and RoI Align layer to generate the bounding box and mask of each area in the ear image to form the predicted ear identification image. The cross entropy loss function is then used to calculate the loss value between the predicted ear identification image and the real ear identification image, and the gradient of the loss to the weight of each layer is calculated layer by layer through back propagation. Based on these gradients, the Adam optimizer is used to optimize the model and adjust the weights to minimize the value of the loss function. The training process is repeated and the hyperparameters (such as learning rate, batch size, etc.) are adjusted until the model can accurately identify the key features in the ear image, such as the ear canal, tympanum, etc. , and can accurately generate pixel-level segmentation masks for each target area. After the training is completed, the model is tested using data that has not been used for training to evaluate its accuracy and robustness in the ear feature recognition task. If the model accuracy reaches the expected level, it will be stored as the ear feature identification model; after the training and optimization of the ear feature identification model, the same process will be applied to the feature identification of the nasal cavity and larynx to construct the nasal cavity feature identification model and the laryngeal feature identification model. All constructed feature identification models (ear, nasal cavity and larynx) will be stored in the feature identification unit for subsequent use;When medical staff use a multifunctional endoscope for examination, the main display will automatically select and load the corresponding dynamic feature identification model according to the real-time function type of the multifunctional endoscope (ear, nasal cavity or throat detection), which is used to analyze the key features in the current image sequence in real time, such as abnormal areas, to help medical staff make diagnostic decisions. Through this process, efficient and automated image identification of the ear, nose and throat areas is achieved, greatly improving the efficiency and accuracy of medical staff. ;
[0023] Furthermore, the ear ROI identification rule is used to perform image identification on the ear endoscopy image set to obtain an ear identification image set, and the method includes: Extracting the ear endoscopic image set with a data volume of 1 / N as an initial endoscopic image set; performing image identification on the initial endoscopic image set using the ear ROI identification rule to obtain an initial identified image set; constructing an identification feature extraction model based on a CNN network, and using the initial identified image set and the initial endoscopic image set to train the identification feature extraction model, performing ROI identification feature extraction on the identification feature extraction model to obtain ear ROI identification features; performing image identification update on the ear endoscopic image set using the ear ROI identification rule and the ear ROI identification features to obtain the ear identified image set.
[0024] Optionally, the system terminal first extracts 1 / N amount of image data from the entire ear endoscopic image set as the initial endoscopic image set, where 1 / N refers to randomly extracting a small part of the image from the original image set, and the size of N can be determined according to actual needs; then, the initial endoscopic image set is processed using pre-defined ear ROI identification rules. These identification rules usually include a basic understanding of the ear anatomical structure, such as the regional positioning of the ear canal, eardrum and other parts. These rules will help the system terminal automatically identify key areas in the ear image and mark them. The marking results will constitute an initial identified image set, which contains a binary mask image corresponding to the initial endoscopic image set. Each mask image indicates the location of each key area in the image, such as the ear canal and eardrum. Subsequently, a identification feature extraction model is constructed based on a convolutional neural network (CNN). CNN is used to extract local features from the input image and gradually learn the high-level features of the image through multi-layer convolution and pooling operations. The system terminal uses the initial identified image set and the initial endoscopic image set as training images. The data is input into the identification feature extraction model. During the training process, the system terminal uses the same method as above to gradually learn how to extract ear features from the image, such as the morphological features of the ear canal and tympanum, through steps such as forward propagation, loss calculation, back propagation, and parameter optimization, to ensure that the model can accurately identify these areas. After the model training is completed, the system terminal uses the trained identification feature extraction model to infer the ear endoscopy image set, extracts the identification features of each ROI in the image, and constitutes the ear ROI identification features. These identification features represent the location information and image features of the ear area, and can accurately determine the ear canal, tympanum and other areas. Afterwards, the ear ROI identification rules and the extracted ear ROI identification features are used to update the image identification of the entire ear endoscopy image set, that is, the ear ROI identification features are identified to the range corresponding to the ear ROI identification rules in the ear endoscopy image, thereby updating the ear endoscopy image set to the ear endoscopy image set. This ear endoscopy image set contains all correctly marked areas in the ear images, and can provide accurate data support for subsequent ear feature analysis.
[0025] The main display performs real-time key feature identification on the endoscopic image sequence through the dynamic feature identification model. When the dynamic feature identification model outputs a real-time risk identification image, the endoscope control unit controls and updates the multifunctional endoscope according to the real-time risk identification image to acquire an updated image sequence.
[0026] In one embodiment, the main display uses a dynamic feature identification model to process the collected endoscopic image sequence in real time and automatically identify key features in the image. When the dynamic feature identification model identifies potential risks or abnormalities, it generates an ear identification image as a real-time risk identification image to indicate areas in the image that require special attention. At this time, the endoscope control unit automatically adjusts the working parameters of the endoscope, such as focal length or brightness, based on the generated risk identification image, so as to more clearly view and analyze these risk areas. In this way, the endoscope can quickly respond to changes in risk images and update the collected image sequence, thereby providing medical staff with higher quality real-time images and assisting in more accurate diagnosis.
[0027] Furthermore, when the dynamic feature identification model outputs a real-time risk identification image, the endoscope control unit controls and updates the multifunctional endoscope according to the real-time risk identification image to acquire an updated image sequence, and the method includes: The real-time risk identification image is loaded onto the display unit; if the endoscope control unit receives a focus target event transmitted back from the display unit, a visual adjustment sequence is constructed according to the depth feature identification and the focus target event of the real-time risk identification image; the endoscope control unit dynamically regulates the image acquisition process of the multifunctional endoscope based on the visual adjustment sequence as a constraint to obtain the updated image sequence.
[0028] Preferably, after obtaining the real-time risk identification image, the system terminal loads this real-time risk identification image into the display unit of the main display, and the real-time risk identification image will be presented on the display unit to help medical staff quickly identify and focus on important areas. Medical staff view the real-time image on the display unit, and manually click or select a specific area as the target event area (such as a lesion, abnormal area) based on the key area in the real-time risk identification image. The display unit will capture the coordinate information of the target event area selected by the medical staff through the interactive system, such as the upper left and lower right corner coordinates of the rectangular frame, or the center coordinates of the point, and encapsulate these coordinate information as a focused target event and transmit it back to the endoscope control unit. The key to the focused target event is to use the area as the focal area for endoscopic image acquisition to ensure that the lens is aimed at the area for clear imaging; after receiving the focused target event, the endoscope control unit will construct a focused target event based on the depth feature identification in the real-time risk identification image and the position information of the focused target event. A visual adjustment sequence is constructed, in which the depth feature identifier indicates the depth information of each area in the image, helps to determine the position of the focus target in the image and its relative depth, and the focus target event provides the specific position of the target area, ensuring that the adjustment of the endoscope focus can accurately align with the area; then, based on the constructed visual adjustment sequence, the endoscope control unit will dynamically adjust the image acquisition process of the endoscope, that is, according to the focus adjustment and brightness adjustment requirements in the sequence, the working parameters of the endoscope are adjusted in real time, so that the image acquisition process always maintains the clarity and appropriate brightness of the focus target area. Through this dynamic regulation, the endoscope can automatically adjust when acquiring images in real time, making the image of the focus area more accurate, while reducing the problem of blur or overexposure; after dynamic regulation, the endoscope will collect new images to form an updated image sequence, which has higher clarity and focus accuracy, ensuring that medical staff can better observe the target area for further diagnosis and analysis. Through this process, the endoscope can dynamically adjust the focus and brightness according to the operation of medical staff and the real-time analysis of the system, automatically optimize the image quality, and improve the diagnostic efficiency and accuracy of medical staff.
[0029] Furthermore, constructing a visual adjustment sequence according to the depth feature identification and the focus target event of the real-time risk identification image, the method includes: The focused target event is subjected to a two-dimensional coordinate transformation to obtain the endoscopic view angle coordinates; the endoscopic center coordinates of the multifunctional endoscope are interactively obtained, and an offset angle is calculated based on the endoscopic center coordinates and the endoscopic view angle coordinates; an adjustment time step is preset; the adjustment time step is used as an adjustment scale constraint, and an angle adjustment sequence is constructed based on the offset angle; the adjustment time step is used as an adjustment scale constraint, and an output brightness adjustment sequence is matched based on the depth feature identifier; the angle adjustment sequence and the brightness adjustment sequence are associated and spliced to obtain the visual adjustment sequence.
[0030] Optionally, the focus target event is usually selected by the medical staff by touching or clicking a certain area on the main display. The position of the area on the main display is a two-dimensional coordinate (x, y). In order for the endoscope to understand and adjust the viewing angle according to the coordinate, a two-dimensional coordinate conversion is required. The two-dimensional coordinate conversion is to map the coordinates (x, y) on the main display to the actual viewing angle coordinate system of the endoscope. The viewing angle coordinates of the endoscope are a position in the three-dimensional space, which is related to the actual viewing angle, focal length and image acquisition direction of the endoscope. The system terminal accurately maps the focus target on the main display to the endoscope through existing geometric conversion methods (such as homography matrix, perspective projection, etc.). The endoscope view coordinates of the focus target event are determined from the view coordinates to provide an accurate basis for subsequent lens adjustment. Subsequently, the system terminal interacts with the multifunctional endoscope to obtain the endoscope center coordinates of the multifunctional endoscope. The endoscope center coordinates refer to the center position of the endoscope lens or the center point of the field of view, indicating the center position of the current endoscope view. Once the center coordinates of the focus target event and the endoscope are obtained, the system terminal calculates the difference between the endoscope view coordinates and the endoscope center coordinates according to the endoscope center coordinates and the converted endoscope view coordinates to obtain the deviations on the three coordinate axes, and then calculates the ratio of the deviation on the y-axis to the deviation on the x-axis. Perform an inverse tangent calculation to obtain the horizontal offset angle, perform an inverse tangent calculation on the ratio of the deviation on the z-axis and the distance between the two coordinates on the xy plane to obtain the vertical offset angle, and then add the horizontal offset angle and the vertical offset angle to the offset angle to guide the adjustment process of the endoscope lens focal length to bring the target event area into focus; when adjusting the focal length and brightness of the endoscope lens, it is necessary to manage the adjustment process in time, so the system terminal will preset an adjustment time step, which refers to the time interval and adjustment amount corresponding to each adjustment, to ensure that the focal length and brightness of the lens change smoothly and continuously, thereby avoiding image jitter. In the case of unclear or unclear images, the specific size can be set according to actual needs; after obtaining the offset angle and adjustment time step, the system terminal will use the adjustment time step as the adjustment scale constraint, and calculate the ratio of the horizontal offset angle to the adjustment amount of each adjustment time step to determine the number of adjustment time steps required for horizontal adjustment. Similarly, the number of adjustment time steps required for vertical adjustment is calculated based on the vertical offset angle, and then these adjustment time steps are arranged in chronological order to form an angle adjustment sequence. This angle adjustment sequence describes each adjustment step required for the endoscope's focal length to smoothly rotate from the current viewing angle to the target event area in order to achieve precise focus;Similar to the angle adjustment of focal length, the brightness adjustment sequence is generated based on the depth feature identifier of the endoscopic image. The depth feature identifier indicates the relative depth of the target event area, especially when some areas in the image are darker or brighter, the brightness needs to be adjusted. According to the depth feature identifier, the endoscope control unit calculates the final brightness that needs to be adjusted by multiplying the depth information of the depth feature identifier with the brightness adjustment coefficient. This brightness adjustment coefficient is a constant set according to actual needs and equipment performance, and is used to control the sensitivity of brightness changes. The system terminal calculates the difference between the final brightness and the current brightness to obtain the offset brightness, and calculates the required adjustment time steps in the same way as above and combines them in sequence to form a brightness adjustment sequence. This brightness adjustment sequence describes how the brightness should change during the entire image acquisition process of the endoscope lens to ensure that the target event area can be obtained at every moment. Appropriate exposure; after that, the system terminal associates and splices the angle adjustment sequence and the brightness adjustment sequence to generate a comprehensive visual adjustment sequence, which includes every detail of the endoscope's focal angle adjustment and brightness adjustment, ensuring that the endoscope maintains the clarity and appropriate brightness of the target event area while adjusting the focal length, so as to achieve the best observation effect. It should be noted that during the adjustment of the focal angle and brightness, the actual angle of the endoscope will not be changed, nor will the spatial position of the endoscope be changed. The image clarity and field of view are affected only by the focal length adjustment function of the lens to ensure that the target event area is in the focal position; through the above process, accurate focus adjustment and brightness optimization can be achieved, ensuring that the endoscope lens is always focused on the target event area selected by the medical staff, and making necessary brightness adjustments according to the image depth information, so as to provide the best imaging effect for diagnosis. ;
[0031] Furthermore, if the endoscope control unit does not receive the focus target event transmitted back by the display unit, the visual adjustment sequence is constructed according to the depth feature identification of the real-time risk identification image.
[0032] Optionally, under normal circumstances, the endoscope control unit will adjust the focus, brightness and other parameters of the lens according to the focus target event sent back by the display unit. However, if the display unit fails to send back the focus target event, it means that the current focal length meets the target event area that the medical staff needs to pay attention to. At this time, the endoscope control unit will no longer adjust the focal length angle, but will generate a brightness adjustment sequence based on the depth feature identification, and directly use the generated brightness adjustment sequence as the visual adjustment sequence to ensure the stability of the image quality and the clarity of the target event area.
[0033] After receiving the updated image sequence, the feature identification unit runs the dynamic feature identification model to perform image feature identification and outputs an updated identification sequence.
[0034] In one embodiment, after the endoscope is adjusted and optimized, it will re-capture images to generate an updated image sequence, and transmit this updated image sequence to the feature identification unit. After receiving the updated image sequence, the feature identification unit will pre-process the updated image sequence and input the processed image sequence into the dynamic feature identification model for the same image feature identification as mentioned above, generating multiple identification images to form an updated identification sequence. This updated identification sequence contains the abnormal areas marked in the image, helping subsequent medical staff to make a more accurate diagnosis.
[0035] Furthermore, after receiving the updated image sequence, the feature identification unit runs the dynamic feature identification model to perform image feature identification and outputs an updated identification sequence. The method includes: Taking the focused target event as a constraint, the target ROI area and the target ROI features are located in the real-time risk identification image; after receiving the updated image sequence, the feature identification unit pre-processes the updated image sequence with the target ROI area and the target ROI features as constraints to obtain a regional image sequence; the regional image sequence is loaded into the dynamic feature identification model for image feature identification, and an updated identification sequence is output.
[0036] Optionally, the real-time risk identification image indicates the possible risk areas in the image, such as abnormal areas, etc. On this basis, the system terminal uses the focus on the target event as a constraint to locate the target ROI area in the real-time risk identification image, that is, the target event area that the medical staff needs to focus on, and extracts the feature information of the target ROI area, such as shape, texture, color or other important features related to the target area, as the target ROI feature; after the feature identification unit receives the updated image sequence, the system terminal uses the located target ROI area and target ROI features as constraints to preprocess the image. The goal of the preprocessing process is to enhance the visibility of the target ROI area so that subsequent feature recognition is more accurate, including each updated image in the updated image sequence. The new image is cropped to retain only the image data within the target ROI area, and the target ROI features are used to verify the cropped image data to determine whether the cropped image data contains these features, and images that do not contain these features are eliminated, and only cropped images that meet the target ROI features are retained to form a regional image sequence; then, the system terminal inputs the regional image sequence into the dynamic feature identification model for further processing, and the dynamic feature identification model will perform more accurate feature identification on these images, generate identification images corresponding to each image in the regional image sequence, and add these identification images to the updated identification sequence in the order in the regional image sequence, so as to provide support for the diagnosis and analysis of medical staff and improve the efficiency and reliability of medical staff's diagnosis.
[0037] After interactively obtaining a plurality of display precision information of a plurality of secondary displays, the transmission control unit transmits the update identification sequence to the plurality of secondary displays according to the plurality of display precision information.
[0038] In one embodiment, after the system terminal completes the image analysis and generates an update identification sequence, the transmission control unit interacts with multiple sub-displays to obtain the display accuracy information of each sub-display. The display accuracy information includes parameters such as resolution, color display capability, and refresh rate. Based on the acquired accuracy information, the transmission control unit adaptively adjusts the update identification sequence to ensure that the image data transmitted to each sub-display is in line with its hardware capabilities and meets the needs of multi-party collaborative diagnosis or teaching.
[0039] Further, after interactively obtaining multiple display accuracy information of multiple secondary displays, the transmission control unit transmits the update identification sequence to the multiple secondary displays according to the multiple display accuracy information. The method includes: The update identification sequence is subjected to smooth transition processing of image feature changes to obtain an ROI image feature sequence; the update identification sequence is subjected to smooth transition processing of regional feature changes to obtain an ROI regional feature sequence; the ROI image feature sequence and the ROI regional feature sequence are subjected to frame alignment, and then the ROI image feature sequence and the ROI regional feature sequence are fused to obtain a dynamic ROI feature image; and the dynamic ROI feature image is transmitted to the multiple sub-displays after bandwidth optimization is performed according to the multiple display accuracy information.
[0040] Preferably, the update identification sequence includes all regions and feature information in the image that have been feature-identified. However, between consecutive frames of the image sequence, the feature changes may be abrupt or discontinuous, which may affect the smoothness of the image. Therefore, the system terminal will perform smooth transition processing on the regional feature changes on the update identification sequence to eliminate these mutations and ensure that the transition of image features is smoother. For example, interpolation processing is performed on the color, edge, shape, etc. of the feature (taking the average of adjacent data) to ensure smooth transition of feature changes between different frames. After this processing, the generated ROI image feature sequence is an image feature sequence that has been smoothed, ensuring that the feature changes in the target area are smooth and coherent. In addition, the image will also be smoothed. The target event region is continuously smoothed to make the target region in adjacent image frames change more naturally and avoid sudden feature jumps or boundary mutations. For example, the boundary contour of the target event region is smoothly adjusted by Gaussian filtering to ensure that the region contour changes more naturally in consecutive frames and avoid abrupt boundary jumps or discontinuous shape changes. Gaussian filtering is a commonly used image smoothing technology. It achieves image smoothing by calculating the weighted average of each pixel and its surrounding neighborhood pixels. In this way, the edge and color changes of the image are more natural and the discontinuity is reduced. Through this processing, a ROI region feature sequence is generated, which represents the smooth change process of the target event region in consecutive frames. After obtaining the ROI image After aligning the feature sequence and ROI region feature sequence, the system terminal will align the ROI image feature sequence and the ROI region feature sequence frame by frame. This frame-by-frame alignment is to ensure the synchronization of the ROI image feature sequence and the ROI region feature sequence on the time axis so that the two sequences can be correctly fused. After frame alignment, the system terminal will fuse the ROI image feature sequence and the ROI region feature sequence. The fusion process is to combine the image features and the region features to generate a dynamic ROI feature image. This dynamic ROI feature image combines the visual features of the image (such as color, texture, etc.) with the geometric features of the target event area (such as size, shape, etc.) to form a comprehensive image representation. This fusion ensures the image Both the visual information and spatial information of the target event area can be displayed synchronously, thereby providing a more accurate target area analysis; after obtaining the dynamic ROI feature image, the system terminal optimizes the bandwidth according to the display accuracy information of different sub-displays. The display accuracy information refers to the resolution, color depth, refresh rate and other parameters of each sub-display. The bandwidth optimization is to adjust the transmission method of the image data according to the display accuracy of the sub-display. For example, for displays with lower resolution, the image can be compressed or the resolution can be reduced to reduce the data transmission volume and ensure that the image can be displayed smoothly. After completing the bandwidth optimization, the dynamic ROI feature image will be transmitted to multiple sub-displays to ensure that each sub-display presents the most optimized image according to its display capability.
[0041] In summary, the embodiments of the present application have at least the following technical effects: First, the multifunctional endoscope transmits the acquired endoscopic image sequence back to the main display, wherein the multifunctional endoscope integrates ear detection function, nasal cavity detection function and laryngeal detection function; then, the main display runs the feature identification unit to load the dynamic feature identification model according to the real-time function type of the multifunctional endoscope, wherein the CPU of the main display is built-in with an endoscope control unit, a transmission control unit, a display unit and the feature identification unit; further, the main display performs real-time key feature identification on the endoscopic image sequence through the dynamic feature identification model, and when the dynamic feature identification model outputs a real-time risk identification image, the endoscope control unit controls and updates the multifunctional endoscope according to the real-time risk identification image to acquire an updated image sequence; thereafter, after receiving the updated image sequence, the feature identification unit runs the dynamic feature identification model to perform image feature identification and outputs an updated identification sequence; finally, after interactively obtaining multiple display accuracy information of multiple secondary displays, the transmission control unit transmits the updated identification sequence to the multiple secondary displays according to the multiple display accuracy information. It solves the technical problem of low image acquisition and recognition accuracy in the detection process of traditional endoscopes. Through the dynamic feature identification model, the endoscopic image is analyzed and the objects of interest are identified in real time, the image data is updated in time and transmitted to multiple display devices, achieving the technical effect of improving the accuracy and efficiency of image acquisition and identification.
[0042] Embodiment 2, based on the same inventive concept as the intelligent control method of an ear, nose and throat display in the aforementioned embodiment, Figure 3As shown, the present application provides an intelligent control system for an ear, nose and throat display, wherein the system includes: an image return module 11: a multifunctional endoscope returns a sequence of acquired endoscopic images to a main display, wherein the multifunctional endoscope integrates ear detection function, nasal cavity detection function and throat detection function; a model loading module 12: the main display runs a feature identification unit to load a dynamic feature identification model according to the real-time function type of the multifunctional endoscope, wherein the CPU of the main display has an endoscope control unit, a transmission control unit, a display unit and the feature identification unit built in; a control update module 13: the main display updates the dynamic feature identification model through the dynamic feature identification model The dynamic feature identification model performs real-time key feature identification on the endoscopic image sequence. When the dynamic feature identification model outputs a real-time risk identification image, the endoscope control unit controls and updates the multifunctional endoscope according to the real-time risk identification image to acquire an updated image sequence; an image feature identification module 14: after the feature identification unit receives the updated image sequence, it runs the dynamic feature identification model to perform image feature identification and outputs an updated identification sequence; a sequence transmission module 15: after the transmission control unit interactively obtains multiple display accuracy information of multiple sub-displays, it transmits the updated identification sequence to the multiple sub-displays according to the multiple display accuracy information.
[0043] Furthermore, the model loading module 12 is used to execute the following method: The ear detection function is used as a retrieval condition to collect image data and obtain an ear endoscopy image set; an ear ROI identification rule is predefined; the ear ROI identification rule is used to perform image identification on the ear endoscopy image set to obtain an ear identification image set, wherein the ear identification image set includes a plurality of binary mask images; a model architecture of an ear feature identification model is constructed based on a Mask R-CNN model, and the ear endoscopy image set and the ear identification image set are used as training data to perform model parameter optimization to complete the construction of the ear feature identification model; and by analogy, an ear feature identification model, a nasal feature identification model and a laryngeal feature identification model are constructed; the ear feature identification model, the nasal feature identification model and the laryngeal feature identification model are stored in the feature identification unit; the main display runs the feature identification unit according to the real-time function type of the multifunctional endoscope to filter and load the dynamic feature identification model from the ear feature identification model, the nasal feature identification model and the laryngeal feature identification model.
[0044] Furthermore, the model loading module 12 is used to execute the following method: Extracting the ear endoscopic image set with a data volume of 1 / N as an initial endoscopic image set; performing image identification on the initial endoscopic image set using the ear ROI identification rule to obtain an initial identified image set; constructing an identification feature extraction model based on a CNN network, and using the initial identified image set and the initial endoscopic image set to train the identification feature extraction model, performing ROI identification feature extraction on the identification feature extraction model to obtain ear ROI identification features; performing image identification update on the ear endoscopic image set using the ear ROI identification rule and the ear ROI identification features to obtain the ear identified image set.
[0045] Furthermore, the control update module 13 is used to execute the following method: The real-time risk identification image is loaded onto the display unit; if the endoscope control unit receives a focus target event transmitted back from the display unit, a visual adjustment sequence is constructed according to the depth feature identification and the focus target event of the real-time risk identification image; the endoscope control unit dynamically regulates the image acquisition process of the multifunctional endoscope based on the visual adjustment sequence as a constraint to obtain the updated image sequence.
[0046] Furthermore, the control update module 13 is used to execute the following method: The focused target event is subjected to a two-dimensional coordinate transformation to obtain the endoscopic view angle coordinates; the endoscopic center coordinates of the multifunctional endoscope are interactively obtained, and an offset angle is calculated based on the endoscopic center coordinates and the endoscopic view angle coordinates; an adjustment time step is preset; the adjustment time step is used as an adjustment scale constraint, and an angle adjustment sequence is constructed based on the offset angle; the adjustment time step is used as an adjustment scale constraint, and an output brightness adjustment sequence is matched based on the depth feature identifier; the angle adjustment sequence and the brightness adjustment sequence are associated and spliced to obtain the visual adjustment sequence.
[0047] Furthermore, the control update module 13 is used to execute the following method: If the endoscope control unit does not receive the focus target event transmitted back by the display unit, the visual adjustment sequence is constructed according to the depth feature identification of the real-time risk identification image.
[0048] Furthermore, the image feature identification module 14 is used to perform the following method: Taking the focused target event as a constraint, the target ROI area and the target ROI features are located in the real-time risk identification image; after receiving the updated image sequence, the feature identification unit pre-processes the updated image sequence with the target ROI area and the target ROI features as constraints to obtain a regional image sequence; the regional image sequence is loaded into the dynamic feature identification model for image feature identification, and an updated identification sequence is output.
[0049] Furthermore, the sequence transmission module 15 is used to execute the following method: The update identification sequence is subjected to smooth transition processing of image feature changes to obtain an ROI image feature sequence; the update identification sequence is subjected to smooth transition processing of regional feature changes to obtain an ROI regional feature sequence; the ROI image feature sequence and the ROI regional feature sequence are subjected to frame alignment, and then the ROI image feature sequence and the ROI regional feature sequence are fused to obtain a dynamic ROI feature image; and the dynamic ROI feature image is transmitted to the multiple sub-displays after bandwidth optimization is performed according to the multiple display accuracy information.
[0050] Embodiment 3, based on the same inventive concept as the intelligent control method of an ear, nose and throat display in the aforementioned embodiment, the present application provides a storage medium, in which a computer program is stored, and the processor implements the following steps when executing the computer program: the multifunctional endoscope transmits the acquired endoscopic image sequence back to the main display, wherein the multifunctional endoscope integrates ear detection function, nasal cavity detection function and throat detection function; the main display runs the feature identification unit to load the dynamic feature identification model according to the real-time function type of the multifunctional endoscope, wherein the CPU of the main display has built-in endoscope control unit, transmission control unit, display unit and the feature identification model. identification unit; the main display performs real-time key feature identification on the endoscopic image sequence through the dynamic feature identification model, and when the dynamic feature identification model outputs a real-time risk identification image, the endoscope control unit controls and updates the multifunctional endoscope according to the real-time risk identification image to acquire an updated image sequence; after receiving the updated image sequence, the feature identification unit runs the dynamic feature identification model to perform image feature identification and outputs an updated identification sequence; after interactively obtaining multiple display accuracy information of multiple sub-displays, the transmission control unit transmits the updated identification sequence to the multiple sub-displays according to the multiple display accuracy information.
[0051] It should be noted that the above-mentioned sequence of the embodiments of the present application is only for description and does not represent the advantages and disadvantages of the embodiments. And the above-mentioned specific embodiments of this specification are described. The processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0052] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
[0053] This specification and drawings are merely exemplary illustrations of the present application and are deemed to cover any and all modifications, variations, combinations or equivalents within the scope of the present application. Obviously, a person skilled in the art may make various modifications and variations to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the present application and its equivalents, the present application intends to include these modifications and variations.
Claims
1. An intelligent control method for an ENT display, characterized in that: The method comprises: The multifunctional endoscope transmits the acquired endoscopic image sequence back to the main display, wherein the multifunctional endoscope integrates ear detection function, nasal cavity detection function and throat detection function; The main display runs a feature identification unit to load a dynamic feature identification model according to the real-time function type of the multifunctional endoscope, wherein the CPU of the main display has an endoscope control unit, a transmission control unit, a display unit and the feature identification unit built in; The main display performs real-time key feature identification on the endoscopic image sequence through the dynamic feature identification model. When the dynamic feature identification model outputs a real-time risk identification image, the endoscope control unit controls and updates the multifunctional endoscope according to the real-time risk identification image to acquire an updated image sequence; After receiving the updated image sequence, the feature identification unit runs the dynamic feature identification model to perform image feature identification and output an updated identification sequence; After interactively obtaining a plurality of display precision information of a plurality of secondary displays, the transmission control unit transmits the update identification sequence to the plurality of secondary displays according to the plurality of display precision information.
2. The intelligent control method of an ENT display according to claim 1, characterized in that: When the dynamic feature identification model outputs a real-time risk identification image, the endoscope control unit controls and updates the multifunctional endoscope according to the real-time risk identification image to acquire an updated image sequence, and the method includes: loading the real-time risk identification image into the display unit; If the endoscope control unit receives the focus target event transmitted back by the display unit, a visual adjustment sequence is constructed according to the depth feature identification of the real-time risk identification image and the focus target event; The endoscope control unit dynamically regulates the image acquisition process of the multifunctional endoscope based on the visual adjustment sequence to obtain the updated image sequence.
3. The intelligent control method of an ENT display according to claim 2, characterized in that: After receiving the updated image sequence, the feature identification unit runs the dynamic feature identification model to perform image feature identification and outputs an updated identification sequence. The method includes: Using the focused target event as a constraint, locating a target ROI region and a target ROI feature in the real-time risk identification image; After receiving the updated image sequence, the feature identification unit pre-processes the updated image sequence based on the target ROI region and the target ROI feature as constraints to obtain a regional image sequence; The regional image sequence is loaded into the dynamic feature identification model for image feature identification, and an updated identification sequence is output.
4. The intelligent control method of an ENT display according to claim 1, characterized in that: The main display runs a feature identification unit to load a dynamic feature identification model according to the real-time function type of the multifunctional endoscope, and the method includes: Using the ear detection function as a search condition to collect image data, and obtain an ear endoscopy image set; Predefined ear ROI identification rules; Using the ear ROI identification rule to perform image identification on the ear endoscopy image set to obtain an ear identification image set, wherein the ear identification image set includes a plurality of binary mask images; The model architecture of the ear feature identification model is constructed based on the Mask R-CNN model, and the ear endoscopy image set and the ear identification image set are used as training data to perform model parameter optimization to complete the construction of the ear feature identification model; Similarly, an ear feature identification model, a nasal feature identification model, and a laryngeal feature identification model are constructed; storing the ear feature identification model, the nasal feature identification model and the laryngeal feature identification model in the feature identification unit; The main display runs the feature identification unit according to the real-time function type of the multifunctional endoscope to filter and load the dynamic feature identification model from the ear feature identification model, the nasal feature identification model and the laryngeal feature identification model.
5. The intelligent control method of an ENT display according to claim 4, characterized in that: The ear ROI identification rule is used to perform image identification on the ear endoscopy image set to obtain an ear identification image set, and the method includes: Extracting the ear endoscopic image set with a data volume of 1 / N as an initial endoscopic image set; Using the ear ROI identification rule to perform image identification on the initial endoscopic image set to obtain an initial identified image set; After building a marker feature extraction model based on a CNN network and using the initial marker image set and the initial endoscopic image set to train the marker feature extraction model, performing ROI marker feature extraction on the marker feature extraction model to obtain ear ROI marker features; The ear ROI identification rule and the ear ROI identification feature are used to update the image identification of the ear endoscopy image set to obtain the ear identification image set.
6. The intelligent control method of an ENT display according to claim 2, characterized in that: Constructing a visual adjustment sequence according to the depth feature identification and the focus target event of the real-time risk identification image, the method comprises: Performing a two-dimensional coordinate transformation on the focused target event to obtain endoscopy viewing angle coordinates; Interactively obtaining the endoscope center coordinates of the multifunctional endoscope, and calculating the offset angle according to the endoscope center coordinates and the endoscope viewing angle coordinates; Preset adjustment time step; The adjustment time step is used as an adjustment scale constraint, and an angle adjustment sequence is constructed according to the offset angle; The adjustment time step is used as an adjustment scale constraint, and an output brightness adjustment sequence is matched according to the depth feature identifier; The angle adjustment sequence and the brightness adjustment sequence are associated and spliced to obtain the visual adjustment sequence.
7. The intelligent control method of an ENT display according to claim 6, characterized in that: If the endoscope control unit does not receive the focus target event transmitted back by the display unit, the visual adjustment sequence is constructed according to the depth feature identification of the real-time risk identification image.
8. The intelligent control method of an ENT display according to claim 1, characterized in that: After interactively obtaining multiple display accuracy information of multiple secondary displays, the transmission control unit transmits the update identification sequence to the multiple secondary displays according to the multiple display accuracy information. The method includes: Performing smooth transition processing of image feature changes on the update identification sequence to obtain a ROI image feature sequence; Performing smooth transition processing of regional feature changes on the update identification sequence to obtain a ROI regional feature sequence; After performing frame alignment on the ROI image feature sequence and the ROI region feature sequence, the ROI image feature sequence and the ROI region feature sequence are fused to obtain a dynamic ROI feature image; After bandwidth optimization is performed according to the plurality of display accuracy information, the dynamic ROI feature image is transmitted to the plurality of secondary displays.
9. An intelligent control system for an ear, nose and throat display, characterized in that: An intelligent control method for implementing an ENT display according to any one of claims 1 to 8, the system comprising: Image return module: The multifunctional endoscope returns the acquired endoscopic image sequence to the main display, wherein the multifunctional endoscope integrates ear detection function, nasal cavity detection function and throat detection function; Model loading module: the main display runs the feature identification unit to load the dynamic feature identification model according to the real-time function type of the multifunctional endoscope, wherein the CPU of the main display has an endoscope control unit, a transmission control unit, a display unit and the feature identification unit built in; Control update module: the main display performs real-time key feature identification on the endoscopic image sequence through the dynamic feature identification model. When the dynamic feature identification model outputs a real-time risk identification image, the endoscope control unit controls and updates the multifunctional endoscope according to the real-time risk identification image to acquire an updated image sequence; Image feature identification module: after receiving the updated image sequence, the feature identification unit runs the dynamic feature identification model to perform image feature identification and outputs an updated identification sequence; Sequence transmission module: after interactively obtaining multiple display accuracy information of multiple secondary displays, the transmission control unit transmits the update identification sequence to the multiple secondary displays according to the multiple display accuracy information.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, an intelligent control method for an ENT display as described in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
System for dynamically improving medical image acquisition quality
CN102058432A
Endoscope system, endoscope image recognition method and equipment, and storable medium
CN111292318A
Endoscope image processing method and device and storage medium
CN111726506A
Processing system, image processing method, learning method, and processing device
US20230005247A1
System and method for assisting with the diagnosis of otolaryngologic diseases from the analysis of images
US20230274528A1