Intelligent medical display system

The intelligent medical display system addresses image processing speed, stability, and mode switching issues by integrating video acquisition, AI edge computing, and display modules to enhance image quality and surgical accuracy through synchronized diagnostic overlays.

JP2026517957APending Publication Date: 2026-06-02TIANJIN YUJIN ARTIFICIAL INTELLIGENCE MEDICAL TECH CO LTD

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
TIANJIN YUJIN ARTIFICIAL INTELLIGENCE MEDICAL TECH CO LTD
Filing Date
2024-04-29
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Conventional medical display systems suffer from limitations in image processing speed, stability, versatility, image quality degradation, system instability, latency issues, and inability to synchronize and switch between multiple diagnostic modes, which can affect surgical precision and immediacy.

Method used

An intelligent medical display system comprising a video acquisition module, AI edge computing module, marking layer output module, and display module, which processes and overlays lesion target detection, lesion segmentation, and semantic scene detection results onto the original video input, enabling multiple display modes and improved image quality and stability.

Benefits of technology

The system enhances image quality, reduces processing delays, improves surgical accuracy and immediacy, and supports multi-modal and hierarchical diagnosis through synchronized overlay and real-time switching of diagnostic data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026517957000001_ABST
    Figure 2026517957000001_ABST
Patent Text Reader

Abstract

The present invention relates to an intelligent medical display system, which includes a video acquisition module, an AI edge computing module, a marking layer output module, a video overlay module, and a display module. The video acquisition module receives a video input signal and transmits it to the AI ​​edge computing module. The AI ​​edge computing module analyzes the video input signal to acquire lesion target detection messages, lesion segmentation messages, and semantic scene detection messages, and pushes them to the marking layer output module. The marking layer output module performs group marking on the pushed messages, and based on the group marking results, marks the background image to form lesion target detection result videos, lesion segmentation result videos, and semantic scene detection result videos, which are output to the video overlay module. The video overlay module is used to perform video overlay, and the display module is used to display the overlaid video. It can improve image quality and system stability, reduce image processing delay, and improve immediacy and accuracy during surgery.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Cross-reference to Related Applications) This application claims priority to a Chinese patent application filed with the Chinese Patent Office on May 12, 2023, with an application number of 202310536781.5 and an invention title of "Intelligent Medical Display System", and all of its contents are incorporated herein by reference.

[0002] The present invention belongs to the technical field of medical devices, relates to a medical display system, and specifically relates to an intelligent medical display system.

Background Art

[0003] With the rapid development of medical imaging technology and information technology, more and more artificial intelligence-assisted diagnostic devices are emerging in the market. Most of them directly output the video signal of the medical imaging device to an AI edge computing device, and then perform image analysis on an independent device and display the results. However, the prior art has certain limitations in terms of image processing speed, stability, and versatility.

[0004] Specifically, in the prior art, an independent AI edge computing device is adopted to analyze the video signal in real time, and the analyzed image is output to a medical display or output to an independent display screen and arranged side by side with a display for displaying the original image. Such a solution can achieve real-time image analysis, but still has the following defects.

[0005] 1. Single display function: The conventional display can only display the original video signal or the signal after AI edge processing, and cannot realize the switching of various display modes.

[0006] 2. Image quality degradation: The image processed by the AI edge computing device may cause degradation of the quality of the original image.

[0007] 3. System stability and latency issues: Due to the influence of the operating system and AI processing, the system may experience instability and latency issues, which could affect the surgical process.

[0008] 4. Because it fails to immediately capture the physician's attention, the presented information may be overlooked or not effectively received. Furthermore, it is not possible to achieve synchronized overlay and real-time selection / switching of multimodal presented data processed by multiple algorithms.

[0009] Therefore, it is necessary to research and develop a novel intelligent medical display system to address the shortcomings of the conventional technology described above. [Overview of the Initiative] [Problems that the invention aims to solve]

[0010] To overcome the shortcomings of conventional technology, the present invention proposes an intelligent medical display system that can improve image quality and system stability, reduce image processing delays, and enhance immediacy and accuracy during surgery. [Means for solving the problem]

[0011] To achieve the above objectives, the present invention provides the following technical solutions.

[0012] An intelligent medical display system including a video acquisition module, an AI edge computing module, a marking layer output module, a video overlay module, and a display module, The video acquisition module is used to receive video input signals from a medical imaging device and transmit them to the AI ​​edge computing module. The AI ​​edge computing module is used to analyze the video input signal to acquire lesion target detection messages, lesion segmentation messages, and semantic scene detection messages, and to push them to the marking layer output module. The marking layer output module includes an information processing submodule and an image push delivery submodule. The information processing submodule is used to group and mark messages pushed by the AI ​​edge computing module, dividing them into lesion target detection message groups, lesion segmentation message groups, and semantic scene detection message groups. The image push delivery submodule is used to mark background images based on the grouping marking results of the information processing submodule, to form lesion target detection result videos, lesion segmentation result videos, and semantic scene detection result videos, and to output them to the video overlay module. The video overlay module is used to overlay the lesion target detection result video, lesion segmentation result video, and semantic scene detection result video from the image push delivery submodule, and the video input signal from the video acquisition module, and to output the overlaid video to the display module. The aforementioned display module is used to display overlaid video, and is an intelligent medical display system.

[0013] Preferably, the information processing submodule includes a first buffer, and a lesion target detection buffer, a lesion segmentation buffer, and a semantic scene detection buffer connected to the first buffer, wherein the first buffer is connected to the AI ​​edge computing module, receives and temporarily stores messages pushed by the AI ​​edge computing module, analyzes the message header of the message, temporarily stores the lesion target detection message in the lesion target detection buffer, temporarily stores the lesion segmentation message in the lesion segmentation buffer, and temporarily stores the semantic scene detection message in the semantic scene detection buffer, thereby dividing them into a lesion target detection message group, a lesion segmentation message group, and a semantic scene detection message group.

[0014] Preferably, the image push delivery submodule includes a texture channel creation unit, a texture channel context creation unit, a background image drawing unit, a graphics drawing unit, and a video output unit. The texture channel creation unit is used to create a lesion target detection texture channel, a lesion segmentation texture channel, and a semantic scene detection texture channel, and to bind them to three output ports. The texture channel context creation unit is used to create a lesion target detection texture channel context, a lesion segmentation texture channel context, and a semantic scene detection texture channel context for the three output ports, respectively. The background image rendering unit is used to render background images in the lesion target detection texture channel, lesion segmentation texture channel, and semantic scene detection texture channel, respectively. The graphics drawing unit draws a lesion target graphic on the background image in the lesion target detection texture channel based on the lesion target detection message group, draws a lesion division graphic on the background image in the lesion division texture channel based on the lesion division message group, and draws a semantic scene graphic on the background image in the semantic scene detection texture channel based on the semantic scene detection message group. The video output unit is used to form a lesion target detection result video from a background image and a lesion target graphic drawn thereon, to form a lesion segmentation result video from a background image and a lesion segmentation graphic drawn thereon, and to form a semantic scene detection result video from a background image and a semantic scene detection graphic drawn thereon, and to output each of these to the video overlay module via the three output ports.

[0015] Preferably, the video overlay module includes four input terminals and one output terminal, the four input terminals being used to receive lesion target detection result videos, lesion segmentation result videos, and semantic scene detection result videos from the video output unit, and video input signals from the video acquisition module, respectively, and the output terminal being used to output the overlaid video to the display module.

[0016] Preferably, the video overlay module is 1) A step of establishing and initializing a video input signal ring buffer, a lesion detection result video ring buffer, a lesion segmentation result video ring buffer, and a semantic scene detection result video ring buffer, and simultaneously setting the capacity of each of the ring buffers, 2) Create intercept threads for the input channels of the four input terminals, so that when each input terminal receives video input, it can write frame data to the corresponding ring buffer, and when the capacity of the ring buffer is full, pop up header frame data while inserting new frame data, 3) Creating a shifted output, that is, when the video input signal ring buffer first reaches a predetermined storage length, the data of the first frame is output directly without any processing, and the data of the first frame in the video input signal ring buffer is shifted by one frame in time with the data of the first frame in the other three ring buffers. 4) The process of outputting the AI ​​screen after overlay compositing, This performs a video overlay.

[0017] Preferably, step 4) specifically means, 4.1) Using the output time of the first frame data of the video input signal occupied by the shifted output, a graphic overlay is performed on the first frame data of the headers of the lesion detection result video ring buffer, lesion segmentation result video ring buffer, and semantic scene detection result video ring buffer to obtain a synthesized AI graphic presentation signal. 4.2) The method includes overlaying the frame data of the synthesized AI graphic presentation signal and the video input signal to finally obtain the AI ​​screen after the overlay synthesis.

[0018] Preferably, the AI ​​edge computing module includes a lesion target detection submodule, a lesion segmentation submodule, and a semantic scene detection submodule, wherein the lesion target detection submodule is used to recognize the video input signal frame by frame, obtain lesion detection object targets, and convert the lesion detection object targets into messages to obtain the lesion target detection messages; the lesion segmentation submodule is used to recognize the video input signal frame by frame, obtain lesion segmentation object targets, and convert the lesion segmentation object targets into messages to obtain the lesion segmentation messages; and the semantic scene detection submodule is used to recognize the video input signal frame by frame, obtain semantic scene detection results, and convert the semantic scene detection results into messages to obtain the semantic scene detection messages.

[0019] Preferably, the lesion target detection message includes a message header and a message body, where the message header is DET and the message body is the coordinates of a plurality of lesion detection object targets, with the plurality of coordinates separated by the "|" symbol; the lesion segmentation message includes a message header and a message body, where the message header is SEG and the message body is a set of point coordinates of a plurality of lesion segmentation object targets, with the set of point coordinates of multiple groups separated by the "|" symbol; and the semantic scene detection message includes a message header and a message body, where the message header is TXT and the message body is presentation characters corresponding to a plurality of semantic scenes, with the presentation characters of multiple groups separated by the "|" symbol.

[0020] Preferably, the display module has an image input source switching key, and the video collection module, the AI edge computing module, and the video overlay module are all connected to the display module. The image input source switching key is used to determine whether the display content by the display module is from the video collection module, the AI edge computing module, or the video overlay module.

Advantages of the Invention

[0021] Compared with the prior art, the intelligent medical display system of the present invention has one or more of the following beneficial technical effects.

[0022] 1. The present invention provides three display modes, thereby having a multifunctional effect and providing more convenience to medical staff.

[0023] 2. The present invention supports multi-modal diagnosis, hierarchical diagnosis, radiomics joint application of disease diagnosis, and clinical technology practice through AI prompt signals, which is helpful for improving the quality of medical operations of doctors.

[0024] 3. The present invention realizes the overlay of the original video input signal and the marked video output from the AI edge computing module through the video overlay technology, improving the image quality and system stability.

[0025] 4. The present invention reduces the image processing delay and improves the immediacy and accuracy during surgery.

Brief Description of the Drawings

[0026] [Figure 1] It is a configuration schematic diagram of the intelligent medical display system of the present invention. [Figure 2] It is a configuration and operation flowchart of the marking layer output module of the present invention. [Figure 3]This is a schematic diagram of the configuration of the image push delivery submodule of the present invention. [Figure 4] This is a flowchart illustrating how the video overlay module of the present invention performs video overlay. [Modes for carrying out the invention]

[0027] The present invention will be further described below in combination with the drawings and embodiments, and the contents of the embodiments are not intended to limit the scope of protection of the present invention.

[0028] In response to the shortcomings of conventional medical display systems, the present invention provides an intelligent medical display system that can improve image quality and system stability, reduce image processing delays, and enhance immediacy and accuracy during surgery.

[0029] Figure 1 shows a schematic diagram of the configuration of the intelligent medical display system of the present invention. As shown in Figure 1, the intelligent medical display system of the present invention includes a video acquisition module, an AI edge computing module, a marking layer output module, a video overlay module, and a display module.

[0030] Here, the video acquisition module is used to receive video input signals from the medical imaging device and transmit them to the AI ​​edge computing module.

[0031] Specifically, the video acquisition module can receive video input signals from various medical imaging devices, such as endoscopes, ultrasound machines, or laparoscopes. Furthermore, the video acquisition module can transmit the received video input signals to the AI ​​edge computing module via a high-speed PCIe interface.

[0032] The AI ​​edge computing module is used to analyze the video input signal to obtain lesion target detection messages, lesion segmentation messages, and semantic scene detection messages, and to push them to the marking layer output module.

[0033] Specifically, the AI ​​edge computing module includes a lesion target detection submodule, a lesion segmentation submodule, and a semantic scene detection submodule.

[0034] Here, the lesion target detection submodule is configured based on target detection algorithms such as yolov5, yolov7, and faster-rcnn. The construction process for the lesion target detection submodule is as follows.

[0035] 1. Collect medical images containing lesions, mark the medical images with rectangular frames, and obtain a lesion target detection dataset.

[0036] 2. Construct a neural network for detecting lesion targets.

[0037] In this invention, a lesion target detection neural network is constructed using the yolov5 target detection network, which includes three groups of close-stage residual network structures for feature extraction: one group is a feature pyramid scaling structure for feature extraction and multiscale feature map generation; one group is a feature pyramid aggregation structure for feature extraction and multiscale feature map fusion; and the remaining group is a target detection head structure for estimating the position of framing the target.

[0038] 3. The lesion target detection dataset is input to the lesion target detection neural network, and after backpropagation training, a lesion target detection algorithm, i.e., the lesion target detection submodule, is obtained.

[0039] During use, the lesion target detection submodule is used to recognize the video input signal frame by frame, obtain a lesion detection object target, and convert the lesion detection object target into a message to obtain the lesion target detection message.

[0040] The lesion target detection message includes a message header and a message body. Here, the message header is DET, and the message body is the coordinates of multiple lesion detection object targets, separated by the "|" symbol.

[0041] In the present invention, since the detection target is leveled by a rectangular frame, the lesion detection object target is a rectangular frame, the message body is the coordinates of the top-left and bottom-right vertices of the rectangular frame, and the coordinates of the two vertices are separated by the "|" symbol.

[0042] Finally, the lesion target detection submodule pushes the lesion target detection message to the marking layer output module using socket technology.

[0043] The lesion segmentation submodule can be constructed based on image segmentation algorithms such as UNet, FCN, and Deeplab. The construction process for the lesion segmentation submodule is as follows.

[0044] 1. Collect medical images containing lesions, perform contour marking on the medical images, and obtain a lesion segmentation dataset.

[0045] 2. Construct a lesion-segmented neural network.

[0046] The present invention employs a UNet partitioning network to construct a lesion partitioning neural network, which includes a set of encoder modules composed of four groups of convolutional blocks with decreasing scale for extracting features and outputting four sets of feature maps, and a set of decoder structures composed of four groups of convolutional blocks with increasing scale for feature fusion and reconstruction of the partitioning results.

[0047] 3. The lesion segmentation dataset is input to the lesion segmentation neural network, and after backpropagation training, a lesion segmentation algorithm, i.e., the lesion segmentation submodule, is obtained.

[0048] During use, the lesion segmentation submodule recognizes the video input signal frame by frame, obtains a lesion segmentation object target, converts the lesion segmentation object target into a message, and obtains the lesion segmentation message.

[0049] The lesion segmentation message includes a message header and a message body, the message header being a SEG, and the message body being a set of point coordinates for multiple lesion segmentation object targets, with multiple groups of point coordinate sets separated by the "|" symbol.

[0050] In the present invention, since the lesion is divided using a polygon, the lesion division object target is a polygon, the message body is the coordinates of each vertex of the polygon, and the coordinates of each vertex are separated by the "|" symbol.

[0051] Finally, the lesion segmentation submodule pushes the lesion segmentation message to the marking layer output module using socket technology.

[0052] The aforementioned semantic scene detection submodule can be constructed based on semantic scene detection algorithms such as TSN, SlowFast, and X3D. The construction process for the aforementioned semantic scene detection submodule is as follows.

[0053] 1. Medical video segments are collected, labeled and marked, and a semantic scene dataset is obtained. In this embodiment, the labeling and marking would be such that, for a fecal water segment obtained using an endoscope, it would be labeled "Please be careful when lavaging the bowel," and for an ulcer lesion observation segment, it would be labeled "Turn on the staining light and observe. A biopsy can be performed by taking a sample from the inside of the ulcer rim." It should be understood that the above is merely an illustrative explanation and includes, but is not limited to, the above segments and levels.

[0054] 2. Construct a semantic scene detection neural network.

[0055] The present invention constructs a semantic scene detection neural network employing a TSM semantic scene detection network, and the semantic scene detection neural network includes a convolutional neural network with 16 groups of residual structures for receiving an input video segment and extracting features frame by frame to obtain a feature map for each frame; a group of spatiotemporal fusion modules for performing shift fusion on the feature maps obtained for each frame, that is, shifting the first 1 / 3 of the feature map for each frame one frame backward and shifting the last 1 / 3 of the feature map for each frame one frame forward, thereby creating a shift effect on the feature map for each frame, and finally merging the spatiotemporal feature map; and a fully connected layer module for performing semantic category tag prediction on the fused spatiotemporal feature map.

[0056] 3. The semantic scene dataset is input to the semantic scene detection neural network, and after backpropagation training, a semantic scene detection algorithm, i.e., the semantic scene detection submodule, is obtained.

[0057] During use, the semantic scene detection submodule recognizes the video input signal frame by frame, obtains a semantic scene detection result, converts the semantic scene detection result into a message, and obtains the semantic scene detection message.

[0058] The aforementioned semantic scene message includes a message header and a message body, the message header being TXT, and the message body consisting of presentation characters corresponding to multiple semantic scenes, with multiple groups of presentation characters separated by the "|" symbol.

[0059] Finally, the semantic scene detection submodule pushes the semantic scene detection message to the marking layer output module using Socket technology.

[0060] The marking layer output module includes an information processing submodule and an image push delivery submodule.

[0061] Here, the information processing submodule is used to group and mark the messages pushed by the AI ​​edge computing module, dividing them into lesion target detection message groups, lesion segmentation message groups, and semantic scene detection message groups. The image push delivery submodule is used to mark the background image based on the grouping marking results of the information processing submodule, to form lesion target detection result videos, lesion segmentation result videos, and semantic scene detection result videos, and to output them to the video overlay module.

[0062] Specifically, as shown in Figure 2, the information processing submodule includes a first buffer, and a lesion target detection buffer, a lesion segmentation buffer, and a semantic scene detection buffer connected to the first buffer.

[0063] Here, the first buffer is connected to the AI ​​edge computing module, receives and temporarily stores messages pushed by the AI ​​edge computing module, analyzes the message header of the message, temporarily stores the lesion target detection message in the lesion target detection buffer based on the message header, temporarily stores the lesion segmentation message in the lesion segmentation buffer, and temporarily stores the semantic scene detection message in the semantic scene detection buffer, thereby dividing the messages into a lesion target detection message group, a lesion segmentation message group, and a semantic scene detection message group.

[0064] In other words, if the message header of a message pushed by the AI ​​edge computing module is DET, the first buffer determines that the message is a lesion target detection message and stores it in the lesion target detection buffer; if the message header of a message pushed by the AI ​​edge computing module is SEG, the first buffer determines that the message is a lesion segmentation message and stores it in the lesion segmentation buffer; if the message header of a message pushed by the AI ​​edge computing module is TXT, the first buffer determines that the message is a semantic scene detection message and stores it in the semantic scene detection buffer. As a result, the lesion target detection buffer is temporarily stored with multiple lesion target detection messages, i.e., lesion target detection message groups; the lesion segmentation buffer is temporarily stored with multiple lesion segmentation messages, i.e., lesion segmentation message groups; and the semantic scene detection buffer is temporarily stored with multiple semantic scene detection messages, i.e., semantic scene detection message groups. Thus, group marking for pushed messages is realized.

[0065] In the present invention, the effective period of the lesion target detection buffer, lesion segmentation buffer, and semantic scene detection buffer is set to 16.66 ms. It should be noted that 16.66 ms is determined by the access frequency of the image push delivery submodule. In the present invention, the image push delivery submodule employs an access frequency of 60 frames / second, so 1000 / 60 ≈ 16.66. In actual use, the effective period is not limited to this and may vary depending on the access frequency of the image push delivery submodule.

[0066] Simultaneously, upon reception, the first buffer can receive messages pushed by the AI ​​edge computing module via the Socket Client.

[0067] As shown in Figure 3, in the present invention, the image push delivery submodule includes a texture channel creation unit, a texture channel context creation unit, a background image drawing unit, a graphics drawing unit, and a video output unit.

[0068] Here, the texture channel creation unit is used to create a lesion target detection texture channel, a lesion segmentation texture channel, and a semantic scene detection texture channel, and to bind them to three output ports.

[0069] In the present invention, the texture channel creation unit can create the lesion target detection texture channel, lesion segmentation texture channel, and semantic scene detection texture channel using GLFW software. Specifically, the glfwGetMonitorPos function can be used to select a video output port, and the glfwGetMonitorPos function supports the incoming video transmission number, for example, binding the lesion target detection texture channel to output port 1, i.e., receiving getMonitors()[1]. Similarly thereafter, the lesion segmentation texture channel can be bound to output port 2, and the semantic scene detection texture channel can be bound to output port 3.

[0070] The texture channel context creation unit is used to create a lesion target detection texture channel context, a lesion segmentation texture channel context, and a semantic scene detection texture channel context for the three output ports, respectively.

[0071] In this invention, the glfwMakeContextCurrent function in the GLFW software is used to bind the three output ports to an OpenGL context, thereby obtaining the lesion target detection texture channel context, the lesion segmentation texture channel context, and the semantic scene detection texture channel context.

[0072] The background image rendering unit is used to render background images in the lesion target detection texture channel, the lesion segmentation texture channel, and the semantic scene detection texture channel, respectively.

[0073] In this invention, to cycle-render 60 frames per second, the clear value is set using the glClearColor function when initializing the frame buffer for each frame. This function receives four parameters corresponding to the red, green, blue, and alpha (transparency) components, respectively. In this invention, the background is set to white by default, meaning the incoming parameters are 1.0, 1.0, 1.0, and 1.0. In other words, when using the glClearColor function and the incoming parameters are 1.0, 1.0, 1.0, and 1.0, a white background image can be drawn within each texture channel. In this invention, when a white background image is used and a video overlay is performed, the background image does not affect the original video input signal.

[0074] The graphic drawing unit draws a lesion target graphic on the background image in the lesion target detection texture channel based on the lesion target detection message group, draws a lesion division graphic on the background image in the lesion division texture channel based on the lesion division message group, and draws a semantic scene graphic on the background image in the semantic scene detection texture channel based on the semantic scene detection message group.

[0075] Since the lesion target graphic is rectangular and the lesion segmentation graphic is polygonal, and their drawing principles and methods are the same, they will be described together below.

[0076] First, the lesion target detection texture channel reads the lesion target detection message group from the lesion target detection buffer frame by frame, and the lesion division texture channel reads the lesion division message group from the lesion division buffer frame by frame, obtaining the coordinates of each vertex in each message within the corresponding message group, and converting the vertex coordinates into GPU-interpretable vertex buffer objects using the glBufferData function. Next, the glBindBuffer method is called to bind the GPU-interpretable vertex buffer object to the corresponding lesion target detection texture channel context or lesion division texture channel context. Finally, the glDrawArrays method is called to complete the graphic drawing. This method obtains the vertex data from the corresponding context and draws the corresponding graphic onto the background image of the current texture channel.

[0077] Specifically, if the lesion target graphic is a rectangular frame, the format of the vertex coordinates is [(x0min,y0min),(x0max,y0max)],[(x1min,y1min),(x1max,y1max)]...[(xn'min,yn'min),(xn'max,yn'max)]. In the formula, each "[ ]" represents a single graphic object, xmin and ymin represent the coordinates of the upper left corner of each graphic, and xmax and ymax represent the coordinates of the lower right corner of the graphic.

[0078] If the lesion segmentation graphic is a polygon, the format of the coordinate vertices is [(x0,y0),(x1,y1),...,(xn,yn)],...,[(x0,y0),(x1,y1),...,(xn,yn)]. In the formula, each "[ ]" represents a single graphic object and stores the coordinates of each vertex in the graphic.

[0079] Since semantic scene graphics are character graphics, here we will describe separately how to draw character graphics in the semantic scene detection texture channel. For the sake of explanation, the present invention will be described with an example, and assuming that the message body of the semantic scene detection message in the semantic scene detection buffer is "Please be careful when washing your bowels", the specific steps for drawing the corresponding character graphic are as follows.

[0080] First, the system traverses through each character one by one, and based on the character value, it searches for the corresponding texture object in the character texture set.

[0081] Next, the character positions are calculated. Assuming that the character display position of the present invention is in the middle of the bottom of a 1920*1080 screen, the font size is 14 pixels, and there are 7 exemplary characters, the vertex coordinates of the first character are (896,1060),(910,1060). Similarly, assuming that the character spacing is 2 pixels, the vertex coordinates of the second character are (912,1054),(926,1054), and the remaining character calculations are based on these.

[0082] Next, the vertex coordinates of each character are input one by one, and the glBufferData function is used to convert the vertex coordinates into GPU-interpretable vertex buffer objects. The glBindBuffer method is then called to bind the GPU-interpretable vertex buffer objects to the corresponding semantic scene detection texture channel context.

[0083] Finally, the glDrawArrays method is called to complete the character graphic drawing. This method retrieves vertex data from the corresponding context and draws the corresponding texture graphic content onto the background image of the current texture channel.

[0084] Here, it should be understood that the aforementioned character texture set exists as fixed content, that is, it is executed each time the program is initialized, and its specific creation process includes the following:

[0085] First, the FreeType software library is used to load the font file and call FT_Set_Pixel_Sizes to set the desired font and pixel sizes (the fonts belong to general technical characteristics, and regarding the font size, the present invention calculates the font size using the following formula: 10.5 pounds / 72 * 96 dpi = 14 pixels, i.e., one pound is equivalent to 1 / 72 of an inch, and the present invention assumes display on a 96 dpi display).

[0086] Next, a character texture set is created. Specifically, the graphics library used in this invention is OpenGL, and each character in the font file is traversed, the glBindTexture function is called to obtain a texture object, FreeType is used to load the font for each character, and it is converted into bitmap information, the bitmap is copied to the texture object, and the bitmap contains the texture content, coordinates, dimensions, baseline position and character width for each character.

[0087] Finally, a mapping set is established, and in this invention, the data type is stored using a key value, where the key is a character and the value is a texture object.

[0088] The video output unit is used to form a lesion target detection result video from a background image and a lesion target graphic drawn thereon, to form a lesion segmentation result video from a background image and a lesion segmentation graphic drawn thereon, and to form a semantic scene detection result video from a background image and a semantic scene detection graphic drawn thereon, and to output each of these to the video overlay module via the three output ports.

[0089] Specifically, the frame buffer refresh function can be called to produce a continuous output of the corresponding graphics, thereby forming the video output.

[0090] The video overlay module is used to overlay the lesion target detection result video, lesion segmentation result video, and semantic scene detection result video from the image push delivery submodule, as well as the video input signal from the video acquisition module, and to output the overlaid video to the display module.

[0091] Specifically, the video overlay module includes four input terminals and one output terminal. The four input terminals are used to receive lesion target detection result videos, lesion segmentation result videos, and semantic scene detection result videos from the video output unit, and video input signals from the video acquisition module, respectively. The video overlay module performs an overlay on the received video. The output terminal is used to output the overlaid video to the display module.

[0092] In the present invention, the video overlay module may be an FPGA video overlay module. When the FPGA video overlay module performs an overlay on the received video, the overlay order from the bottom layer to the top layer is the video input signal from the video acquisition module, i.e., original video input screen → lesion target detection result video → lesion segmentation result video → semantic scene detection result video, and finally an AI screen after overlay synthesis is obtained.

[0093] As shown in Figure 4, the specific steps for performing a video overlay on the FPGA video overlay module are as follows.

[0094] 1. Establish and initialize the video input signal ring buffer, the lesion detection result video ring buffer, the lesion segmentation result ring buffer, and the semantic scene detection result video ring buffer, and simultaneously set the capacity of each of the ring buffers.

[0095] Each of the aforementioned ring buffers must implement memory chip interface logic. To connect the DDR3 memory chip to the external memory interface of the FPGA video overlay module and generate the memory controller and interface, the FPGA video overlay module uses a DDR3 controller IP core and a Memory Interface Generator (MIG) tool. The MIG tool can generate the memory controller and interface based on the memory chip model number and parameters. Regarding the capacity of the ring buffer, in this invention, assuming that the received signal data is 1920*1080 pixels, the RGB graphics (occupying 3 bytes) are 24 bits, and 3 frames are stored by default, the capacity of the ring buffer is as follows.

[0096] 1920 * 1080 * 3 (bytes) * 3 (frames) = 18,662,400 bytes.

[0097] 2. Create interception threads for 4 sets of channels, that is, interception threads for the input channels of the 4 input terminals. In this way, when each input terminal receives video input, it can intercept and write the respective frame data to the corresponding ring buffer. When the capacity of the ring buffer is full, new frame data is inserted and the header frame data is popped up at the same time.

[0098] 3. Create a shifted output. By creating a shifted output, it is possible to guarantee low delay output of the original image and obtain stable marking information.

[0099] The process for creating the shifted output described above is as follows:

[0100] First, when the video input signal ring buffer reaches a predetermined storage length for the first time, the data of the first frame within it is output directly without any processing.

[0101] In this invention, an HDMI 2.0 output interface can be realized using the IP core of a Xilinx HDMI 2.0 interface. Specifically, the IP core of the DDR memory controller reads video frame data from the DDR and transmits it to a finite state machine (FSM). The FSM controls the order and timing of data reading and transmission to the HDMI interface according to its internal state. At the same time, the FSM controls the generation of synchronization signals, including horizontal synchronization, vertical synchronization, and pixel enable signals. These synchronization signals determine the specific format and time series of the video frames, ultimately causing the HDMI interface to output images in the correct format and speed.

[0102] Next, after outputting the first frame data, subsequent new frame data continues to enter the video input signal ring buffer. At this time, the first frame data in the video input signal ring buffer is 1 frame time-shifted from the first frame data in the remaining three ring buffers.

[0103] 4. Output the AI ​​screen after overlay compositing.

[0104] First, the output time of the first frame data of the video input signal occupied by the shifted output is used to perform a graphic overlay on the first frame data of the headers in the lesion detection result video ring buffer, lesion segmentation result video ring buffer, and semantic scene detection result video ring buffer to obtain a synthesized AI graphic presentation signal. (It should be noted that during this process, whether or not the graphic overlay can be completed within the time of one frame, i.e., 1 second / 60 frames ≈ 16.66 ms, depends on the operating speed of the FPGA chip. For example, an excessively low-grade FPGA chip may not be able to meet this performance requirement, but this invention can meet this performance requirement by using a high-grade FPGA chip.)

[0105] In the present invention, the synthesized AI graphic presentation signal is obtained by the following steps. (1) Data from the first frame is extracted from the lesion target detection result video ring buffer, the lesion segmentation result video ring buffer, and the semantic scene result video ring buffer, respectively, and each pixel is traversed sequentially based on the size information. (2) Determine whether the three obtained pixel values ​​are the background color or not (assuming the background color is white with values ​​of 255,255,255, as mentioned above). (3) If all are background colors, the pixel values ​​at that position in the AI ​​graphic presentation signal are also set to background colors. If one to three of the three pixel values ​​are not background color values, the pixel values ​​at that position in the AI ​​graphic presentation signal are set according to the priority order. The priority order is: 1. Frame data pixel values ​​corresponding to the semantic scene result video → 2. Frame data pixel values ​​corresponding to the lesion segmentation result video → 3. Frame data pixel values ​​corresponding to the lesion detection result video. (4) Finally, the synthesized AI graphic presentation signal is obtained.

[0106] Next, an overlay is performed on the frame data of the video input signal and the synthesized AI graphic presentation signal. Specifically, the frame data of the video input signal is read from the video input signal ring buffer, and a judgment is made for each pixel value in relation to the synthesized AI graphic presentation signal. If the value is the same as the background color (here, as mentioned above, we assume it is white with values ​​of 255,255,255), the pixel of the video input signal frame data is output. If the value is different from the background color, the pixel value is set in the synthesized AI graphic presentation signal, and finally, the AI ​​screen after overlay synthesis is obtained.

[0107] The aforementioned display module is used to display the video after it has been overlaid.

[0108] Preferably, the internal structure of the display module includes a video input terminal chip, a video output terminal chip, and an image input source switching key. The video acquisition module, AI edge computing module, and video overlay module are all connected to the video input terminal chip of the display module. The image input source switching key is used to determine whether the content displayed by the display module is from the video acquisition module, the AI ​​edge computing module, or the video overlay module.

[0109] Thus, the display module can display according to the display mode selected by the user. The user can control the switching display mode using the image input source switching key. Display mode 1 directly displays the original video input signal of the medical imaging device, display mode 2 displays a video signal with marking information that is overlaid on the original video input signal after processing by the AI ​​edge computing module and output by the marking layer output module, and display mode 3 displays the operating system HDMI signal output from the AI ​​edge computing module. Therefore, by providing three display modes, the present invention has a multifunctional effect and provides greater convenience to medical professionals.

[0110] Finally, it should be noted that the above embodiments are merely for illustrating the technical solutions of the present invention and do not limit the scope of protection of the present invention. Those skilled in the art can modify or replace the technical solutions of the present invention with equivalent replacements based on the spirit of the present invention without departing from the essence and scope of the technical solutions of the present invention.

Claims

1. An intelligent medical display system including a video acquisition module, an AI edge computing module, a marking layer output module, a video overlay module, and a display module, The video acquisition module is used to receive video input signals from a medical imaging device and transmit them to the AI ​​edge computing module. The AI ​​edge computing module is used to analyze the video input signal to acquire lesion target detection messages, lesion segmentation messages, and semantic scene detection messages, and to push them to the marking layer output module. The marking layer output module includes an information processing submodule and an image push delivery submodule. The information processing submodule is used to perform group marking on messages pushed by the AI ​​edge computing module and divide them into lesion target detection message groups, lesion segmentation message groups, and semantic scene detection message groups. The image push delivery submodule is used to perform marking on a background image based on the group marking results of the information processing submodule, to form lesion target detection result videos, lesion segmentation result videos, and semantic scene detection result videos, and to output them to the video overlay module. The video overlay module is used to overlay the lesion target detection result video, lesion segmentation result video, and semantic scene detection result video from the image push delivery submodule, and the video input signal from the video acquisition module, and to output the overlaid video to the display module. An intelligent medical display system characterized in that the display module is used to display video after it has been overlaid.

2. The information processing submodule includes a first buffer, and a lesion target detection buffer, a lesion segmentation buffer, and a semantic scene detection buffer connected to the first buffer, wherein the first buffer is connected to the AI ​​edge computing module, receives and temporarily stores messages pushed by the AI ​​edge computing module, analyzes the message header of the messages, temporarily stores the lesion target detection messages in the lesion target detection buffer, temporarily stores the lesion segmentation messages in the lesion segmentation buffer, and temporarily stores the semantic scene detection messages in the semantic scene detection buffer, thereby dividing them into a lesion target detection message group, a lesion segmentation message group, and a semantic scene detection message group, as described in claim 1.

3. The aforementioned image push delivery submodule includes a texture channel creation unit, a texture channel context creation unit, a background image drawing unit, a graphics drawing unit, and a video output unit. The texture channel creation unit is used to create a lesion target detection texture channel, a lesion segmentation texture channel, and a semantic scene detection texture channel, and to bind them to three output ports. The texture channel context creation unit is used to create a lesion target detection texture channel context, a lesion segmentation texture channel context, and a semantic scene detection texture channel context for the three output ports, respectively. The background image rendering unit is used to render background images in the lesion target detection texture channel, lesion segmentation texture channel, and semantic scene detection texture channel, respectively. The graphics drawing unit draws a lesion target graphic on the background image in the lesion target detection texture channel based on the lesion target detection message group, draws a lesion division graphic on the background image in the lesion division texture channel based on the lesion division message group, and draws a semantic scene graphic on the background image in the semantic scene detection texture channel based on the semantic scene detection message group. The intelligent medical display system according to claim 2, characterized in that the video output unit is used to form a lesion target detection result video from a background image and a lesion target graphic drawn thereon, to form a lesion segmentation result video from a background image and a lesion segmentation graphic drawn thereon, to form a semantic scene detection result video from a background image and a semantic scene detection graphic drawn thereon, and to output each of these to the video overlay module via the three output ports.

4. The intelligent medical display system according to claim 3, wherein the video overlay module includes four input terminals and one output terminal, the four input terminals being used to receive lesion target detection result videos, lesion segmentation result videos, and semantic scene detection result videos from the video output unit, and video input signals from the video acquisition module, and the output terminal being used to output the overlaid video to the display module.

5. The aforementioned video overlay module is 1) A step of establishing and initializing a video input signal ring buffer, a lesion detection result video ring buffer, a lesion segmentation result video ring buffer, and a semantic scene detection result video ring buffer, and simultaneously setting the capacity of each of the ring buffers, 2) Create intercept threads for the input channels of the four input terminals, so that when each input terminal receives video input, it can write frame data to the corresponding ring buffer, and when the capacity of the ring buffer is full, insert new frame data and pop up header frame data, 3) Create a shifted output; that is, when the video input signal ring buffer first reaches a predetermined storage length, output the data of the first frame directly without any processing, and at this time, shift the data of the first frame in the video input signal ring buffer by one frame in time with the data of the first frame in the other three ring buffers. 4) The process of outputting the AI ​​screen after overlay compositing, The intelligent medical display system according to claim 4, characterized by performing a video overlay.

6. The aforementioned step 4) specifically means, 4.1) Using the output time of the first frame data of the video input signal occupied by the shifted output, a graphic overlay is performed on the first frame data of the headers of the lesion detection result video ring buffer, lesion segmentation result video ring buffer, and semantic scene detection result video ring buffer to obtain a synthesized AI graphic presentation signal. 4.2) The intelligent medical display system according to claim 5, characterized in that it includes overlaying the frame data of the synthesized AI graphic presentation signal and the video input signal to finally obtain the AI ​​screen after the overlay synthesis.

7. The AI ​​edge computing module includes a lesion target detection submodule, a lesion segmentation submodule, and a semantic scene detection submodule, wherein the lesion target detection submodule is used to recognize the video input signal frame by frame, obtain a lesion detection object target, convert the lesion detection object target into a message, and obtain the lesion target detection message; the lesion segmentation submodule is used to recognize the video input signal frame by frame, obtain a lesion segmentation object target, convert the lesion segmentation object target into a message, and obtain the lesion segmentation message; and the semantic scene detection submodule is used to recognize the video input signal frame by frame, obtain a semantic scene detection result, convert the semantic scene detection result into a message, and obtain the semantic scene detection message, characterized in that the intelligent medical display system is as described in claim 1.

8. The intelligent medical display system according to claim 7, characterized in that the lesion target detection message includes a message header and a message body, the message header being DET, the message body being the coordinates of a plurality of lesion detection object targets, the plurality of coordinates being separated by the "|" symbol, the lesion segmentation message includes a message header and a message body, the message header being SEG, the message body being a set of point coordinates of a plurality of lesion segmentation object targets, the set of point coordinates of a plurality of groups being separated by the "|" symbol, and the semantic scene detection message includes a message header and a message body, the message header being TXT, the message body being presentation characters corresponding to a plurality of semantic scenes, the presentation characters of a plurality of groups being separated by the "|" symbol.

9. The intelligent medical display system according to claim 1, characterized in that the display module has an image input source switching key, the video acquisition module, the AI ​​edge computing module and the video overlay module are all connected to the display module, and the image input source switching key is used to determine whether the content displayed by the display module is from the video acquisition module, the AI ​​edge computing module or the video overlay module.