Multifunctional endoscope camera system and collaborative control method
Patent Information
- Application Number
- CN202610832763.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-10
- Publication Date
- 2026-09-15
Smart Images

Figure CN122744686A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, specifically to a multifunctional endoscopic imaging system and its collaborative control method. Background Technology
[0002] Endoscopic imaging systems are core equipment in modern minimally invasive surgery. With evolving clinical needs, single-function systems are insufficient to meet the requirements of combined multi-endoscopic surgeries and multimodal image fusion diagnosis and treatment. Current endoscopic imaging systems mostly employ an integrated, closed architecture, with tightly coupled functional modules that cannot be independently upgraded or expanded. More importantly, when clinicians attempt to expand functionality by simply connecting external, independent ultrasound, AI processing, or other devices, they face the following prominent problems:
[0003] 1. Lack of a unified communication and scheduling mechanism: Different devices use their own independent communication interfaces and proprietary protocols, resulting in chaotic control commands, data synchronization failures, and the inability of the system to work collaboratively as a whole.
[0004] 2. Difficulty in accurately fusing multimodal data: There are random delays in the arrival time of different source data such as endoscopic images, ultrasound images, and AI analysis results. Direct superposition will cause frame misalignment, resulting in augmented reality (AR) navigation or annotation deviation, and loss of clinical reference value.
[0005] 3. Rigid resource allocation: When processing multiple 4K videos and AI inference tasks simultaneously, it is impossible to dynamically allocate computing, bandwidth and other resources according to the surgical scenario, which can easily cause delays or frame drops in core tasks and affect surgical safety.
[0006] Therefore, there is an urgent need for a collaborative control method that can treat various extended devices as controlled units and perform task scheduling, frame synchronization, and multimodal data fusion through a unified protocol, so as to achieve multifunctional on-demand allocation and dynamic expansion and reduce equipment upgrade costs. Summary of the Invention
[0007] In view of the above, the purpose of this invention is to provide a multifunctional endoscopic camera system and a collaborative control method. By defining a unified communication protocol, topology-based task scheduling, and a multi-source data frame synchronization and fusion mechanism, the system solves the synchronization, scheduling, and communication bottlenecks when multiple modules work together, providing high-precision and low-latency technical support for multi-scope combined surgery and AI-assisted diagnosis and treatment.
[0008] To achieve the aforementioned objectives, the present invention adopts the following technical solution.
[0009] A multifunctional endoscopic imaging system and a collaborative control method are disclosed. The system includes a host device for constituting the endoscopic system and at least one expansion device for constituting a multifunctional expansion system. The host device serves as a host computer. The method includes the following steps: S1. System Initialization and Topology Discovery: After the host device is powered on, it broadcasts a device discovery command to the extension interface through a unified communication protocol, receives device identification and function description information returned by each access extension device, and dynamically constructs the functional topology diagram of the current system. S2. Working mode configuration and resource allocation: Based on the composite working mode selected by the user and the functional topology diagram, the host generates a collaborative workflow, parses the composite working mode into a sequence of atomic tasks, and allocates the atomic tasks to at least one target execution unit according to the computing power and bandwidth resources of each extended device. The target execution unit includes the processing module of the host device itself or the extended device. S3. Multi-source data frame synchronization and fusion scheduling: The host receives image data streams from different sources, performs frame alignment of the multi-source data streams using timestamps; according to the collaborative workflow, it generates a synthetic video stream containing at least two of the following: main screen, auxiliary screen, AI analysis annotation, and ultrasound image, and allocates the synthetic video stream to the output paths of different monitors in real time. S4. Distributed processing and result feedback: When an atomic task requires AI processing, the host sends the corresponding data stream to the data intelligence module through a unified communication protocol. After the data intelligence module completes the processing, it sends the labeled data or inference results back to the host in the form of metadata for fusion display in step S3.
[0010] Furthermore, the unified communication protocol described in step S1 defines a frame structure for multiplexing control signaling, data streams, and video streams. This frame structure sequentially includes: a frame header field, a module ID field, a data type field, a timestamp field, a data length field, a data payload field, and a CRC checksum field. The data type field is used to distinguish between control commands, real-time video streams, AI metadata, and raw ultrasound data. The control commands include mode switching commands, resource allocation commands, and time synchronization commands.
[0011] Furthermore, the composite working mode is at least one of multi-scope combination, AI assistance, and ultrasound fusion. The specific configuration of the collaborative workflow generated in step S2 is as follows: control the image extension module to access at least one auxiliary endoscope signal and transmit it back to the host; control the ultrasound module to start continuous acquisition and transmit the ultrasound image stream back to the host; the host combines the main endoscope video stream and the auxiliary endoscope video stream and distributes it to the digital intelligence module; the host receives the AI annotation information transmitted back from the digital intelligence module, performs secondary fusion with the ultrasound image stream and the endoscope composite video stream, and outputs it.
[0012] Furthermore, the fusion scheduling in step S3 employs a multimodal image fusion control method, specifically:
[0013] The host receives endoscopic video stream, ultrasound image stream, and AI analysis result stream returned by the digital intelligence module. The AI analysis result includes, but is not limited to, the boundary coordinates, classification attributes, and confidence level of the lesion area given by the auxiliary diagnosis.
[0014] The image and video processing module uses the coordinate information in the AI analysis results to dynamically map the ultrasound image to the corresponding anatomical position in the endoscopic field of view according to the preset spatial registration parameters, and performs color enhancement and outline drawing on the AI-annotated area.
[0015] If there is a time discrepancy between ultrasound images and endoscopic images, the timestamps in their respective data packets are used to perform frame alignment in the fusion buffer, with the alignment discrepancy not exceeding 500 microseconds.
[0016] In the synthesized multimodal fusion image, the transparency and display priority of the ultrasound layer and the AI annotation layer can be adjusted by the user in real time via the touch screen, or automatically switched by the system according to the surgical steps.
[0017] Furthermore, the spatial registration parameters are obtained through preoperative calibration or dynamically updated by the digital intelligence module based on feature point matching between real-time endoscopic images and ultrasound images, using an online registration algorithm.
[0018] A multifunctional endoscopic camera system for implementing the above method includes a host device and at least one expansion device. The host device is a host computer relative to the expansion device, and the two are connected via a high-speed communication cable supporting the unified communication protocol. The host device includes:
[0019] The interface board module provides at least one endoscope camera interface, at least one main monitor interface, a touch screen interface, and a standardized expansion interface for communicating with expansion devices; the expansion interface supports hot-swapping and can automatically identify the module ID and functional description information of the connected device.
[0020] The image and video processing module, connected to the interface board module, is used to process images acquired by the endoscope camera in real time and perform frame synchronization, picture-in-picture composition, multimodal fusion and output path allocation of multi-source images.
[0021] An embedded control and storage module is connected to the image and video processing module and the interface board module. It has an embedded protocol parsing and scheduling unit, which is used to generate and parse the unified communication protocol, perform system initialization and topology discovery, task parsing and resource allocation, and is responsible for recording surgical process videos and capturing images.
[0022] The touchscreen module and the main monitor provide the human-computer interaction interface and the real-time display of the surgical scene, respectively.
[0023] The extended device includes one or more of a digital intelligence module, an ultrasound module, and an imaging extended module. Each extended module has a built-in protocol parsing engine that can respond to the host's control commands and transmit data back as required.
[0024] The digital intelligence module has a built-in high-performance computing unit, which is used to run artificial intelligence algorithms to realize functions such as AI-assisted diagnosis, surgical navigation, 3D reconstruction, multimodal image fusion, intelligent robotic arm, and voice control. It also encapsulates the processed AI metadata in the frame structure of the unified communication protocol and sends it back to the host device through the protocol parsing engine.
[0025] The ultrasound module has a built-in ultrasound imaging unit, which is used to connect the ultrasound probe to realize the ultrasound endoscopy imaging function, and encapsulates the raw ultrasound data through the protocol parsing engine and transmits it to the host device or the digital module.
[0026] The image expansion module has a built-in multi-channel video input interface, which is used to receive at least one endoscope camera signal and encapsulate the video stream through the protocol parsing engine before transmitting it to the host device, so as to realize the fusion display of multiple screens, picture-in-picture display or screen switching function.
[0027] The secondary monitor can be connected to the image expansion module, digital intelligence module, and ultrasound module to assist in displaying expanded image, AI-assisted diagnostic image, and ultrasound image.
[0028] At least one extended camera device, connected to an image extension module, is used to provide additional endoscopic images during multi-scope combined surgery.
[0029] Furthermore, the mode switching state machine within the embedded control and storage module includes at least five states: standby, idle, configuration, synchronous operation, and exit. The transition from the idle state to the configuration state is triggered by the user selecting a composite working mode. In the configuration state, the host sends configuration instructions to each extended device one by one. After all devices are ready and time synchronization is completed, the system transitions to the synchronous operation state, where multi-source data frame synchronization and fusion scheduling are performed.
[0030] Compared with the prior art, the beneficial effects of the present invention, which adopts the aforementioned technical solution, are as follows:
[0031] (1) By defining a unified communication protocol and frame structure, the multi-channel multiplexing and precise synchronization of control flow, data flow and video flow between the host and each extended device are realized, which completely solves the problem of instruction conflict and data delay when multiple devices work together, and ensures millisecond-level collaborative response of each module in complex surgical mode;
[0032] (2) The resource allocation method based on topology discovery and adaptive task scheduling can dynamically analyze the composite working mode into atomic tasks according to the clinical scenario and intelligently allocate computing power and bandwidth resources. Without replacing the hardware, it greatly improves the system's concurrent processing efficiency for complex tasks such as multi-scope combination, AI assistance, and ultrasound fusion, and effectively avoids frame loss and delay in core tasks.
[0033] (3) By using multi-source data frame synchronization and multi-modal image fusion control technology, multi-source heterogeneous images such as endoscopy, ultrasound and AI analysis results are precisely aligned in time and space, realizing augmented reality-level surgical navigation and lesion area visualization, significantly improving the accuracy and clinical reference value of multi-modal image fusion, and reducing the complexity of surgical operation and the risk of misjudgment. Attached Figure Description
[0034] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0035] Figure 1 This is a schematic diagram of a multifunctional endoscopic camera system according to an embodiment of the present invention. Detailed Implementation
[0036] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the described embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0037] Unless otherwise defined, the technical or scientific terms used in this invention shall have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in this invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0038] To keep the following description of the embodiments of the present invention clear and concise, detailed descriptions of some known functions and components have been omitted.
[0039] like Figure 1 As shown, a multifunctional endoscope camera system includes a host device 1 for constituting the basic functions of the endoscope system and at least one extension device 2 for constituting a multifunctional extension system. The host device is a host computer relative to the extension device, and the two are connected by a high-speed communication cable supporting a unified communication protocol.
[0040] The host device includes: an interface board module 11, an image and video processing module 12, an embedded control and storage module 13, a touch screen module 14, a main monitor 15, and at least one camera device 16. The interface board module 11 has hot-swappable expansion interfaces and can automatically identify the module ID and functional description information of the connected devices. The image and video processing module 12 is used to perform image preprocessing, multi-source frame synchronization, and multimodal image fusion. The embedded control and storage module 13 has a built-in protocol parsing and scheduling unit, used to generate and parse frames of the unified communication protocol, perform system initialization and topology discovery, task parsing and resource allocation, and implement mode switching state machine management.
[0041] The expansion module includes one or more of the following modules: a digital intelligence module 21, an ultrasound module 22, an image expansion module 23, a secondary monitor 24, and at least one extended camera device 25. Each expansion module has a built-in protocol parsing engine, capable of responding to host commands and encapsulating and transmitting data as required.
[0042] To address the challenges of interconnectivity and collaborative scheduling among multiple devices, this invention designs a unified communication protocol based on a master-slave architecture. All extended devices are connected to the host via a customized multi-core cable that simultaneously supports power supply and high-speed data communication. The protocol operates on a 10 Gigabit serial link at the underlying level, enabling multi-channel multiplexing of control, data, and video streams. The protocol employs a request-response model, with the host acting as the master controller and each extended module acting as a slave device responding to commands.
[0043] The unified communication protocol defines a basic frame structure. Each frame contains the following fields: a frame header field (4 bytes long, using a fixed value) used to identify the start of a frame and achieve frame synchronization; a module ID field (2 bytes long) used to identify the source or target module. In this embodiment, 0x01 represents the host, 0x10 represents the AI module, 0x20 represents the ultrasound module, and 0x30 represents the image extension module; a data type field (1 byte long) used to distinguish the payload type. In this embodiment, 0x01 is the control command, 0x02 is the real-time video stream, 0x03 is the AI metadata, 0x04 is the ultrasound raw data, and 0x05 is the response message; a timestamp field (8 bytes long, using a 64-bit system timestamp in nanoseconds, timed uniformly by the host, used to achieve frame synchronization of multi-source data); a data length field (4 bytes long) used to indicate the number of bytes in the data payload; a data payload field carrying the actual transmitted data content; and a CRC check field (4 bytes long) used to perform cyclic redundancy check on the entire frame to ensure the correctness of transmission.
[0044] All data streams must be encapsulated into this frame format before transmission. The AI metadata payload uses JSON or a custom binary format, containing lesion boundary coordinates, category, confidence level, and timestamps of associated image frames to ensure spatiotemporal consistency in the fused display.
[0045] The protocol parsing and scheduling unit embedded in the host manages a finite state machine to handle mode switching under different surgical scenarios. This ensures that all modules start in an orderly manner, are strictly synchronized, and reliably exit under complex modes, avoiding data loss or deadlock due to state chaos. The state machine contains at least five states: standby, idle, configuring, synchronously running, and exiting. Its state transition process is as follows:
[0046] After the system completes its power-on self-test, the state machine enters the "standby" state. In the standby state, the host automatically performs initialization operations, broadcasts a topology discovery frame to the expansion interface, collects the ID and status information of each access module, and then the state machine transitions to the "idle" state.
[0047] When a user selects a multi-mirror combination, AI, ultrasound, or at least one composite working mode on the touchscreen, the state machine is triggered to transition from the "idle" state to the "configuration" state. In the configuration state, the host first parses the configuration file corresponding to the selected mode, generating a list of atomic tasks containing multiple subtasks. Typical tasks include activating the imaging extension module, starting continuous acquisition by the ultrasound module, and configuring the digital intelligence module to load lesion detection and instrument recognition models. Subsequently, the host sends configuration commands to each relevant extension device: after sending an activation command to the imaging extension module and receiving confirmation, it allocates video input channels and sets auxiliary camera parameters; it sends a start acquisition command to the ultrasound module, specifying the output target as the host, and the ultrasound module immediately begins acquisition and encapsulates the ultrasound data stream for transmission; it sends a model loading command and working parameters to the digital intelligence module, which loads the corresponding model from local storage and returns to the ready state.
[0048] Once all subtasks are ready, the host performs resource verification. If the current system bandwidth or computing power is insufficient, the host issues an alarm and performs a downgrade configuration; if the verification passes, the host broadcasts a "time synchronization" frame to all modules, resets the clock counters of all modules, and then the state machine enters the "synchronous operation" state.
[0049] In synchronous operation mode, the host continuously receives data streams from each module, performs multi-source frame synchronization and fusion scheduling, generates a composite video stream, and outputs it to the corresponding monitor. If the user chooses to pause or switch to another mode, the state machine will enter the "exiting" state. In the exiting state, the host sends exit commands to each module in sequence, stopping data stream transmission and releasing allocated resources for each module. When all modules return idle confirmation, the state machine returns to the "idle" state, waiting to receive new user commands.
[0050] Under the "multi-mirror combination + AI + ultrasound" composite mode, the multimodal image fusion control method proposed in this invention is specifically implemented as follows:
[0051] The image and video processing module 12 internally establishes a multi-source fusion buffer. The main endoscope video stream from camera 16, the auxiliary endoscope video stream from image extension module 23, the ultrasound image stream from ultrasound module 22, and the AI metadata stream returned by digital intelligence module 21 each enter their respective circular buffer queues;
[0052] The scheduling unit continuously reads the timestamps of the data packets at the head of each buffer queue. Using the timestamp T_master of the main endoscope frame as a reference, it selects the latest frame from other data sources whose timestamp falls within the [T_master - 500μs, T_master + 500μs] window for alignment. If no matching frame exists in the current alignment window for any of the main endoscope video stream, auxiliary endoscope video stream, ultrasound image stream, or AI metadata stream, that source reuses its previously aligned frame or discards the current frame to be aligned to avoid lag.
[0053] The AI model running in the digital intelligence module 21 detects the lesion area in real time from the synthetic endoscope image, and the generated AI metadata includes: the coordinates of the vertices of the region polygon (based on the coordinate system of the main endoscope image), the lesion type and confidence level, the position of the instrument tip, etc.
[0054] The image and video processing module 12 pre-stores a spatial registration matrix generated by preoperative CBCT or ultrasound calibration. Using this matrix, the two-dimensional sector section of the ultrasound image is dynamically mapped to the corresponding position on the main endoscopic screen. Highlighting and pseudo-color enhancement are applied to the AI-annotated lesion boundaries to form a semi-transparent overlay layer.
[0055] In the final composite image, the layer priorities from high to low are: instrument tip annotation, AI lesion contour enhancement, ultrasound sub-image, auxiliary endoscope picture-in-picture, and main endoscope background. Users can independently adjust the transparency of the ultrasound layer and AI annotation layer via the slider on the touchscreen 14. In key steps such as "anatomical exposure," the system can automatically increase the transparency of the ultrasound layer and hide the AI annotation to reduce visual interference.
[0056] To optimize resources, the scheduling unit reallocates tasks based on the scenario. For example, in a "polyp removal" scenario, if the host determines that the image expansion module is not in use, it allocates its built-in image processing unit via control commands to share the picture-in-picture scaling and color correction tasks of the host's image and video processing module, while reserving all of the host's FPGA resources for the multimodal fusion pipeline. Simultaneously, the computing power of the digital intelligence module is prioritized for the polyp detection model, increasing the frame rate and reducing inference latency. The entire process is automatically completed by control commands using a unified communication protocol, requiring no human intervention.
[0057] The above description is merely an embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent modifications made based on the content of the present invention specification and drawings, or direct or indirect applications in related technical fields, are similarly included within the patent protection scope of the present invention.
Claims
1. A multi-functional endoscope camera system and a cooperative control method, the system including a host device for constituting an endoscope system and an extension device for constituting a multi-functional extension system, the host device being a host computer with respect to the extension device, characterized by, The method includes: S1. System Initialization and Topology Discovery: After the host device is powered on, it broadcasts a device discovery command to the extension interface through a unified communication protocol, receives device identification and function description information returned by each access extension device, and dynamically constructs the functional topology diagram of the current system. S2. Working mode configuration and resource allocation: Based on the composite working mode selected by the user and the functional topology diagram, the host generates a collaborative workflow, parses the composite working mode into a sequence of atomic tasks, and allocates the atomic tasks to at least one target execution unit according to the computing power and bandwidth resources of each extended device. The target execution unit includes the processing module of the host device itself or the extended device. S3. Multi-source data frame synchronization and fusion scheduling: The host receives image data streams from different sources, performs frame alignment of the multi-source data streams using timestamps; according to the collaborative workflow, it generates a synthetic video stream containing at least two of the following: main screen, auxiliary screen, AI analysis annotation, and ultrasound image, and allocates the synthetic video stream to the output paths of different monitors in real time. S4. Distributed processing and result feedback: When an atomic task requires AI processing, the host sends the corresponding data stream to the data intelligence module through the unified communication protocol. After the data intelligence module completes the processing, it sends the labeled data or inference results back to the host in the form of metadata for fusion display in step S3.
2. The multi-functional endoscope camera system and collaborative control method according to claim 1, characterized by: The unified communication protocol described in step S1 defines a frame structure for multiplexing control signaling, data streams, and video streams. The frame structure sequentially includes: a frame header field, a module ID field, a data type field, a timestamp field, a data length field, a data payload field, and a CRC check field. The data type field is used to distinguish between control commands, real-time video streams, AI metadata, and raw ultrasound data. The control commands include at least mode switching commands, resource allocation commands, and time synchronization commands.
3. The multi-functional endoscope camera system and collaborative control method according to claim 1, characterized by: The composite working mode is at least one of multi-scope combination, AI assistance, and ultrasound fusion. The specific configuration of the collaborative workflow generated in step S2 is as follows: control the image expansion module to access at least one auxiliary endoscope signal and transmit it back to the host; control the ultrasound module to start continuous acquisition and transmit the ultrasound image stream back to the host; the host combines the main endoscope video stream and the auxiliary endoscope video stream and distributes it to the digital intelligence module; the host receives the AI annotation information transmitted back from the digital intelligence module, performs secondary fusion with the ultrasound image stream and the endoscope composite video stream, and outputs it.
4. The multi-functional endoscope camera system and collaborative control method according to claim 1, characterized by: The fusion scheduling in step S3 employs a multimodal image fusion control method, including: receiving an endoscopic video stream, an ultrasound image stream, and an AI analysis result stream transmitted back by the digital intelligence module, wherein the AI analysis result includes the boundary coordinates and attribute information of the lesion area; using the coordinate information in the AI analysis result, dynamically mapping the ultrasound image to the corresponding anatomical position in the endoscopic field of view according to preset spatial registration parameters, and visually enhancing the AI-annotated area; if there is a temporal deviation between different source data streams, using the timestamps in each data packet to perform frame alignment in the fusion buffer, so that the alignment deviation is lower than a preset threshold; the synthesized multimodal fused image allows the user to adjust the display attributes of the ultrasound layer and the AI annotation layer in real time, or the system can automatically switch the display priority of the layers according to the surgical steps.
5. The multifunctional endoscopic imaging system and collaborative control method according to claim 4, characterized in that: The spatial registration parameters are obtained through preoperative calibration or dynamically updated by the digital intelligence module based on feature point matching between real-time endoscopic images and ultrasound images, using an online registration algorithm.
6. A multifunctional endoscopic imaging system for performing the method according to any one of claims 1 to 5, characterized in that, It includes a host device and at least one expansion device, wherein the host device acts as a host computer, and the two are connected via a high-speed communication cable supporting a unified communication protocol. The host device includes: The interface board module provides standardized expansion interfaces that support hot-swapping and automatically identify the module ID and functional description information of the access device. The image and video processing module performs image preprocessing, multi-source frame synchronization, and multimodal image fusion. The embedded control and storage module, with a built-in protocol parsing and scheduling unit, generates and parses frames of the unified communication protocol, performs system initialization and topology discovery, task parsing and resource allocation, and implements mode switching state machine management. The extended devices include at least one of a digital intelligence module, an ultrasound module, and an image extension module. Each extension module has a built-in protocol parsing engine capable of responding to host commands and encapsulating and transmitting data as required.
7. A multifunctional endoscopic imaging system according to claim 6, characterized in that, The mode switching state machine in the embedded control and storage module includes at least five states: standby, idle, configuration, synchronous operation, and exit. The transition from the idle state to the configuration state is triggered by the user selecting a composite working mode. In the configuration state, the host sends configuration instructions to each extended device one by one. After all devices are ready and time synchronization is completed, the system transitions to the synchronous operation state, where multi-source data frame synchronization and fusion scheduling are performed.