Endoscopic image data processing method and system based on foot pedal signals
Patent Information
- Application Number
- CN202611003838.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-07
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]本发明提供一种基于脚踏信号的内镜图像数据处理方法和系统,用以解决现有技术中的基于脚踏信号的内镜图像数据处理方法仅支持单一触发逻辑,人机交互效率低,灵活性和准确差的缺陷
[0015]本发明提供的一种基于脚踏信号的内镜图像数据处理方法和系统,通过接收脚踏输入设备发送的脚踏操作信号;提取脚踏操作信号的按压持续时间,在按压持续时间大于或等于预设的长按时间阈值的情况下,确定脚踏操作信号为长按脚踏操作信号;确定不同脚踏操作信号与不同功能指令之间的映射关系,基于长按脚踏操作信号和映射关系,生成目标特征分析触发指令;响应于目标特征分析触发指令,将内镜设备输出的当前视频流数据中的当前图像序列数据输入图像特征分析模型,得到结构化图像特征数据,提升了人机交互的效率,提高了图像数据处理的灵活性和准确性;通过在用户界面上,将结构化图像特征数据与当前视频流数据进行同步渲染显示,为操作人员提供了直观、即时的视觉反馈,提升了用户体验。
Smart Images

Figure CN122824972A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image data processing technology, and in particular to an endoscopic image data processing method and system based on foot pedal signals. Background Technology
[0002] In recent years, endoscopic equipment has been widely used in various internal space exploration and inspection scenarios. When operating an endoscope, operators typically need to use both hands to hold the endoscope handle for precise posture and field of view adjustments, making it difficult to free their hands to operate the control panel of external equipment to complete image storage and analysis commands. Therefore, foot switches have been widely introduced as auxiliary input devices into endoscopic data processing systems to achieve contactless human-machine interaction.
[0003] However, existing endoscopic foot pedal control technologies are often limited in function, typically supporting only simple, single-trigger hardware-level operations. With the widespread adoption of image processing technology, endoscopic systems are integrating increasingly complex video stream analysis functions. Configuring a separate physical pedal for each function could easily lead to accidental activation by operators in confined working environments; conversely, relying on a single pedal leaves existing systems lacking intelligent analysis of foot pedal signal characteristics. This results in low human-computer interaction efficiency when operators need to flexibly switch between image acquisition and feature analysis functions, easily disrupting the continuity of probing operations and failing to meet the urgent needs of modern intelligent endoscopic systems for efficient and accurate image data processing. Summary of the Invention
[0004] This invention provides a method and system for processing endoscopic image data based on foot pedal signals, which solves the shortcomings of existing endoscopic image data processing methods based on foot pedal signals, which only support a single trigger logic, have low human-computer interaction efficiency, and are inflexible and inaccurate.
[0005] This invention provides a method for processing endoscopic image data based on foot pedal signals, comprising: Receive foot pedal operation signals sent by the foot pedal input device; Extract the pressing duration of the foot pedal operation signal. If the pressing duration is greater than or equal to a preset long press time threshold, determine that the foot pedal operation signal is a long press foot pedal operation signal. Determine the mapping relationship between different foot pedal operation signals and different function commands, and generate a target feature analysis trigger command based on the long press foot pedal operation signal and the mapping relationship; In response to the target feature analysis trigger command, the current image sequence data in the current video stream data output by the endoscope device is input into a pre-trained image feature analysis model to obtain structured image feature data output by the image feature analysis model; the current image sequence data includes multiple video frame image data within the current time window; On the user interface, the structured image feature data and the current video stream data are rendered and displayed synchronously. After receiving the foot pedal operation signal sent by the foot pedal input device, the method further includes: Extract the pressing frequency of the foot pedal operation signal, and determine the type of the foot pedal operation signal based on the pressing frequency; When the pressing frequency indicator is a single press, the foot pedal operation signal is determined to be a single foot pedal operation signal; If the pressing frequency indication is two consecutive presses, the foot pedal operation signal is determined to be a double foot pedal operation signal; When the pressing frequency indication is N consecutive pressings, the foot pedal operation signal is determined to be a continuous foot pedal operation signal, where N is a positive integer greater than 2.
[0006] In some embodiments, determining the mapping relationship between different foot pedal operation signals and different function commands includes: Obtain the current system's operating mode identifier, which indicates the current system's preset operating mode. The preset operating mode includes at least one of image acquisition mode, video recording mode, and AI analysis mode. Based on the operating mode identifier, the mapping relationship between different foot pedal operation signals and function commands is determined; under different preset operating modes, the same type of foot pedal operation signal is mapped to different function commands.
[0007] In some embodiments, after determining that the foot pedal operation signal is a single foot pedal operation signal, the method further includes: Based on the single foot pedal operation signal and the mapping relationship, an image acquisition command is generated; In response to the image acquisition command, extract the current frame image data from the current video stream data; Add a timestamp and / or identification information to the current frame image data, store the current frame image data in the database, and display a message indicating successful image data acquisition on the user interface.
[0008] In some embodiments, after determining that the foot pedal operation signal is a double foot pedal operation signal, the method further includes: Based on the dual foot pedal operation signals and the mapping relationship, a video recording start command or a video recording stop command is generated. In response to the video recording start command, the endoscopic device is controlled to perform a video recording start operation for the current video stream data; or, in response to the video recording stop command, the endoscopic device is controlled to perform a video recording stop operation for the current video stream data. The video recording status is displayed on the user interface. Save the current video stream data within the recording time period as a video file.
[0009] In some embodiments, after determining that the pedal operation signal is a continuous pedal operation signal, the method further includes: Based on the continuous foot pedal operation signals and the mapping relationship, a keyframe marking instruction is generated; In response to the keyframe marking instruction, the current frame image data is obtained from the current video stream data, the current frame image data is determined as keyframe image data, and a timestamp and / or identification information is added to the keyframe image data.
[0010] In some embodiments, receiving a foot pedal operation signal sent by a foot pedal input device and extracting the pressing duration of the foot pedal operation signal includes: Receives a foot pedal operation signal sent by the foot pedal input device, wherein the foot pedal operation signal is received at the current moment; Calculate the time interval between the foot pedal operation signal and the historical foot pedal operation signal, wherein the historical foot pedal operation signal was received at the previous time before the current time and sent by the foot pedal input device; If the time interval is determined to be greater than or equal to a preset anti-shake time threshold, the foot pedal operation signal is determined to be a valid trigger signal, and the pressing duration of the foot pedal operation signal is extracted.
[0011] In some embodiments, the structured image feature data includes: feature data of the target region of interest, location information, and feature confidence.
[0012] In some embodiments, the step of synchronously rendering and displaying the structured image feature data and the current video stream data includes: Based on the location information, a visual identifier box is generated and overlaid at the corresponding position of the current video stream data; In the associated area of the visual identifier box, the feature data and the feature confidence level are rendered and displayed; In the preset status bar area of the user interface, auxiliary prompts and operation information for the current image sequence data are dynamically updated.
[0013] In some embodiments, generating a target feature analysis trigger command based on the long-press foot pedal operation signal and the mapping relationship includes: Obtain the current system's operating status identifier, which indicates whether the current system is in a preset mode restriction state; When the working status indicator indicates that the current system is not in a preset mode restriction state, a target feature analysis trigger command is generated based on the long press foot pedal operation signal and the mapping relationship.
[0014] The present invention also provides an endoscopic image data processing system based on foot pedal signals, comprising: The receiving unit is used to receive foot pedal operation signals sent by the foot pedal input device; The determining unit is used to extract the pressing duration of the foot pedal operation signal, and determine the foot pedal operation signal as a long press foot pedal operation signal if the pressing duration is greater than or equal to a preset long press time threshold. The generation unit is used to determine the mapping relationship between different foot pedal operation signals and different function commands, and to generate a target feature analysis trigger command based on the long press foot pedal operation signal and the mapping relationship. The data processing unit is configured to, in response to the target feature analysis trigger command, input the current image sequence data from the current video stream data output by the endoscopic device into a pre-trained image feature analysis model to obtain structured image feature data output by the image feature analysis model; the current image sequence data includes multiple video frame image data within the current time window; The display unit is used to synchronously render and display the structured image feature data and the current video stream data on the user interface. The system further includes a signal classification unit, used to extract the pressing frequency of the foot pedal operation signal after the receiving unit receives the foot pedal operation signal sent by the foot pedal input device, and determine the type of the foot pedal operation signal based on the pressing frequency; if the pressing frequency indicates a single press, the foot pedal operation signal is determined to be a single foot pedal operation signal; if the pressing frequency indicates two consecutive presses, the foot pedal operation signal is determined to be a double foot pedal operation signal; if the pressing frequency indicates N consecutive presses, the foot pedal operation signal is determined to be a continuous foot pedal operation signal, where N is a positive integer greater than 2.
[0015] This invention provides a method and system for processing endoscopic image data based on foot pedal signals. The method involves receiving foot pedal operation signals from a foot pedal input device; extracting the pressing duration of the foot pedal operation signal; determining a long-press foot pedal operation signal if the pressing duration is greater than or equal to a preset long-press time threshold; determining the mapping relationship between different foot pedal operation signals and different functional commands; generating a target feature analysis trigger command based on the long-press foot pedal operation signal and the mapping relationship; responding to the target feature analysis trigger command, inputting the current image sequence data from the current video stream data output by the endoscopic device into an image feature analysis model to obtain structured image feature data. This improves the efficiency of human-computer interaction and enhances the flexibility and accuracy of image data processing. By synchronously rendering and displaying the structured image feature data and the current video stream data on the user interface, the method provides operators with intuitive and immediate visual feedback, improving the user experience. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0017] Figure 1 This is a flowchart illustrating the endoscopic image data processing method based on foot pedal signals provided by the present invention.
[0018] Figure 2 This is a schematic diagram of the process for extracting the pressing duration of the foot pedal operation signal provided by the present invention.
[0019] Figure 3 This is a schematic diagram of the endoscopic image data processing system based on foot pedal signals provided by the present invention.
[0020] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0022] Figure 1This is a flowchart illustrating the endoscopic image data processing method based on foot pedal signals provided by the present invention, as shown below. Figure 1 As shown, the method includes: Step 110: Receive the foot pedal operation signal sent by the foot pedal input device.
[0023] The foot input device refers to a hardware control peripheral located below the endoscope operating table for operators to press with their feet. The foot input device can be a single-pedal device or a multi-pedal combination device. A physical or logical communication connection can be established between the foot input device and the processing host executing the method of this embodiment via a wired universal serial bus, serial port, or wireless Bluetooth. In medical endoscopy, doctors often need to trigger operations immediately upon observing critical lesions; therefore, low-latency signal capture is crucial to avoid missing important diagnostic information.
[0024] Among them, the foot pedal operation signal refers to the electrical signal or digital pulse signal generated when the pedal of the foot pedal input device is pressed down or released.
[0025] Optionally, the start time stamp corresponding to the pedal being pressed and the end time stamp corresponding to the pedal being released are obtained from the foot input device. The difference between the end time stamp and the start time stamp is the pressing duration of the foot operation signal.
[0026] Step 120: Extract the pressing duration of the foot pedal operation signal. If the pressing duration is greater than or equal to the preset long press time threshold, determine the foot pedal operation signal as a long press foot pedal operation signal.
[0027] Optionally, the start time stamp corresponding to the pedal being pressed and the end time stamp corresponding to the pedal being released are obtained from the foot input device. The difference between the end time stamp and the start time stamp is the pressing duration of the foot operation signal.
[0028] The long press time threshold is a preset time limit parameter used to distinguish between short presses and continuous presses. The long press time threshold can be a specific value such as 0.5 seconds, 1 second, or 2 seconds, and its value can be flexibly configured in the system settings interface according to the operator's pressing habits.
[0029] Optionally, the extracted press duration is compared with a long press time threshold. If the press duration is greater than or equal to the long press time threshold, the current operation is determined to be an intentionally sustained pedaling action, and the pedal operation signal is then classified as a long press pedal operation signal.
[0030] Step 130: Determine the mapping relationship between different foot pedal operation signals and different function commands. Based on the long press foot pedal operation signal and the mapping relationship, generate target feature analysis trigger commands.
[0031] Specifically, the different foot pedal operation signals include at least a single foot pedal operation signal, a double foot pedal operation signal, a continuous foot pedal operation signal, and a long press foot pedal operation signal; the different function commands include at least an image acquisition command, a video recording start command, a video recording stop command, a keyframe marking command, a target feature analysis trigger command, and an auxiliary prompt command.
[0032] Optionally, the mapping relationship between different foot pedal operation signals and different function commands can be determined based on user input; different mapping rule configuration files can be loaded according to the preferences of different departments and users to improve the adaptability and personalization of the system.
[0033] Optionally, by calling the system's internal message queue or application programming interface, a software-level control command is issued in the background to activate the artificial intelligence analysis module. This command is the target feature analysis trigger instruction.
[0034] Step 140: In response to the target feature analysis trigger command, input the current image sequence data in the current video stream data output by the endoscope device into the pre-trained image feature analysis model to obtain the structured image feature data output by the image feature analysis model.
[0035] The endoscopic equipment can include rigid endoscopes, flexible endoscopes, and other detection devices used to acquire continuous images of the internal space. Current video stream data refers to the collection of high-definition digital video frames acquired in real time by the image sensors at the front end of the endoscopic equipment and transmitted to the processing host.
[0036] The current image sequence data consists of multiple consecutive image frames extracted from the current video stream. Upon receiving a target feature analysis trigger command, the system extracts the current image sequence data from the memory buffer and feeds it into the image feature analysis model.
[0037] The image feature analysis model is a mathematical model based on deep learning networks, such as convolutional neural networks, pre-trained using a large number of endoscopic sample images. This model can perform multi-dimensional feature extraction and classification regression calculations on the input image pixels, ultimately outputting the results. Structured image feature data refers to machine-readable data that has undergone formatted and encapsulated processing, containing a set of intelligent parsing results of the model on target regions in the image.
[0038] In some embodiments, the structured image feature data includes: feature data of the target region of interest, location information, and feature confidence.
[0039] The target region of interest refers to a tissue area with abnormal morphology or specific anatomical significance during endoscopic exploration, such as polyps, tumors, bleeding points, or abnormal vascular networks.
[0040] Feature data refers to the descriptive information generated by the image feature analysis model after intelligently parsing the aforementioned target region of interest. This includes the target's size, shape, color intensity, and predicted feature classification results. Location information refers to the spatial positioning data of the target region of interest within the two-dimensional pixel plane of the corresponding image frame in the current video stream data. This is specifically represented as the set of vertex pixel coordinates of the bounding box or a combination of the center point coordinates and the target's width and height. Feature confidence refers to the probability score given by the image feature analysis model for its output feature classification results. It quantifies the model's confidence in the recognition result and can be expressed as a percentage or a floating-point number between zero and one. By parsing and encapsulating the underlying tensor data output by the model, structured image feature data containing the above three types of information is generated.
[0041] Step 150: On the user interface, the structured image feature data and the current video stream data are rendered and displayed synchronously.
[0042] The user interface refers to the graphical interactive view displayed on a medical monitor or display screen connected to the processing host. The display method specifically employs a multi-layer overlay technique. The bottom layer continuously plays the current video stream data, while the upper layer performs geometric drawing and text label rendering based on the parsed structured image feature data. Synchronous rendering requires that the drawing frame rate of the upper-layer feature data and the refresh frame rate of the bottom-layer video stream maintain a high degree of consistency in timestamps, ensuring that the drawn feature markers can accurately follow the movement of the corresponding targets in the video frame and move in real time, avoiding screen tearing or marker delay.
[0043] It should be noted that the technical solution of this invention is particularly applicable to endoscopic diagnostic and treatment scenarios such as bronchoscopy, gastroscopy, and laparoscopy. In these scenarios, doctors need to operate the endoscopic equipment with both hands to complete tasks such as observation, positioning, and operation, and cannot directly contact traditional input devices such as keyboards and mice; at the same time, the treatment environment requires strict sterile conditions, and any contact between the hands and non-sterile equipment may pose a risk of cross-infection. Therefore, this invention uses a foot-operated input device as a human-computer interaction medium to achieve contactless control. Doctors can complete functions such as image acquisition, video recording, and AI analysis triggering solely through foot movements while maintaining sterile operation with both hands throughout the process, effectively ensuring a sterile environment and improving operational efficiency.
[0044] Understandably, by refining structured image feature data into three core dimensions—feature data, location information, and feature confidence—it can provide comprehensive data support for the subsequent synchronous rendering of the user interface. This not only enables operators to intuitively locate the precise spatial position and attribute details of the target area, but also helps them quickly assess the reliability of the algorithm's prompts based on feature confidence, thereby effectively improving the scientific nature and decision-making efficiency of human-machine collaborative processing of endoscopic image data.
[0045] In this embodiment of the invention, by accurately extracting and determining the pressing duration of the foot pedal operation signal, the long press action is intelligently bound to a complex image feature analysis function, realizing a vertical expansion in the single physical pedal control dimension. This allows the operator to seamlessly trigger the intelligent analysis of the background algorithm model with only a single long press of the foot pedal while continuously operating the endoscope handle with both hands. The structured image feature data obtained from the analysis is then accurately and synchronously fused and displayed with the real-time image, reducing the cognitive load and operational cost of human-computer interaction, improving the efficiency of human-computer interaction, enhancing the efficiency, flexibility, and accuracy of image data processing, and improving the user experience.
[0046] In some embodiments, after receiving the foot pedal operation signal sent by the foot pedal input device, the method further includes: Extract the pressing frequency of the foot pedal operation signal, and determine the type of foot pedal operation signal based on the pressing frequency; When the press frequency indicator is single press, the foot pedal operation signal is determined to be a single foot pedal operation signal; When the pressing frequency indicator is two consecutive presses, the foot pedal operation signal is determined to be a double foot pedal operation signal; When the pressing frequency indicator is N consecutive presses, the foot pedal operation signal is determined to be a continuous foot pedal operation signal, where N is a positive integer greater than 2.
[0047] Among them, the pressing frequency refers to the cumulative number of times the operator performs the physical action of pressing down and fully releasing the foot input device within the preset effective judgment time window.
[0048] Optionally, using the system's low-level hardware interface or interrupt handler, an internal timer is started upon the first trigger signal of the foot pedal operation signal, such as a rising or falling edge of the voltage level. This timer has a specific duration, such as one second or 1.5 seconds, as a decision time window. The number of complete pulse signals received within this decision time window is continuously monitored and counted; this number represents the extracted press frequency. Based on the magnitude of the extracted press frequency, the corresponding category of the foot pedal operation signal is determined through a preset logical branch structure.
[0049] Specifically, at the end of the aforementioned determination time window, if the press frequency indication is a single press (meaning the system records only one complete pulse signal), the input event is determined as a single foot pedal operation signal. If the press frequency indication is two consecutive presses (meaning two complete pulse signals are captured consecutively within the time window), the processor determines the input event as a double foot pedal operation signal. If the press frequency indication is N consecutive presses (meaning the operator performs rapid multiple presses, and the system counts three or more complete pulse signals), the input event is determined as a continuous foot pedal operation signal.
[0050] In this embodiment of the invention, by introducing an automatic extraction mechanism for the frequency of pressing and a signal type determination mechanism during the signal reception stage, it is possible to accurately identify and distinguish the different rhythmic foot-pressing actions of the operator. Without adding additional physical hardware pedals, the input instruction set of a single foot pedal device is enriched by utilizing the continuous pressing characteristics in the time dimension. This provides the necessary data foundation for flexibly mapping various different endoscopy business processing logics for different types of foot pedal signals, effectively improving the expansion capability of human-computer interaction.
[0051] In some embodiments, determining the mapping relationship between different foot pedal operation signals and different function commands includes: Obtain the current system's working mode identifier, which indicates the current system's preset working mode. The preset working mode includes at least one of image acquisition mode, video recording mode, and AI analysis mode. Based on the working mode identifier, the mapping relationship between different foot pedal operation signals and different function commands is determined; under different preset working modes, the same type of foot pedal operation signal is mapped to different function commands.
[0052] For example, at the beginning of a bronchoscopy, the system defaults to image acquisition mode, which allows doctors to capture key images at any time; when the doctor performs a double-step operation to start video recording, the system automatically switches to video recording mode, at which point the interface displays the recording status; when the AI analysis results are returned, the system can automatically switch to AI analysis mode, which allows doctors to review the analysis results.
[0053] Optionally, based on the current operating mode identifier, the corresponding mapping rule table is loaded. The mapping rule table defines the correspondence between different types of foot pedal operation signals and function commands.
[0054] Specifically, in different preset working modes, the same type of foot pedal operation signal is mapped to different functional instructions, thereby realizing flexible control with one pedal for multiple uses.
[0055] Optionally, the mapping relationship is stored in the system database in the form of a lookup table. The mapping result of the function instruction can be quickly determined by querying the lookup table entry corresponding to the current working mode identifier.
[0056] In this embodiment of the invention, by obtaining the current system's working mode identifier, which indicates the current preset working mode of the system, and based on the working mode identifier, the mapping relationship between different foot pedal operation signals and different function commands is determined, thereby realizing flexible and multifunctional control of foot pedal operation and improving the user experience.
[0057] In some embodiments, after determining that the foot pedal operation signal is a single foot pedal operation signal, the method further includes: Based on a single foot pedal operation signal and mapping relationship, an image acquisition command is generated; In response to an image acquisition command, extract the current frame image data from the current video stream data; Add a timestamp and / or identification information to the current frame image data, store the current frame image data in the database, and display a message indicating successful image data acquisition on the user interface.
[0058] The image acquisition command is a software-level control command used to trigger still image capture. The current frame image data refers to a high-definition still digital image captured from the high-speed buffer of the current video stream data at the instant the image acquisition command takes effect.
[0059] The timestamp refers to the specific year, month, day, hour, minute, and second when the image acquisition action occurred. Identification information may include the patient's basic identity information, the endoscope device number, or the code of the current operator.
[0060] Optionally, the extracted current frame image data is bound and encapsulated with the corresponding timestamp and identification information, and then persistently stored in a local hard disk database or a cloud server database. Simultaneously, a visual feedback mechanism, such as a text pop-up or status icon, is displayed in a designated area of the user interface via the graphics rendering engine, such as the upper right corner or bottom of the screen, to indicate successful image data acquisition.
[0061] In this embodiment of the invention, by assigning image capture function to a single foot pedal operation signal and binding and persistently storing the captured image with metadata, coupled with intuitive interface feedback, it can help operators quickly retain high-definition images of key lesions without interrupting the main endoscopic operation, thereby improving the standardization of medical record data collection and the efficiency of clinical evidence collection.
[0062] In some embodiments, after determining that the foot pedal operation signal is a double foot pedal operation signal, the method further includes: Based on the dual foot pedal operation signals and mapping relationship, a video recording start command or a video recording stop command is generated. In response to a video recording start command, control the endoscopic device to perform a video recording start operation for the current video stream data; or, in response to a video recording stop command, control the endoscopic device to perform a video recording stop operation for the current video stream data. Displays the video recording status on the user interface; Save the current video stream data within the recording time period as a video file.
[0063] Optionally, a state machine is maintained internally to record the current recording state. Upon receiving a double foot pedal operation signal, this state machine is queried. If the system is currently in a non-recording state, a video recording start command is generated to trigger the underlying multimedia framework to begin continuously writing the current video stream data to the storage medium; if the system is currently in a recording state, a video recording stop command is generated to trigger the multimedia framework to end the current data stream writing operation.
[0064] Optionally, during recording, the video recording status can be dynamically displayed by drawing, for example, a flashing red dot or dynamically timed recording time text in a prominent position on the user interface. Once the video recording stop operation is complete, the consecutive video frames within the recording period are encoded and compressed to generate a video file with a universal format.
[0065] In this embodiment of the invention, by mapping the double-click foot pedal action to a start and stop switch for video recording, a reusable state flip control mechanism is realized, providing operators with an extremely convenient way to retain dynamic images. This ensures that the continuous dynamic images during the internal detection process are completely recorded, while the interface status display effectively avoids the occurrence of missed or incorrect recordings.
[0066] In some embodiments, after determining that the foot pedal operation signal is a continuous foot pedal operation signal, the method further includes: Based on continuous foot pedal operation signals and mapping relationships, keyframe marking instructions are generated. In response to the keyframe marking instruction, the current frame image data is obtained from the current video stream data, the current frame image data is identified as keyframe image data, and timestamps and / or identification information are added to the keyframe image data.
[0067] Keyframe marking commands are control commands used to highlight key images in a continuous video data stream or among numerous acquired images. Keyframe image data typically corresponds to highly valuable images of rare lesions or confirmation images of crucial steps in clinical exploration.
[0068] The method for acquiring the current frame image data is similar to that described for a single snapshot. After identifying it as keyframe image data, a timestamp is added, and specific tag attributes with high priority are written into the identification information, such as machine-readable identification information like a focus marker or an anomaly warning data bit.
[0069] In this embodiment of the invention, by using the relatively special physical action of continuous stomping to trigger keyframe marking, operators can quickly add stars to high-value images without interrupting the overall detection rhythm. This facilitates the rapid retrieval and location of key evidence in the subsequent data review process, effectively improving the overall efficiency of detection summary and data analysis.
[0070] Figure 2 This is a schematic diagram illustrating the process for extracting the pressing duration of the foot pedal operation signal provided by the present invention. Figure 2 As shown, in some embodiments, receiving a foot pedal operation signal sent by a foot pedal input device and extracting the pressing duration of the foot pedal operation signal includes: Step 210: Receive the foot pedal operation signal sent by the foot pedal input device. The foot pedal operation signal is received at the current moment.
[0071] The current moment refers to the instant when the foot switch level changes or is triggered. The reception of the foot switch operation signal relies on the system's real-time interrupt response mechanism or high-frequency status polling mechanism to ensure that the signal can be captured with low latency.
[0072] Step 220: Calculate the time interval between the foot pedal operation signal and the historical foot pedal operation signal. The historical foot pedal operation signal was received at the previous moment before the current moment and was sent by the foot pedal input device.
[0073] The historical foot pedal operation signal refers to the last foot pedal action pulse generated immediately before the current moment on the system timeline and recorded in system memory. The time interval is the difference between the timestamp corresponding to the current moment and the timestamp recorded when the historical foot pedal operation signal was generated. This time interval value is accurately obtained by reading the historical timestamp data persistently stored in the cache and performing a mathematical difference operation with the currently obtained moment. In a sterile operating environment, doctors wear special shoe covers or surgical slippers, and their foot pedal rhythm differs from that in a regular office environment. Therefore, accurate time interval calculation helps distinguish between intentional operation instructions and unintentional touches.
[0074] Step 230: If the time interval is greater than or equal to the preset anti-shake time threshold, the foot pedal operation signal is determined as a valid trigger signal, and the pressing duration of the foot pedal operation signal is extracted.
[0075] The anti-shake time threshold is a pre-set, extremely short time limit parameter, such as fifty or one hundred milliseconds, to filter out vibrations in the mechanical contact springs of the foot pedal device during the moment of closure or unintentional slight touches by the operator. Considering that doctors may experience slight foot movements due to body positioning while concentrating on operating the endoscope, this anti-shake mechanism effectively avoids erroneous commands caused by accidental touches, ensuring that critical operations such as AI-assisted diagnosis are only triggered when the doctor has a clear intention.
[0076] Optionally, the calculated time interval is compared with the anti-shake time threshold. If the time interval is greater than or equal to the anti-shake time threshold, it indicates that there is a reasonable interval between two adjacent level changes that conforms to the normal physical movement of the human body. This is determined to be a valid trigger signal, indicating that it represents the operator's genuine intention to step on the pedal. The signal is then allowed to proceed, triggering the subsequent action of extracting the foot pedal operation signal and the pressing duration. Conversely, if the time interval is less than the anti-shake time threshold, it is determined to be an invalid noise pulse signal and discarded directly. This software anti-shake method does not require additional hardware filtering circuitry and is particularly suitable for the requirements of cleaning, disinfection, and aseptic maintenance of medical equipment.
[0077] In this embodiment of the invention, by introducing a timestamp-based anti-shake determination mechanism at the initial stage of signal reception, high-frequency glitches caused by equipment aging or minor accidental touches can be efficiently intercepted and filtered out using pure software logic without adding additional hardware anti-shake filtering circuits. This ensures from the data source that all foot pedal commands entering the subsequent core business processing flow have clear and genuine interaction intentions, effectively avoiding false triggering of the endoscope system due to invalid signals or frequent occupation of computing resources, improving the accuracy of underlying command recognition and the stability of overall operation. It can fully adapt to the needs of doctors operating endoscopes with both hands and engaging in non-contact interaction in medical scenarios.
[0078] In some embodiments, the structured image feature data is rendered and displayed synchronously with the current video stream data, including: Based on location information, a visual identifier box is generated and overlaid at the corresponding position in the current video stream data; In the associated area of the visual identifier box, feature data and feature confidence are rendered and displayed; In the preset status bar area of the user interface, auxiliary prompts and operation information for the current image sequence data are dynamically updated.
[0079] Among them, visual identifiers refer to the geometric boundaries used to highlight and outline the target area of interest in real-time video footage, such as a striking rectangular frame, a circular frame, or a polygonal outline dynamically drawn along the edge of the lesion.
[0080] Optionally, based on the two-dimensional coordinate parameters in the obtained position information, the above-mentioned graphic boundary is calculated and drawn in real time in the upper transparent layer of the video display buffer, so that it visually and accurately covers the corresponding pixel position of the target organization in the underlying video stream, thereby achieving visual locking of the target.
[0081] The associated area refers to the preset pixel space that is close to the edge of the visual identifier or has a clear directional relationship with it, such as the blank area directly above the identifier, the lower right corner, or the side blank area connected by a guide line.
[0082] Optionally, using a graphics rendering engine, feature data describing target attributes, such as the organization's specific size, prediction category, and feature confidence levels representing the model's judgment confidence, such as percentage values, are plotted in real time in the form of text with a specific color and font within the associated area.
[0083] The preset status bar area refers to a fixed information display panel in the user interface that is independent of the main video playback window, such as the top sidebar or bottom floating bar of the screen. Auxiliary prompts may include the model's current running status, the cumulative number of detected abnormal targets, or high-risk warning text. Operational information may include indications of the current foot pedal input device's operating mode, descriptions of the current foot pedal combo function, or interactive prompts suggesting the next step.
[0084] Optionally, the above-mentioned prompt text and status icons, as well as the current image sequence data and the trigger status of the foot pedal operation signal, are refreshed synchronously.
[0085] In this embodiment of the invention, by overlaying visual identifier boxes at corresponding spatial positions in the video frame and intuitively displaying feature data and confidence levels in the associated areas, and supplemented by real-time updates of global status bar information, a layered augmented reality interface display mechanism is constructed. This allows operators to quickly obtain spatial positioning and attribute identification assistance provided by the algorithm model while keeping a close eye on the real-time endoscope image, effectively avoiding operational distraction caused by shifting gaze. At the same time, the synchronous update of global status information further improves the information completeness of the human-computer interaction interface and the overall safety of operation.
[0086] In some embodiments, a target feature analysis trigger command is generated based on the long-press foot pedal operation signal and the mapping relationship, including: Obtain the current system's operating status identifier, which indicates whether the current system is under preset mode restrictions. When the working status indicator indicates that the current system is not in the preset mode restriction state, a target feature analysis trigger command is generated based on the long press foot pedal operation signal and mapping relationship.
[0087] The operating status identifier is a digital code or enumerated status bit used to characterize the current operating stage of the system or the system resource usage. Optionally, the current system operating status identifier includes at least the operating status of the endoscope equipment.
[0088] Optionally, the working status identifier can be queried in real time by calling the system state machine application interface or reading the global status register. The preset mode restriction state refers to specific working scenarios that are not allowed or suitable for high-computational image feature analysis, which are pre-configured by the system. For example, the endoscope equipment is in a high-load continuous video writing state, the image sensor is in white balance calibration state, or the image is in an offline viewing state where the image is frozen.
[0089] Specifically, after obtaining the working status identifier, it is logically matched against a pre-stored constraint status table in the system memory. If the working status identifier indicates that the current system is not in a preset mode constraint state, it means that the current system's computing resources and video stream input state meet the prerequisites for calling the pre-trained image feature analysis model. Subsequently, based on the long-press foot pedal operation signal, a target feature analysis trigger command is generated and issued. Conversely, if the current system is detected to be in a preset mode constraint state, the foot pedal trigger action is actively intercepted, and a corresponding mode mutual exclusion prompt can be displayed on the user interface to avoid conflicts in underlying commands.
[0090] In this embodiment of the invention, by introducing a mechanism for obtaining the working status identifier and verifying the mode restriction status before generating the target feature analysis trigger instruction, a global logical security lock is added to the long press operation of the foot pedal. This effectively prevents the processing task conflict or system computing power overload crash caused by the operator habitually pressing the foot pedal when the system is in a special working mode or high load operation phase. This improves the operational stability of the endoscopic image data processing system and the security of software and hardware scheduling in complex clinical application scenarios.
[0091] The endoscopic image data processing system based on foot pedal signals provided by the present invention will be described below. The endoscopic image data processing system based on foot pedal signals described below can be referred to in correspondence with the endoscopic image data processing method based on foot pedal signals described above.
[0092] Figure 3 This is a schematic diagram of the endoscopic image data processing system based on foot pedal signals provided by the present invention. Figure 3 As shown, the present invention also provides an endoscopic image data processing system based on foot pedal signals. The endoscopic image data processing system 300 based on foot pedal signals includes: The receiving unit 310 is used to receive the foot pedal operation signal sent by the foot pedal input device; The determining unit 320 is used to extract the pressing duration of the foot pedal operation signal. If the pressing duration is greater than or equal to a preset long press time threshold, the foot pedal operation signal is determined to be a long press foot pedal operation signal. The generation unit 330 is used to determine the mapping relationship between different foot pedal operation signals and different function commands, and to generate target feature analysis trigger commands based on the long press foot pedal operation signal and the mapping relationship. Data processing unit 340 is used to respond to a target feature analysis trigger command by inputting the current image sequence data in the current video stream data output by the endoscope device into a pre-trained image feature analysis model to obtain structured image feature data output by the image feature analysis model; the current image sequence data includes multiple video frame image data within the current time window; Display unit 350 is used to synchronously render and display structured image feature data and current video stream data on the user interface; The system also includes a signal classification unit, which, after the receiving unit 310 receives the foot pedal operation signal sent by the foot pedal input device, extracts the pressing frequency of the foot pedal operation signal and determines the type of the foot pedal operation signal based on the pressing frequency; if the pressing frequency indication is a single press, the foot pedal operation signal is determined to be a single foot pedal operation signal; if the pressing frequency indication is two consecutive presses, the foot pedal operation signal is determined to be a double foot pedal operation signal; if the pressing frequency indication is N consecutive presses, the foot pedal operation signal is determined to be a continuous foot pedal operation signal, where N is a positive integer greater than 2.
[0093] Optionally, the mapping relationship between different foot pedal operation signals and different function commands is determined, including: Obtain the current system's working mode identifier, which indicates the current system's preset working mode. The preset working mode includes at least one of image acquisition mode, video recording mode, and AI analysis mode. Based on the working mode identifier, the mapping relationship between different foot pedal operation signals and different function commands is determined; under different preset working modes, the same type of foot pedal operation signal is mapped to different function commands.
[0094] Optionally, the system further includes an image acquisition unit, which is used for: After determining that the foot pedal operation signal is a single foot pedal operation signal, an image acquisition command is generated based on the single foot pedal operation signal and the mapping relationship. In response to an image acquisition command, extract the current frame image data from the current video stream data; Add a timestamp and / or identification information to the current frame image data, store the current frame image data in the database, and display a message indicating successful image data acquisition on the user interface.
[0095] Optionally, the system further includes a video recording control unit, which is used for: After determining that the foot pedal operation signal is a double foot pedal operation signal, a video recording start command or a video recording stop command is generated based on the double foot pedal operation signal and the mapping relationship. In response to a video recording start command, control the endoscopic device to perform a video recording start operation for the current video stream data; or, in response to a video recording stop command, control the endoscopic device to perform a video recording stop operation for the current video stream data. Displays the video recording status on the user interface; Save the current video stream data within the recording time period as a video file.
[0096] Optionally, the system further includes a keyframe marking unit, which is used for: After determining that the foot pedal operation signal is a continuous foot pedal operation signal, a keyframe marking instruction is generated based on the continuous foot pedal operation signal and the mapping relationship. In response to the keyframe marking instruction, the current frame image data is obtained from the current video stream data, the current frame image data is identified as keyframe image data, and timestamps and / or identification information are added to the keyframe image data.
[0097] Optionally, the foot pedal operation signal sent by the foot pedal input device is received, and the pressing duration of the foot pedal operation signal is extracted, including: Receives foot pedal operation signals sent by the foot pedal input device, and the foot pedal operation signals are received at the current moment; Calculate the time interval between the foot pedal operation signal and the historical foot pedal operation signal. The historical foot pedal operation signal was received at the previous moment of the current moment and sent by the foot pedal input device. If the time interval is greater than or equal to the preset anti-shake time threshold, the foot pedal operation signal is determined as a valid trigger signal, and the pressing duration of the foot pedal operation signal is extracted.
[0098] Optionally, the structured image feature data includes: feature data of the target region of interest, location information, and feature confidence.
[0099] Optionally, the structured image feature data is rendered and displayed synchronously with the current video stream data, including: Based on location information, a visual identifier box is generated and overlaid at the corresponding position in the current video stream data; In the associated area of the visual identifier box, feature data and feature confidence are rendered and displayed; In the preset status bar area of the user interface, auxiliary prompts and operation information for the current image sequence data are dynamically updated.
[0100] Optionally, based on the long-press foot pedal operation signal and mapping relationship, a target feature analysis trigger command is generated, including: Obtain the current system's working status identifier and mapping relationship. The working status identifier is used to indicate whether the current system is in a preset mode restriction state. When the working status indicator indicates that the current system is not in the preset mode restriction state, a target feature analysis trigger command is generated based on the long press foot pedal operation signal.
[0101] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention, such as... Figure 4 As shown, the electronic device may include a processor 410, a communication interface 420, a memory 430, and a communication bus 440, wherein the processor 410, the communication interface 420, and the memory 430 communicate with each other through the communication bus 440. The processor 410 can call logical instructions in the memory 430 to execute an endoscopic image data processing method based on foot pedal signals. This method includes: receiving a foot pedal operation signal sent by a foot pedal input device; extracting the pressing duration of the foot pedal operation signal, and determining the foot pedal operation signal as a long-press foot pedal operation signal if the pressing duration is greater than or equal to a preset long-press time threshold; determining the mapping relationship between different foot pedal operation signals and different functional instructions, and generating a target feature analysis trigger instruction based on the long-press foot pedal operation signal and the mapping relationship; and responding to the target feature analysis trigger instruction by inputting the current image sequence data from the current video stream data output by the endoscopic device into a pre-trained image feature analysis module. The system obtains structured image feature data output by the image feature analysis model; on the user interface, the structured image feature data and the current video stream data are synchronously rendered and displayed; after receiving the foot pedal operation signal sent by the foot pedal input device, the system further includes: extracting the pressing frequency of the foot pedal operation signal, and determining the type of the foot pedal operation signal based on the pressing frequency; if the pressing frequency indication is a single press, the foot pedal operation signal is determined to be a single foot pedal operation signal; if the pressing frequency indication is two consecutive presses, the foot pedal operation signal is determined to be a double foot pedal operation signal; if the pressing frequency indication is N consecutive presses, the foot pedal operation signal is determined to be a continuous foot pedal operation signal, where N is a positive integer greater than 2.
[0102] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0103] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0104] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0105] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for processing endoscopic image data based on foot pedal signals, characterized in that, include: Receive foot pedal operation signals sent by the foot pedal input device; Extract the pressing duration of the foot pedal operation signal. If the pressing duration is greater than or equal to a preset long press time threshold, determine that the foot pedal operation signal is a long press foot pedal operation signal. Determine the mapping relationship between different foot pedal operation signals and different function commands, and generate a target feature analysis trigger command based on the long press foot pedal operation signal and the mapping relationship; In response to the target feature analysis trigger command, the current image sequence data in the current video stream data output by the endoscope device is input into a pre-trained image feature analysis model to obtain structured image feature data output by the image feature analysis model; the current image sequence data includes multiple video frame image data within the current time window; On the user interface, the structured image feature data and the current video stream data are rendered and displayed synchronously. After receiving the foot pedal operation signal sent by the foot pedal input device, the method further includes: Extract the pressing frequency of the foot pedal operation signal, and determine the type of the foot pedal operation signal based on the pressing frequency; When the pressing frequency indicator is a single press, the foot pedal operation signal is determined to be a single foot pedal operation signal; If the pressing frequency indication is two consecutive presses, the foot pedal operation signal is determined to be a double foot pedal operation signal; When the pressing frequency indication is N consecutive pressings, the foot pedal operation signal is determined to be a continuous foot pedal operation signal, where N is a positive integer greater than 2.
2. The endoscopic image data processing method based on foot pedal signals according to claim 1, characterized in that, Determining the mapping relationship between different foot pedal operation signals and different function commands includes: Obtain the current system's operating mode identifier, which indicates the current system's preset operating mode. The preset operating mode includes at least one of image acquisition mode, video recording mode, and AI analysis mode. Based on the operating mode identifier, the mapping relationship between different foot pedal operation signals and different function commands is determined; under different preset operating modes, the same type of foot pedal operation signal is mapped to different function commands.
3. The endoscopic image data processing method based on foot pedal signals according to claim 1, characterized in that, After determining that the foot pedal operation signal is a single foot pedal operation signal, the method further includes: Based on the single foot pedal operation signal and the mapping relationship, an image acquisition command is generated; In response to the image acquisition command, extract the current frame image data from the current video stream data; Add a timestamp and / or identification information to the current frame image data, store the current frame image data in the database, and display a message indicating successful image data acquisition on the user interface.
4. The endoscopic image data processing method based on foot pedal signals according to claim 1, characterized in that, After determining that the foot pedal operation signal is a double foot pedal operation signal, the method further includes: Based on the dual foot pedal operation signals and the mapping relationship, a video recording start command or a video recording stop command is generated. In response to the video recording start command, the endoscopic device is controlled to perform a video recording start operation for the current video stream data; or, in response to the video recording stop command, the endoscopic device is controlled to perform a video recording stop operation for the current video stream data. The video recording status is displayed on the user interface. Save the current video stream data within the recording time period as a video file.
5. The endoscopic image data processing method based on foot pedal signals according to claim 1, characterized in that, After determining that the foot pedal operation signal is a continuous foot pedal operation signal, the method further includes: Based on the continuous foot pedal operation signals and the mapping relationship, a keyframe marking instruction is generated; In response to the keyframe marking instruction, the current frame image data is obtained from the current video stream data, the current frame image data is determined as keyframe image data, and a timestamp and / or identification information is added to the keyframe image data.
6. The endoscopic image data processing method based on foot pedal signals according to claim 1, characterized in that, The process of receiving a foot pedal operation signal from a foot pedal input device and extracting the pressing duration of the foot pedal operation signal includes: Receives a foot pedal operation signal sent by the foot pedal input device, wherein the foot pedal operation signal is received at the current moment; Calculate the time interval between the foot pedal operation signal and the historical foot pedal operation signal, wherein the historical foot pedal operation signal was received at the previous time before the current time and sent by the foot pedal input device; If the time interval is determined to be greater than or equal to a preset anti-shake time threshold, the foot pedal operation signal is determined to be a valid trigger signal, and the pressing duration of the foot pedal operation signal is extracted.
7. The endoscopic image data processing method based on foot pedal signals according to claim 1, characterized in that, The structured image feature data includes: feature data of the target region of interest, location information, and feature confidence level.
8. The endoscopic image data processing method based on foot pedal signals according to claim 7, characterized in that, The step of synchronously rendering and displaying the structured image feature data and the current video stream data includes: Based on the location information, a visual identifier box is generated and overlaid at the corresponding position of the current video stream data; In the associated area of the visual identifier box, the feature data and the feature confidence level are rendered and displayed; In the preset status bar area of the user interface, auxiliary prompts and operation information for the current image sequence data are dynamically updated.
9. The endoscopic image data processing method based on foot pedal signals according to claim 2, characterized in that, The step of generating a target feature analysis trigger command based on the long-press foot pedal operation signal and the mapping relationship includes: Obtain the current system's operating status identifier, which indicates whether the current system is in a preset mode restriction state; When the working status indicator indicates that the current system is not in a preset mode restriction state, a target feature analysis trigger command is generated based on the long press foot pedal operation signal and the mapping relationship.
10. An endoscopic image data processing system based on foot pedal signals, characterized in that, include: The receiving unit is used to receive foot pedal operation signals sent by the foot pedal input device; The determining unit is used to extract the pressing duration of the foot pedal operation signal, and determine the foot pedal operation signal as a long press foot pedal operation signal if the pressing duration is greater than or equal to a preset long press time threshold. The generation unit is used to determine the mapping relationship between different foot pedal operation signals and different function commands, and to generate a target feature analysis trigger command based on the long press foot pedal operation signal and the mapping relationship. The data processing unit is configured to, in response to the target feature analysis trigger command, input the current image sequence data from the current video stream data output by the endoscopic device into a pre-trained image feature analysis model to obtain structured image feature data output by the image feature analysis model; the current image sequence data includes multiple video frame image data within the current time window; The display unit is used to synchronously render and display the structured image feature data and the current video stream data on the user interface. The system further includes a signal classification unit, used to extract the pressing frequency of the foot pedal operation signal after the receiving unit receives the foot pedal operation signal sent by the foot pedal input device, and determine the type of the foot pedal operation signal based on the pressing frequency; if the pressing frequency indicates a single press, the foot pedal operation signal is determined to be a single foot pedal operation signal; if the pressing frequency indicates two consecutive presses, the foot pedal operation signal is determined to be a double foot pedal operation signal; if the pressing frequency indicates N consecutive presses, the foot pedal operation signal is determined to be a continuous foot pedal operation signal, where N is a positive integer greater than 2.