Remote medical teaching platform-oriented multi-terminal collaborative interaction method and system, and medium

By using signal desensitization and fusion and multimodal annotation technology in the collaborative platform, the problem of lack of collaboration between experts and students in remote medical teaching has been solved, realizing efficient teaching flow generation and a rich learning experience for students, thereby improving the quality and efficiency of remote medical teaching.

CN121397285APending Publication Date: 2026-01-23SCANMED CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511346581.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-19
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

In telemedicine teaching platforms, the annotations and explanations by experts are independent of those by students, lacking close collaboration. This results in a monotonous and uninterrupted learning experience for students, reducing learning efficiency and teaching effectiveness.

Method used

By using a collaborative platform to perform time-synchronized signal desensitization and fusion, a desensitized and fused data packet stream is generated and layered into expert-level and student-level code streams. Experts perform multimodal collaborative annotation to generate structured annotation data packets. Combined with spatially bound annotation dynamic rendering, a hybrid output teaching stream is output, supporting split-screen perspective switching on the student side.

Benefits of technology

It enabled efficient collaboration between experts and trainees, enhanced the diversity and comprehensibility of teaching content, improved the trainees' learning experience and interactivity, and ensured data security and patient privacy protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121397285A_ABST
    Figure CN121397285A_ABST
Patent Text Reader

Abstract

The invention provides a remote medical teaching platform-oriented multi-terminal collaborative interaction method and system and a medium, and relates to the technical field of intelligent medical systems.The method comprises the steps that after a collaborative middle platform receives an original operation video stream and vital sign data, a desensitization fusion data packet stream is obtained through signal desensitization fusion; after the hierarchical coding is compressed into an expert-level code stream and a student-level code stream, the code streams are distributed to M expert ends and a plurality of student ends by adopting two-dimensional transmission quality parameter mapping; m expert ends carry out multi-modal collaborative labeling and output M groups of structured labeling data packets; dynamic rendering of space binding labels is executed, after mixed output teaching streams are obtained, the mixed output teaching streams are distributed to a plurality of student ends, and the student ends support split-screen view angle switching. The technical problems that in remote medical teaching in the prior art, labeling and explanation of an expert end are often independent of a student end, so that learning experience of the student is single, interaction is lacked, and learning efficiency and teaching effect are reduced are solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the technical field of intelligent medical systems, in particular to a multi-terminal collaborative interaction method and system for a remote medical teaching platform and a medium. BACKGROUND

[0002] With the continuous development of intelligent medical systems, remote medical teaching platforms have gradually become an important part of medical education. Traditional medical teaching relies on face-to-face practical teaching, especially in the guidance and training of surgical procedures. Through video live broadcast, recording, and online assistance by experts, remote medical teaching provides a convenient and flexible learning method for students. However, in the existing remote medical teaching platform, the annotation and explanation of the expert terminal are often independent of the student terminal. Students can only passively receive the expert annotation content, lacking real-time interaction and feedback. There is a lack of collaborative interaction between experts and students. The annotated content is usually static graphics or voice, without more levels and rich teaching methods, resulting in a single learning experience and lack of interaction for students, thereby reducing learning efficiency and teaching effectiveness. SUMMARY

[0003] The application provides a multi-terminal collaborative interaction method and system for a remote medical teaching platform and a medium, aiming to solve the technical problem that in the existing remote medical teaching, the annotation and explanation of the expert terminal are often independent of the student terminal, and experts and students lack close collaboration, resulting in a single learning experience and lack of interaction for students, thereby reducing learning efficiency and teaching effectiveness.

[0004] The first aspect of the application provides a multi-terminal collaborative interaction method for a remote medical teaching platform. The method comprises: after receiving the original surgical video stream and vital sign data returned by the operating room terminal, the collaborative middle station obtains the desensitization fusion data packet stream by performing time sequence synchronization signal desensitization fusion; after layering and encoding compression of the desensitization fusion data packet stream into expert-level code stream and student-level code stream, the expert-level code stream and student-level code stream are distributed to M expert terminals and multiple student terminals using double-dimensional transmission quality parameter mapping; after synchronously receiving the expert-level code stream, the M expert terminals perform multi-modal collaborative annotation of the expert-level code stream based on pre-set annotation permission allocation, and output M groups of structured annotation data packets; after performing spatially bound annotation dynamic rendering on the M groups of structured annotation data packets, a mixed output teaching stream is obtained, and the mixed output teaching stream is distributed to the multiple student terminals, wherein the mixed output teaching stream and the student-level code stream support split-screen perspective switching at the student terminal.

[0005] In a second aspect, the application discloses a multi-terminal collaborative interaction system for a remote medical teaching platform, which is used for the multi-terminal collaborative interaction method for the remote medical teaching platform. The system comprises a signal desensitization fusion module, which is used for obtaining a desensitization fusion data packet stream by performing time sequence synchronization signal desensitization fusion after receiving an original surgery video stream and vital sign data returned by an operating room terminal in cooperation with a middle station; a mapping and distribution module, which is used for performing hierarchical encoding compression on the desensitization fusion data packet stream to obtain an expert-level code stream and a student-level code stream, and then mapping and distributing the expert-level code stream and the student-level code stream to M expert terminals and a plurality of student terminals by using double-dimensional transmission quality parameters; a collaborative labeling module, which is used for performing multi-modal collaborative labeling on the expert-level code stream based on preset labeling permission allocation after the M expert terminals synchronously receive the expert-level code stream, and outputting M groups of structured labeling data packets; and a teaching stream distribution module, which is used for performing spatially bound labeling dynamic rendering on the M groups of structured labeling data packets to obtain a mixed output teaching stream, and then distributing the mixed output teaching stream to the plurality of student terminals, wherein the mixed output teaching stream and the student-level code stream support split-screen perspective switching at the student terminals.

[0006] In a third aspect, the application discloses a storage medium having a computer program stored thereon, and the computer program is executed by a processor to implement the steps of the multi-terminal collaborative interaction method for the remote medical teaching platform in the first aspect.

[0007] The one or more technical solutions provided in the application have at least the following beneficial effects: The original operation video stream and vital sign data transmitted back from the operating room end are received by the coordination platform, and through time-synchronous signal desensitization fusion, the privacy information of the patient can be effectively protected, and the time-synchronous data is ensured, which ensures the seamless docking of real-time video stream and vital sign data in the operation process, and eliminates the potential risk of privacy leakage through desensitization processing, so that teaching and analysis data can be safely shared; through hierarchical encoding compression of the desensitized fusion data packet stream into expert-level stream and student-level stream, and using double-dimensional transmission quality parameter mapping, the data transmission quality requirements of different terminals can be optimized, the expert end receives high-definition expert-level video stream, and the student end receives appropriate quality video stream according to the bandwidth and equipment requirements, while ensuring the stability and quality of transmission, which effectively balances the transmission quality between different ports, ensuring that experts and students can smoothly receive suitable teaching video content under different network conditions, thereby improving the user experience and adaptability of the remote medical teaching platform; under the cooperation of multiple experts, based on the preset annotation permission, experts can perform multi-modal collaborative annotation on the operation video stream, and generate structured annotation data packets, which ensures the comprehensive analysis of the operation process by experts, realizes the efficient cooperation of the expert end, and the annotation content is output in a structured manner, which is convenient for subsequent rendering and teaching flow production, and enhances the diversity and intelligibility of teaching content; through spatial binding dynamic rendering of the M group of structured annotation data packets, a hybrid output teaching flow with high accuracy standard and voice explanation is generated, and students can obtain comprehensive teaching content of the operation video and expert annotation, so that the learning experience of students is more rich and intuitive, in addition, the student end supports split-screen view switching, which can not only watch real-time operation video transmitted back by the expert end, but also receive fused real-time teaching video, which enhances the learning experience of students, so that students can not only observe the operation process in real time, but also check the annotation and explanation of experts in the teaching video, thereby improving the individualization and interactivity of learning. In general, the intelligent medical system realizes intelligent medical teaching, stronger interactivity, ensures data security and patient privacy protection, and greatly improves the quality and efficiency of remote medical teaching.

[0008] The above description is only a summary of the technical solutions of the present application, in order to more clearly understand the technical means of the present application, the specific embodiments of the present application can be implemented according to the content of the description, and in order to make the above and other purposes, characteristics and advantages of the present application more obvious and easy to understand, the following specific embodiments of the present application are described. BRIEF DESCRIPTION OF DRAWINGS

[0009] Figure 1 The multi-end collaborative interaction method flowchart for the remote medical teaching platform provided by the embodiments of the present application.

[0010] Figure 2 A multi-terminal collaborative interaction system structure schematic diagram for a remote medical teaching platform is provided in the embodiments of the present application.

[0011] Label explanation: signal desensitization fusion module 10, mapping distribution module 20, collaborative annotation module 30, teaching flow distribution module 40. DETAILED DESCRIPTION

[0012] The embodiments of the present application provide a multi-terminal collaborative interaction method, system and medium for a remote medical teaching platform, which solves the technical problem that in the remote medical teaching of the prior art, the annotation and explanation of the expert end are often independent of the student end, the expert and the student lack close cooperation, the learning experience of the student is single and lacks interaction, and the learning efficiency and teaching effect are reduced.

[0013] After introducing the basic principles of the present application, various non-limiting embodiments of the present application will be specifically introduced in combination with the drawings of the specification. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.

[0014] Embodiment one, as shown in the embodiments of the present application, a multi-terminal collaborative interaction method for a remote medical teaching platform is provided, the method comprises: Figure 1 After receiving the original surgery video stream and vital sign data returned by the operating room end, the collaborative middle station obtains the desensitization fusion data packet stream by performing time sequence synchronization signal desensitization fusion.

[0015] The collaborative middle station receives the original surgery video stream and vital sign data from the operating room end, wherein the original surgery video stream is a real-time recorded surgery process video, which contains visual information about the surgery details of the patient and the surgery environment; the vital sign data is derived from the medical equipment in the operating room, which provides real-time vital sign information about the patient, such as heart rate, blood pressure, body temperature, etc.

[0016] ​The collaborative middle station synchronizes the two kinds of data in time, ensuring that the original surgery video stream and vital sign data can be accurately aligned. Since the timestamps of the video stream and vital sign data may not match completely, a time synchronization operation needs to be performed to ensure that they can be processed in the same time sequence. After time synchronization, desensitization processing is performed. The original surgery video stream may contain sensitive information such as the patient's face and private parts, and the vital sign data may contain private information such as patient ID that can identify individuals. By applying desensitization algorithms such as Gaussian blur, data encryption, etc., all sensitive information is removed or encrypted, thereby protecting the privacy of the patient. After desensitization processing, the collaborative middle station fuses the original surgery video stream and vital sign data. The fusion process includes formatting or packaging the two kinds of data so that they can be transmitted as a whole stream. Finally, a desensitized fusion data packet stream is output.

[0017] After the desensitized fusion data packet stream is hierarchically encoded and compressed into an expert-level code stream and a student-level code stream, the expert-level code stream and the student-level code stream are distributed to M expert terminals and multiple student terminals using a dual-dimensional transmission quality parameter mapping.

[0018] The desensitized fusion data packet stream is hierarchically encoded and compressed into two code streams. The expert-level code stream is a high-quality video stream that is mainly transmitted to the expert terminal. Since the expert terminal usually needs high-definition video for diagnosis and analysis, the expert-level code stream needs to have high resolution and low delay. The student-level code stream is a relatively low-quality video stream that is transmitted to the student terminal. Since the task of the student terminal is to learn the surgery process and does not require high-definition pictures, the student-level code stream can have low resolution and high delay.

[0019] The dual-dimensional transmission quality parameters are set according to the transmission quality requirements of the expert terminal and the student terminal. For example, the expert terminal is set to 4K resolution to ensure that the expert terminal can see high-definition pictures suitable for precise analysis of the surgery process with a delay of less than or equal to 200 ms, ensuring that the expert can see the details of the surgery process in real time and not be affected by too much delay when making decisions. The student terminal is set to 1080p resolution, which is suitable for students to learn and observe, and does not need the same high definition as the expert terminal. The delay is less than or equal to 500 ms, ensuring that the student terminal can receive the teaching video stream within a reasonable delay to meet the real-time learning needs.

[0020] Through dual-dimensional transmission quality parameter mapping, the hierarchically encoded expert-level code stream and student-level code stream are distributed to M expert terminals and multiple student terminals. The two kinds of code streams are transmitted according to different quality requirements to ensure that each port can receive a video stream that meets its needs.

[0021] The M expert terminals perform multi-modal collaborative labeling on the expert-level code stream based on a preset labeling permission distribution after synchronously receiving the expert-level code stream, and output M groups of structured labeling data packets.

[0022] M expert terminals synchronously receive expert-level code streams, and each expert terminal performs labeling operations according to a preset labeling permission distribution. The preset labeling permission includes different permission levels, such as whether to allow voice labeling, graphic labeling, video marking, etc. The permissions of each expert terminal can be different, and the specific distribution is based on the role and task requirements. For example, some experts can perform in-depth graphic labeling, such as marking the surgical site, while other experts can only perform simple voice labeling. Each expert terminal performs multi-modal collaborative labeling on the expert-level code stream, which means that the labeling process is not limited to visual labeling, but can also combine voice, graphics, and other modalities. Each expert terminal generates a structured labeling data packet through the above labeling behavior. These structured labeling data packets not only contain timestamp information, but also include structured information such as labeling type, labeling position, and labeling content, which can be used later.

[0023] The M groups of structured labeling data packets are subjected to spatially bound labeling dynamic rendering to obtain a mixed output teaching stream, and the mixed output teaching stream is distributed to the multiple student terminals, wherein the mixed output teaching stream and the student-level code stream support split-screen view switching at the student terminal.

[0024] The M groups of structured labeling data packets are subjected to spatially bound labeling dynamic rendering, which includes dynamically rendering the information labeled by different expert terminals to the surgical video stream, and maintaining the spatial consistency of these labels and the original video, that is, the labeling position should correspond to the actual position in the video image. After rendering, all labels and video content are synthesized and output as a mixed output teaching stream, which combines the collaborative labeling information of multiple expert terminals and has high educational value, helping students understand the surgical steps and important details.

[0025] The mixed output teaching stream is distributed to the multiple student terminals, allowing students to watch the teaching video labeled by multiple experts in real time. At the student terminal, the mixed output teaching stream and the previous student-level code stream support split-screen view switching, and students can switch between different views according to their needs and learning progress to flexibly choose to watch different content, for example, students can choose to watch the video and the expert's labeling commentary at the same time to achieve multi-angle learning and interaction.

[0026] Further, after receiving the original surgical video stream and vital sign data from the operating room terminal, the collaborative middleware obtains a desensitization fusion data packet stream by performing time sequence synchronization signal desensitization fusion, and the method comprises: The original surgery video stream is frame-level parsed to obtain a surgery video frame sequence, and the surgery video frame sequence is visually desensitized based on organ feature recognition to obtain a desensitized video frame sequence; the vital sign data is sliced based on sampling points to obtain a vital sign slice sequence, and the vital sign slice sequence is subjected to PII anonymization processing to obtain a desensitized vital sign slice sequence; the desensitized vital sign slice sequence is frame-level aligned to the desensitized video frame sequence through timestamp matching, and data correlation packaging is performed for time series binding, and the desensitized fusion data packet stream is output.

[0027] Frame-level parsing refers to decomposing a complete original surgery video stream into individual frames, each frame being a static image representing a time point in the video. The parsing process includes extracting each frame image to obtain a surgery video frame sequence. Organ feature recognition is achieved through image recognition technology, such as convolutional neural networks, to identify human organs or sensitive areas in video frames, such as eyes, faces, private parts, etc. This process relies on deep learning algorithms for target detection and segmentation of video frames to identify sensitive areas that need to be desensitized. Visual desensitization refers to applying desensitization operations to the identified sensitive areas to protect patient privacy, such as using Gaussian blur to blur sensitive areas so that they cannot be clearly identified. Ultimately, by performing visual desensitization on each frame, a desensitized video frame sequence is obtained to ensure that sensitive information is not leaked.

[0028] Vital sign data is a continuous time series data. By sampling point slicing, the data is cut into multiple data slices according to a preset time window or sampling interval, with each data slice corresponding to vital sign data for a specific time period. Combining multiple data slices results in a vital sign slice sequence.

[0029] PII (Personally Identifiable Information) anonymization processing is to protect patient privacy by removing all sensitive information that may expose patient identity, including direct identifiers such as patient ID, name, address, etc., which can uniquely identify a person; quasi-identifiers such as age, gender, birth date, etc., which, although not individually completely identifying a person, can help identify individuals in specific contexts. By generalizing sensitive information, such as generalizing specific age to age range or generalizing birth date to month or year, each record shares at least K other records with the same characteristics, thus avoiding the leakage of individual identity. Finally, a K-anonymity dataset is generated, output as a desensitized vital sign slice sequence, which ensures that even if the data is leaked, it cannot be traced back to a specific patient.

[0030] Although both vital sign data and surgical video data are based on the same surgical procedure, their collection frequencies and timestamps can differ. By comparing the timestamps of video frames and the timestamps of vital sign data, the vital sign data slice corresponding to each video frame can be determined. At this point, the timestamps of vital sign data and video frames are accurately matched, allowing each video frame and vital sign data to be accurately matched.

[0031] After completing the timestamp matching, the desensitized vital sign slice sequence and the desensitized video frame sequence are associated and packaged in chronological order, meaning that each video frame and its corresponding vital sign data are processed as a whole, ensuring that they remain synchronized during subsequent transmission and analysis. The final output is a desensitized fusion data packet stream containing synchronized surgical video data and vital sign data, with all sensitive information desensitized to ensure data security and privacy.

[0032] Furthermore, the M expert terminals perform multi-modal collaborative annotation on the expert-level code stream based on pre-set annotation permission allocation after synchronously receiving the expert-level code stream, outputting M sets of structured annotation data packets. The method comprises: The M expert terminals are pre-assigned roles and permissions, resulting in M role annotation permissions. The first expert terminal performs annotation tool interaction on the decoded and rendered expert-level code stream based on the first role annotation permission, performs annotation space coordinate binding, and outputs the first set of structured annotation data packets. If the first role annotation permission includes voice annotation permission, the first voice annotation packet is generated through voiceprint separation and text transcription during real-time voice collection. The first voice annotation packet is stored in the first set of structured annotation data packets based on the timestamp.

[0033] M expert terminals refer to multiple experts participating in the surgical teaching process. To manage and allocate the responsibilities and permissions of each expert, the roles and permissions of these experts are first pre-assigned, that is, before each expert participates in the annotation process, they are assigned different annotation permissions based on factors such as their identity, experience, and task requirements. For example, some experts can only perform graphical annotation, some experts can also perform voice annotation, and other experts have advanced permissions to edit, delete, or modify annotations. Each expert terminal obtains corresponding role annotation permissions based on its role and permission pre-assignment, which defines the types of annotations and the scope of operations that each expert can perform.

[0034] The first expert terminal is any one of the M expert terminals, serving as the current analysis object. The first expert terminal performs annotation operations according to the first role annotation permission. The annotation tool interaction refers to the process in which the expert uses the annotation tool to perform annotation on the surgical video stream, including: a graphical annotation tool, such as drawing an arrow, framing a region, marking a surgical step, etc.; a text tool, used to add a textual explanation in the video to annotate important surgical details; and an audio tool, used to perform annotation through voice recognition (on the premise of having voice annotation permission).

[0035] When performing annotation, the expert binds the annotation to spatial information in the video (i.e., the coordinate position in the video), which means that the content of the annotation must accurately match the corresponding part in the video frame. For example, if the expert marks a surgical site, the annotated arrow or framed region should accurately correspond to the position of the site in the video frame, so that the annotation and the video content can be consistent when the annotation is rendered.

[0036] After the first expert terminal performs annotation tool interaction and completes spatial coordinate binding, a first set of structured annotation data packets is generated, which contains detailed information of the annotation, such as annotation type, annotation position, timestamp, etc.

[0037] If the first role annotation permission includes voice annotation permission, the expert can use voice to perform annotation, rather than being limited to graphical annotation or textual annotation. During voice annotation, the expert's voice input is captured through real-time voice acquisition. After voice acquisition, the voice signal is separated from background noise and the expert's voice signal through voiceprint separation technology, ensuring that the voice annotation is clear and accurate. The collected voice content is converted into text through text transcription technology, facilitating subsequent processing and analysis. This process typically uses a voice recognition algorithm to convert voice signals into text. Based on the voice information after voiceprint separation and text transcription, a first voice annotation packet is generated, which contains the expert's voice commentary content and related timestamp information, so as to be synchronized with other annotation information in the video stream in the future.

[0038] The first voice annotation packet is associated with the timestamp of the video frame. Through the timestamp, the expert's voice annotation can be accurately matched with the annotation content in the video frame, ensuring that the student can hear the expert's voice commentary and see the corresponding annotation at the same time when watching the video. After timestamp association, the first voice annotation packet is stored in the first set of structured annotation data packets as part of the data stream for further use by experts and students.

[0039] Further, after performing spatial binding annotation dynamic rendering on the M sets of structured annotation data packets to obtain a mixed output teaching stream, the mixed output teaching stream is distributed to the plurality of student terminals. The method comprises: extracting M sets of structured annotation data packets, each set containing M speech annotation packets and M sets of spatial coordinate binding information for M graphic annotation layers; performing multi-expert annotation conflict resolution after spatial-temporal mapping alignment of the M sets of spatial coordinate binding information by overlapping the M graphic annotation layers; outputting a mixed graphic annotation layer teaching stream; performing multi-source speech fusion after empty set elimination of the M speech annotation packets; outputting a mixed speech annotation layer teaching stream; performing spatial superposition synthesis after time sequence alignment of the mixed graphic annotation layer teaching stream and the mixed speech annotation layer teaching stream; and distributing the mixed output teaching stream to the multiple student terminals according to the student dimension transmission quality parameter in the two-dimensional transmission quality parameter.

[0040] From the M sets of structured annotation data packets, M speech annotation packets are extracted, each containing an expert's speech commentary, specifically including speech content and speech timestamp information. At the same time, M sets of spatial coordinate binding information for M graphic annotation layers are extracted from the M sets of structured annotation data packets. The graphic annotation layers include annotated graphic elements such as arrows, boxed regions, lines, etc., and the spatial coordinate binding information is the position of the annotated graphic elements in the video frame, for example, the start and end coordinates of an arrow, the coordinates of the four corners of a boxed region, etc.

[0041] Spatial-temporal mapping alignment refers to the unified alignment of all experts' graphic annotation layers to ensure that annotations at the same position in different expert terminals can be accurately corresponded. When multiple experts make annotations, due to different observation angles of each expert, the annotations made will also be different. Therefore, in this process, all graphic annotation layers are overlapped, i.e., each expert's annotations are aligned in time and space to ensure annotations at the same time point and the same spatial position.

[0042] After overlapping the graphic annotation layers, annotation conflicts may occur, for example, multiple experts make different annotations at the same position, such as different arrows, annotation content, etc. In order to ensure that the final output teaching stream is clear and consistent, conflict resolution is needed. For example, if multiple experts' annotations conflict, the annotations of certain experts can be selected as the final annotations based on factors such as experts' experience and authority. After conflict resolution, the final output is a mixed graphic annotation layer teaching stream that integrates all experts' annotation information and eliminates conflicts, ensuring that the annotation content is clear and consistent.

[0043] In the process of providing voice annotation by multiple experts, some experts may not annotate some video segments, resulting in some empty voice annotation packages. These empty voice annotation packages have no actual content, so empty set elimination is needed, that is, to delete the part that does not contain any valid voice annotation from the M voice annotation packages. Multi-source voice fusion is to fuse all valid voice annotation packages into a unified voice annotation stream, including: fusing voice annotation packages in chronological order to avoid confusion of voice annotations in different time periods, and merging the voice annotation contents if multiple experts have annotated the same video to ensure consistent voice content. After empty set elimination and voice fusion, the final output is a voice annotation layer mixed teaching stream, which contains the voice annotation information of all experts, ensuring that students can hear the explanations of all experts and providing a richer learning experience.

[0044] The graphical annotation layer mixed teaching stream and the voice annotation layer mixed teaching stream are time-aligned to ensure that the annotation contents (whether graphical annotation or voice annotation) in the two teaching streams are accurately synchronized in time, for example, the graphical annotation on a certain video frame needs to be synchronized with the voice annotation. Spatial superposition synthesis is to synthesize the time-aligned graphical annotation layer and voice annotation layer, and finally output a mixed output teaching stream containing visual and auditory annotations. In the spatial superposition synthesis process, graphical annotations are embedded into the corresponding pictures according to the spatial coordinates of the video, while voice annotations are played at the appropriate time points. Students can fully understand the surgical process by watching the graphical annotations and listening to the explanations of the voice annotations during the learning process.

[0045] In the dual-dimensional transmission quality parameter, the student dimension transmission quality parameter is set according to the transmission quality requirement of the student end, for example, the student end sets 1080p resolution, which is suitable for student learning and observation, and does not need the same high definition as the expert end, and the delay is less than or equal to 500ms, which ensures that the student end can receive the teaching video stream within a reasonable delay and meet the real-time learning needs. According to the student dimension transmission quality parameter, the mixed output teaching stream is distributed to multiple student ends to ensure that the student ends can receive the mixed output teaching stream that meets their needs.

[0046] Further, the method further comprises: The three-screen display engine for spatiotemporal synchronization is pre-deployed at the student end, wherein the three-screen display engine is used to control the split-screen display of the teaching guidance screen, the standard comparison screen, and the real-time observation screen; the historical cache backtracking decoding of the student-level code stream is performed according to the output timestamp of the mixed output teaching stream, and a synchronous comparison teaching stream is output; the mixed output teaching stream and the synchronous comparison teaching stream are mapped and loaded to the teaching guidance screen and the standard comparison screen; the real-time observation screen receives and codes the screen display of the student-level code stream returned by the collaborative middle station in real time, wherein the relative time offset amount is superimposed and displayed on the real-time observation screen according to the time deviation of the standard comparison screen and the teaching guidance screen.

[0047] The three-screen display engine for spatiotemporal synchronization is pre-deployed at the student end, wherein the three-screen display engine is used to control the split-screen display of the teaching guidance screen, the standard comparison screen, and the real-time observation screen; the historical cache backtracking decoding of the student-level code stream is performed according to the output timestamp of the mixed output teaching stream, and a synchronous comparison teaching stream is output; the mixed output teaching stream and the synchronous comparison teaching stream are mapped and loaded to the teaching guidance screen and the standard comparison screen; the real-time observation screen receives and codes the screen display of the student-level code stream returned by the collaborative middle station in real time, wherein the relative time offset amount is superimposed and displayed on the real-time observation screen according to the time deviation of the standard comparison screen and the teaching guidance screen.

[0048] The mixed output teaching stream contains all the information of the graphical and voice annotations. In order to ensure the time synchronization between all the annotation information and the video content, timestamps are used for control. At the student end, the student-level code stream and the mixed output teaching stream are time-aligned. Since the student-level code stream may be delayed or out of synchronization, the student-level code stream is decoded by historical cache backtracking according to the output timestamp of the mixed output teaching stream, and the data is compared with the timestamp in the mixed output teaching stream to ensure the synchronization of the two streams in time. The output synchronous comparison teaching stream contains all the time-aligned content, so that the student can see the corresponding teaching video and annotations at the correct time.

[0049] The mixed output teaching stream is displayed on the teaching guidance screen, which contains the annotations and teaching content of the expert; the synchronous comparison teaching stream is displayed on the standard comparison screen to show the reference standard or comparison information. The two streams are respectively mapped and loaded to the corresponding screens according to the display timestamp and spatial coordinate information, so that the student can see the graphics, voice, and video synchronized with the teaching content.

[0050] The real-time observation screen is used to display the video stream received by the student terminal in real time. The student can observe the real-time content occurring during the operation on this screen and cooperate with the middle station to return the student-level stream to the student terminal. The student-level stream is the real-time video stream received by the student terminal. In this process, the code screen display is enhanced to ensure that the student can see clear real-time video and the display content is coordinated with other screens.

[0051] There is a time deviation between the standard comparison screen and the teaching guidance screen. In particular, if the displayed content has different network delays or video processing delays, in order to provide consistent display on the real-time observation screen, the display of the standard comparison screen and the teaching guidance screen on the real-time observation screen is relatively time offset according to the time deviation. This means that the display content of the real-time observation screen will be adjusted according to the time deviation of the standard comparison screen and the teaching guidance screen, to ensure that the video and annotation information on different screens are synchronized when viewed by the student terminal.

[0052] Further, the mixed output teaching stream and the student-level stream support local perspective multi-scale switching operations on the student terminal.

[0053] Local perspective switching means that the student terminal can switch between multiple perspectives or regions based on different learning needs to better observe the video content. For example, in a surgical video, the student can switch between the main perspective and the detailed perspective of the surgical site, or switch between the annotations of multiple experts to obtain more comprehensive information. Multi-scale switching means that the student can adjust the scale size of the view according to different display needs to view different details or a wider view. For example, the student can view the content at different viewing distances through gesture operations or interface buttons. Through a large-scale view, the student can see the panorama of the surgical process and understand the general surgical steps. Through a small-scale view, the student can focus on a specific surgical site to view the details of expert annotations or observe specific operations.

[0054] Further, the surgical video frame sequence is visually desensitized based on organ feature recognition to obtain a desensitized video frame sequence, and the method comprises: A lightweight segmentation neural network is trained based on a medical privacy region annotation dataset. The lightweight segmentation neural network is used to perform semantic segmentation on the surgical video frame sequence to identify marked human privacy organ regions and text sensitive regions, and generate a privacy region mask sequence. The privacy region mask sequence is processed frame by frame to perform Gaussian blur processing on the human privacy organ region and pixel covering processing on the text sensitive region, and a desensitization processing frame sequence is output. The desensitization processing frame sequence and the surgical video frame sequence are spatially synthesized to output the desensitized video frame sequence.

[0055] The medical privacy region annotation dataset is a dataset containing medical image data such as surgical video frames, with annotations of privacy regions, including sensitive areas such as patients' private parts, faces, organ features, etc. These privacy regions require desensitization processing in surgical videos. The lightweight segmentation neural network is an optimized neural network specifically designed for image segmentation tasks. Compared with traditional full-function neural networks, lightweight networks usually have less computational demand and lower parameter quantity, so they are suitable for real-time video processing. Training the lightweight segmentation neural network using the medical privacy region annotation dataset enables it to automatically identify privacy regions in surgical videos. The network will learn how to segment human privacy organ regions such as genitals, faces, and text-sensitive regions such as medical text or identifiers in the video, such as patient names and medical record numbers.

[0056] Semantic segmentation is an image processing technique that assigns each pixel in an image to a specific class. In this step, semantic segmentation is applied to the sequence of surgical video frames to identify and label human privacy organ regions and text-sensitive regions using the trained lightweight segmentation neural network. After semantic segmentation, the lightweight segmentation neural network generates a privacy region mask for each surgical video frame. This privacy region mask is a binary image where sensitive regions are marked as 1, indicating areas that need to be desensitized, and the rest are marked as 0, indicating areas that do not need to be desensitized.

[0057] For the identified human privacy organ regions, Gaussian blur processing is applied. Gaussian blur is an image processing technique that blurs sensitive areas by averaging the pixels in the image, protecting the privacy of patients. This blurring process makes it impossible for viewers to clearly see the details of the human privacy organ regions, especially the face and private parts, thus achieving the purpose of privacy protection. For text-sensitive regions, pixel covering processing is used, which involves covering the text-sensitive regions with other unrelated image content such as fill colors or patterns, thus completely hiding these sensitive information. This processing ensures that text information is not leaked while not affecting other parts of the video.

[0058] After completing the Gaussian blur processing and pixel covering processing, the desensitization processing frame sequence is obtained. The desensitization processing frame sequence is a modified version of the original video, where all sensitive regions have been appropriately desensitized to ensure patient privacy is protected.

[0059] The desensitization processing frame sequence is spatially combined with the original surgery video frame sequence, and spatial combination refers to replacing the desensitized sensitive area with the processed version, such as blurring or covering, in each video frame, and the other areas remain unchanged. The final result of spatial combination is a new desensitized video frame sequence, which contains all the contents of the surgery video, but all the sensitive areas have been desensitized. This video can be safely used for teaching, research or sharing without worrying about leaking patient privacy.

[0060] Further, the PII anonymization processing is performed on the vital sign slice sequence to obtain a desensitized vital sign slice sequence, and the method comprises: Based on the identifier type, the vital sign slice sequence field is separated into a direct identifier field set and a quasi-identifier field set. The metadata stripping processing is performed on the direct identifier field set to obtain an anonymized direct field set. The numerical interval generalization processing is performed on the quasi-identifier field set to obtain a generalized quasi-identifier set. The anonymized direct field set and the generalized quasi-identifier set are mapped and overlaid to the vital sign slice sequence to obtain the desensitized vital sign slice sequence.

[0061] Vital sign data usually contains multiple types of fields, some of which are direct identifiers that can directly identify individuals, such as patient names, IDs, addresses, etc. Direct identifiers usually need to be completely removed or strongly protected; others are quasi-identifiers, which cannot directly identify individuals alone but may have identification potential when combined with other data, such as age, gender, date of birth, etc. Quasi-identifiers can be generalized or processed in other ways. Based on the identifier type, the fields in the vital sign slice sequence are classified to obtain a direct identifier field set and a quasi-identifier field set.

[0062] Metadata stripping is to remove information that can reveal the identity of the patient, for example, for direct identifier fields such as patient ID, name, etc. If these information are not removed, the privacy of the patient will be leaked. The metadata stripping processing is performed on the direct identifier field set to remove or replace all personal identity information in these fields with an anonymous identifier. After metadata stripping, an anonymized direct field set is obtained, in which all fields that can directly identify the patient have been removed or replaced, ensuring the privacy and security of the data.

[0063] For quasi-identifier field sets such as age, birth date, etc., simply removing these fields cannot effectively protect privacy, because they may still reveal the identity of an individual when combined with other data, therefore, numerical interval generalization is used to process quasi-identifier field sets, specifically, the numerical values in the quasi-identifier fields are converted to more extensive ranges or categories, for example, a specific age (such as 35 years old) is generalized to an age range (such as 30-40 years old), by processing the quasi-identifier field sets with numerical interval generalization, a generalized quasi-identifier set is obtained, which contains each quasi-identifier field replaced by a more extensive category or interval, thereby reducing the identifiability of the information. Through the generalization process, the identification ability of the quasi-identifier field is reduced, so that even if these data are combined with other information, the individual cannot be accurately identified, which can protect the privacy of patients while still retaining the statistical value of the data.

[0064] The anonymized direct field set and the generalized quasi-identifier set are mapped and overlaid onto the original vital sign slice sequence, that is, the anonymized direct field set replaces the direct identifier fields in the original data, and the generalized quasi-identifier set replaces the quasi-identifier fields in the original data, after the above processing, the sensitive fields in the original vital sign slice sequence have been replaced or generalized, forming a desensitized vital sign slice sequence, this new sequence contains data that has been processed for privacy protection, and can be used for analysis and sharing without revealing the identity of the patient, achieving the protection of patient privacy while retaining the value of the data for subsequent processing, research and analysis.

[0065] In summary, the multi-terminal collaborative interaction method for a remote medical teaching platform provided by the embodiments of the present application has the following technical effects: The collaborative platform receives the raw surgical video stream and vital sign data transmitted from the operating room. Through time-synchronized signal desensitization and fusion, it effectively protects patient privacy while ensuring data time synchronization. This ensures seamless integration of real-time video streams and vital sign data during surgery and eliminates potential privacy leak risks through desensitization processing, enabling secure sharing of teaching and analysis data. By layering and encoding the desensitized fused data packet stream into expert-level and student-level streams and employing dual-dimensional transmission quality parameter mapping, optimization can be performed according to the data transmission quality requirements of different terminals. The expert terminal receives a high-definition expert-level video stream, while the student terminal receives a video stream of appropriate quality based on bandwidth and equipment requirements, ensuring both transmission stability and quality. This layered encoding and quality parameter mapping effectively balances the transmission quality between different ports, ensuring that experts and students can smoothly receive suitable teaching video content under different network conditions, thereby improving the user experience and adaptability of the remote medical teaching platform. With the collaboration of experts and the client, based on preset annotation permissions, experts can perform multimodal collaborative annotation on surgical video streams and generate structured annotation data packages. This process ensures comprehensive analysis of the surgical procedure by experts, achieving efficient collaborative work on the expert side. The annotation content is output in a structured manner, which facilitates subsequent rendering and production of teaching streams, enhancing the diversity and comprehensibility of the teaching content. By performing spatially bound annotation dynamic rendering on M sets of structured annotation data packages, a hybrid output teaching stream with high accuracy standards and voice explanations is generated. Students can simultaneously obtain comprehensive teaching content from surgical videos and expert annotations, making their learning experience richer and more intuitive. In addition, the student side supports split-screen perspective switching, allowing them to watch real-time surgical videos transmitted from the expert side as well as receive fused and processed real-time teaching videos. This design enhances the student learning experience, enabling students to not only observe the surgical procedure in real time but also view expert annotations and explanations in the teaching videos, thereby improving the personalization and interactivity of learning. Overall, this intelligent medical system, through efficient data processing, flexible annotation, and multi-terminal collaboration, makes medical teaching more intelligent and interactive, while ensuring data security and patient privacy protection, greatly improving the quality and efficiency of remote medical teaching.

[0066] Example 2, based on the same inventive concept as the multi-terminal collaborative interaction method for remote medical teaching platforms in the foregoing examples, such as... Figure 2 As shown in the embodiment of this application, a multi-terminal collaborative interaction system for a remote medical teaching platform is provided, the system comprising: The signal desensitization fusion module 10 is used for, after receiving the original surgery video stream and vital sign data returned by the operating room end, obtaining a desensitization fusion data packet stream by performing time synchronization signal desensitization fusion in cooperation with the intermediate station; the mapping and distribution module 20 is used for, after the desensitization fusion data packet stream is hierarchically encoded and compressed into an expert-level code stream and a student-level code stream, the expert-level code stream and the student-level code stream are mapped and distributed to M expert ends and multiple student ends using double-dimensional transmission quality parameters; the cooperative labeling module 30 is used for, after the M expert ends synchronously receive the expert-level code stream, the multi-modal cooperative labeling of the expert-level code stream is performed based on a preset labeling permission distribution, and M groups of structured labeling data packets are output; the teaching stream distribution module 40 is used for, after the M groups of structured labeling data packets are subjected to spatially bound labeling dynamic rendering, a mixed output teaching stream is obtained, and the mixed output teaching stream is distributed to the multiple student ends, wherein the mixed output teaching stream and the student-level code stream support split-screen view switching at the student end.

[0067] Further, the signal desensitization fusion module 10 is used for performing the following operation steps: After the original surgery video stream is frame-level parsed to obtain a surgery video frame sequence, the surgery video frame sequence is visually desensitized based on organ feature recognition to obtain a desensitized video frame sequence; after the vital sign data is sliced based on a sampling point to obtain a vital sign slice sequence, the vital sign slice sequence is subjected to PII anonymization processing to obtain a desensitized vital sign slice sequence; the desensitized vital sign slice sequence is frame-level aligned to the desensitized video frame sequence through timestamp matching, time-bound data correlation packaging is performed, and the desensitized fusion data packet stream is output.

[0068] Further, the cooperative labeling module 30 is used for performing the following operation steps: The M expert ends are pre-assigned with role permissions to obtain M role labeling permissions; a first expert end performs an annotation tool interaction process on the expert-level code stream after decoding and rendering based on a first role labeling permission, performs annotation spatial coordinate binding, and outputs a first group of structured labeling data packets; if the first role labeling permission includes voice annotation permission, a first voice annotation packet is generated through voiceprint separation and text transcription in a real-time voice collection process; the first voice annotation packet is stored in the first group of structured labeling data packets based on timestamp association.

[0069] Further, the teaching stream distribution module 40 is used for performing the following operation steps: extracting M sets of spatial coordinate binding information of M speech annotation packets and M graphic annotation layers in the M sets of structured annotation data packets; performing multi-expert annotation conflict resolution after spatiotemporal mapping alignment of the M sets of spatial coordinate binding information by overlapping the M graphic annotation layers, and outputting a graphic annotation layer mixed teaching stream; performing multi-source speech fusion after empty set elimination of the M speech annotation packets, and outputting a speech annotation layer mixed teaching stream; performing spatial superposition synthesis after time sequence alignment of the graphic annotation layer mixed teaching stream and the speech annotation layer mixed teaching stream to obtain the mixed output teaching stream; and distributing the mixed output teaching stream to the multiple student terminals according to the student dimension transmission quality parameter in the two-dimension transmission quality parameter.

[0070] Further, the teaching stream distribution module 40 is configured to perform the following operation steps: deploying a spatiotemporal synchronization three-screen display engine at the student terminal, wherein the three-screen display engine is configured to control split-screen display of a teaching guidance screen, a standard comparison screen and a real-time observation screen; performing historical cache back decoding of the student-level code stream according to an output timestamp of the mixed output teaching stream, and outputting a synchronization comparison teaching stream; mapping and loading the mixed output teaching stream and the synchronization comparison teaching stream to the teaching guidance screen and the standard comparison screen; and the real-time observation screen receives and codes and displays the student-level code stream returned by the collaborative hub in real time, wherein the real-time observation screen performs superimposed display of a relative time offset according to a time deviation between the standard comparison screen and the teaching guidance screen.

[0071] Further, the mixed output teaching stream and the student-level code stream support local perspective multi-scale switching operation at the student terminal.

[0072] Further, the signal desensitization fusion module 10 is configured to perform the following operation steps: training a lightweight segmentation neural network based on a medical privacy region annotation data set; performing semantic segmentation on the surgical video frame sequence using the lightweight segmentation neural network to identify marked human privacy organ regions and text sensitive regions, and generating a privacy region mask sequence; performing Gaussian blur processing on the human privacy organ regions and pixel covering processing on the text sensitive regions on the privacy region mask sequence frame by frame to output a desensitization processing frame sequence; and spatially synthesizing the desensitization processing frame sequence and the surgical video frame sequence to output the desensitization video frame sequence.

[0073] Further, the signal desensitization fusion module 10 is configured to perform the following operation steps: Based on the identifier type, the vital sign slice sequence field is separated into a direct identifier field set and a quasi-identifier field set; metadata stripping processing is performed on the direct identifier field set to obtain an anonymized direct field set; numerical interval generalization processing is performed on the quasi-identifier field set to obtain a generalized quasi-identifier set; and the anonymized direct field set and the generalized quasi-identifier set are mapped and overlaid to the vital sign slice sequence to obtain the desensitized vital sign slice sequence.

[0074] The person skilled in the art can clearly understand the multi-terminal collaborative interaction system for the remote medical teaching platform in the embodiments from the foregoing detailed description of the multi-terminal collaborative interaction method for the remote medical teaching platform. Since the multi-terminal collaborative interaction system corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method part description.

[0075] Embodiment three provides a storage medium having a computer program stored thereon, and the computer program is executed by a processor to implement any step of the method in embodiment one.

[0076] The technical features of the above embodiments can be combined in any manner. To make the description concise, not all possible combinations of the technical features in the above embodiments are described, however, as long as the combinations of the technical features do not exist contradictions, they should be considered as the scope of the present disclosure.

[0077] The above description of the disclosed embodiments enables a person skilled in the art to implement or use the present application. Various modifications to the embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A multi-terminal collaborative interaction method for a telemedicine teaching platform, characterized in that, The method comprises: The cooperative middle station receives the original operation video stream and vital sign data returned by the operating room end, performs signal desensitization fusion through time synchronization, and obtains a desensitization fusion data packet stream; After the desensitization fusion data packet stream is hierarchically encoded and compressed into an expert-level code stream and a student-level code stream, the expert-level code stream and the student-level code stream are distributed to M expert ends and multiple student ends by using double-dimensional transmission quality parameter mapping; After the M expert ends synchronously receive the expert-level code stream, the M expert ends perform multi-modal cooperative labeling on the expert-level code stream based on preset labeling permission distribution, and output M groups of structured labeling data packets; After spatially bound labeling dynamic rendering is performed on the M groups of structured labeling data packets, a mixed output teaching stream is obtained, and the mixed output teaching stream is distributed to the multiple student ends, wherein the mixed output teaching stream and the student-level code stream support split-screen view switching at the student end.

2. The multi-end collaborative interaction method for a telemedicine teaching platform according to claim 1, wherein, The cooperative middle station receives the original operation video stream and vital sign data returned by the operating room end, performs signal desensitization fusion through time synchronization, and obtains a desensitization fusion data packet stream, and the method comprises: The original operation video stream is frame-level parsed to obtain a sequence of operation video frames, and the sequence of operation video frames is visually desensitized based on organ feature recognition to obtain a sequence of desensitized video frames; The vital sign data is sliced based on sampling points to obtain a sequence of vital sign slices, and the sequence of vital sign slices is subjected to PII anonymization processing to obtain a sequence of desensitized vital sign slices; The sequence of desensitized vital sign slices is frame-level aligned to the sequence of desensitized video frames through timestamp matching, time-bound data correlation packaging is performed, and the desensitization fusion data packet stream is output.

3. The multi-end collaborative interaction method for a telemedicine teaching platform according to claim 1, wherein, After the M expert ends synchronously receive the expert-level code stream, the M expert ends perform multi-modal cooperative labeling on the expert-level code stream based on preset labeling permission distribution, and output M groups of structured labeling data packets, the method comprises: Role permission pre-distribution is performed on the M expert ends to obtain M role labeling permissions; A first expert end performs an annotation tool interaction process on the expert-level code stream after decoding and rendering based on a first role labeling permission, performs annotation spatial coordinate binding, and outputs a first group of structured labeling data packets; If the first role labeling permission includes voice annotation permission, a first voice annotation packet is generated through voiceprint separation and text transcription in a real-time voice collection process; The first voice annotation packet is stored in the first group of structured labeling data packets based on timestamp association.

4. The multi-end collaborative interaction method for a telemedicine teaching platform according to claim 1, wherein, After spatially bound labeling dynamic rendering is performed on the M groups of structured labeling data packets, a mixed output teaching stream is obtained, and the mixed output teaching stream is distributed to the multiple student ends, and the method comprises: M spatial coordinate binding information of M voice annotation packets and M graphic annotation layers in the M groups of structured labeling data packets is extracted; After time-space mapping alignment of the M spatial coordinate binding information is performed by overlapping the M graphic annotation layers, multi-expert annotation conflict resolution is performed, and a graphic annotation layer mixed teaching stream is output; After empty set elimination is performed on the M voice annotation packets, multi-source voice fusion is performed, and a voice annotation layer mixed teaching stream is output; After the graphic annotation layer mixed teaching stream and the voice annotation layer mixed teaching stream are time-aligned, spatial superposition synthesis is performed to obtain the mixed output teaching stream; According to a student dimension transmission quality parameter in the dual-dimension transmission quality parameter, the mixed output teaching stream is distributed to the multiple student terminals.

5. The multi-end collaborative interaction method for a telemedicine teaching platform according to claim 1, wherein, The method further comprises: A spatiotemporal synchronization three-screen display engine is pre-deployed at the student terminal, wherein the three-screen display engine is used to control split-screen display of a teaching guidance screen, a standard comparison screen and a real-time observation screen; According to an output timestamp of the mixed output teaching stream, historical cache backtracking decoding of the student-level code stream is performed to output a synchronous comparison teaching stream; The mixed output teaching stream and the synchronous comparison teaching stream are mapped and loaded to the teaching guidance screen and the standard comparison screen; The real-time observation screen receives and adds coding screen display to the student-level code stream returned by the collaborative middle station in real time, wherein according to a time deviation between the standard comparison screen and the teaching guidance screen, relative time deviation superposition display is performed on the real-time observation screen.

6. The multi-end collaborative interaction method for a telemedicine teaching platform according to claim 1, wherein, The mixed output teaching stream and the student-level code stream support local perspective multi-scale switching operation at the student terminal.

7. The multi-end collaborative interaction method for a telemedicine teaching platform according to claim 2, wherein, Based on organ feature recognition, visual desensitization is performed on the surgery video frame sequence to obtain a desensitized video frame sequence, and the method comprises: A lightweight segmentation neural network is trained based on a medical privacy region annotation dataset; Semantic segmentation is performed on the surgery video frame sequence by using the lightweight segmentation neural network to identify marked human privacy organ regions and text sensitive regions, and a privacy region mask sequence is generated; By performing Gaussian blur processing on the human privacy organ regions and pixel coverage processing on the text sensitive regions on the privacy region mask sequence frame by frame, a desensitization processing frame sequence is output; The desensitization processing frame sequence and the surgery video frame sequence are spatially synthesized to output the desensitized video frame sequence.

8. The multi-end collaborative interaction method for a telemedicine teaching platform according to claim 2, wherein, PII anonymization processing is performed on the vital sign slice sequence to obtain a desensitized vital sign slice sequence, and the method comprises: Based on identifier types, fields of the vital sign slice sequence are separated into a direct identifier field set and a quasi-identifier field set; Metadata stripping processing is performed on the direct identifier field set to obtain an anonymized direct field set; Numerical interval generalization processing is performed on the quasi-identifier field set to obtain a generalized quasi-identifier set; The anonymized direct field set and the generalized quasi-identifier set are mapped and overlaid to the vital sign slice sequence to obtain the desensitized vital sign slice sequence.

9. A multi-terminal collaborative interaction system for a telemedicine teaching platform, characterized in that, A multi-terminal collaborative interaction method for implementing the remote medical teaching platform of any one of claims 1-8, the system comprising: A signal desensitization fusion module, configured to, after receiving an original surgery video stream and vital sign data returned by an operating room terminal, perform time-synchronized signal desensitization fusion by the collaborative middle station to obtain a desensitization fusion data packet stream; A mapping distribution module, configured to, after layer-encoding and compressing the desensitization fusion data packet stream into an expert-level code stream and a student-level code stream, map and distribute the expert-level code stream and the student-level code stream to M expert terminals and multiple student terminals by using dual-dimension transmission quality parameters. A collaborative labeling module is configured to perform multi-modal collaborative labeling on the expert-level code stream based on a preset labeling permission distribution after the M expert terminals synchronously receive the expert-level code stream, and output M sets of structured labeling data packets. A teaching stream distribution module is configured to perform spatially bound dynamic rendering on the M sets of structured labeling data packets, obtain a mixed output teaching stream, and then distribute the mixed output teaching stream to the multiple student terminals, wherein the mixed output teaching stream and the student-level code stream support split-screen view switching at the student terminals.

10. A storage medium having stored thereon a computer program, characterized in that The computer program, when executed by a processor, implements the steps of the multi-terminal collaborative interaction method for the remote medical teaching platform according to any one of claims 1 to 8.