Medical video sharing method, device, equipment, storage medium and product

By using intelligent analysis models to automatically identify and label medical videos, the problem of low efficiency in manual analysis in existing technologies is solved, and efficient and accurate medical video analysis and information provision are achieved.

CN122179624APending Publication Date: 2026-06-09ZHUOYI (SHENZHEN) TECHNOLOGY CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHUOYI (SHENZHEN) TECHNOLOGY CO LTD
Filing Date
2026-04-21
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

Current medical video recording and transmission technologies cannot automatically identify surgical instruments, diseased organs, and operational stages, requiring manual viewing and analysis, which is inefficient and prone to subjective bias.

Method used

By receiving video upload requests initiated by the user interaction layer, the system performs intelligent analysis on the initial medical video, uses a deep learning-based intelligent analysis model to perform frame-by-frame image segmentation and semantic segmentation, identifies and labels medical entities, surgical stages and risk information, and saves the analysis results locally for users to browse.

Benefits of technology

It achieves automated medical video analysis, improving efficiency, avoiding subjective bias, and providing detailed information on medical entities, surgical stages, and risk-related information, making it easier for users to quickly locate and understand the information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122179624A_ABST
    Figure CN122179624A_ABST
Patent Text Reader

Abstract

This application discloses a medical video sharing method, apparatus, device, storage medium, and product, relating to the field of medical technology. The medical video sharing method includes: receiving a video upload request initiated by a user interaction layer, wherein the video upload request includes an initial medical video; performing intelligent analysis on the initial medical video to obtain a target medical video labeled with analysis results, wherein the analysis results include medical entity information, surgical stage, and risk-related information; and saving the target medical video locally for users in the user interaction layer connected to the cloud server to browse the target medical video. This application eliminates the need for manual viewing and analysis, avoiding inefficiency and subjective bias.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of medical technology, and in particular to a method, apparatus, device, storage medium and product for sharing medical videos. Background Technology

[0002] In related technologies, medical videos are usually recorded, stored, and transmitted through recording and broadcasting systems. However, the above solutions cannot automatically identify surgical instruments, diseased organs, and operation stages, requiring manual viewing and analysis, which is inefficient and prone to subjective bias. Summary of the Invention

[0003] The main objective of this application is to provide a method, apparatus, device, storage medium, and product for sharing medical videos, aiming to solve the technical problems of low efficiency and susceptibility to subjective bias caused by the need for manual viewing and analysis.

[0004] To achieve the above objectives, this application proposes a medical video sharing method, which includes: Receive video upload requests initiated by the user interaction layer, where the video upload request includes the initial medical video; The initial medical video is subjected to intelligent analysis to obtain a target medical video labeled with the analysis results, wherein the analysis results include medical entity information, surgical stage and risk-related information; The target medical video is saved locally so that users in the user interaction layer connected to the cloud server can browse the target medical video.

[0005] In one embodiment, the step of performing intelligent analysis on the initial medical video to obtain a target medical video labeled with the analysis results includes: The initial medical video is segmented frame by frame to obtain a sequence of segmented medical images; The segmented medical image sequence is analyzed using a preset intelligent analysis model to obtain the analysis results to be labeled. The analysis results to be labeled are then labeled on the corresponding video timeline in the initial medical video to obtain the target medical video labeled with the analysis results.

[0006] In one embodiment, the step of analyzing the segmented medical image sequence using a preset intelligent analysis model to obtain the analysis results to be labeled includes: The medical image sequence is preprocessed to obtain preprocessed image frames; The preprocessed image frame is semantically segmented using a preset intelligent analysis model to identify and segment medical entities in the image frame, and an analysis result labeled with the medical entities is obtained. The preset intelligent analysis model is a deep learning-based image segmentation model, which includes at least medical entities such as surgical instruments, organs and tissues, and lesion areas. The analysis results are post-processed and optimized to obtain the optimized analysis results to be labeled.

[0007] In one embodiment, the step of performing frame-by-frame analysis on the initial medical video after preliminary analysis using a preset intelligent analysis model to obtain the analysis results to be labeled includes, prior to the step of extracting key regions from the target image and performing image recognition on the key regions to obtain medical entity information: Acquire historical medical videos and perform preliminary analysis on the historical medical videos to obtain the anonymized historical medical videos and preliminary analysis results; Key regions in desensitized historical medical videos are identified, and image recognition is performed on the key regions to obtain medical entity information. The key regions include surgical instrument regions and organ lesion regions. The desensitized historical medical videos were divided into stages and risks were identified, resulting in surgical stages and risky actions. For the risky actions, generate operation improvement suggestions, and based on the operation improvement suggestions and the risky actions, determine risk-related information; Based on the preliminary analysis results, the medical entity information, the surgical stage, and the risk-related information, the historical medical videos are annotated to obtain the annotated historical medical videos. The labeled historical medical videos are input into the initial intelligent analysis model for iterative training to obtain a preset intelligent analysis model that meets the accuracy requirements.

[0008] In one embodiment, the step of using a preset intelligent analysis model to perform semantic segmentation on the preprocessed image frame, identifying and segmenting medical entities in the image frame, and obtaining analysis results labeled with the medical entities includes: The pre-defined intelligent analysis model is used to output a set of candidate targets, including surgical stage and instrument category, for the preprocessed image frames, and the prior confidence of the candidate target set is given. The candidate target set is transformed into region cues in the image space, and the model is driven by the region cues to perform pixel-level segmentation and instantiation separation. Morphological processing is used to improve the quality of the segmentation boundary, resulting in an image sequence with pixel-level masks. Cross-frame tracking is performed on the pixel-level mask to obtain a mask sequence with trajectory sequence number and an image sequence with occurrence duration; Spatiotemporal alignment and conflict fusion correction are performed on the mask sequence with trajectory sequence number and the image sequence with occurrence duration to obtain the corrected image sequence; Based on the corrected image sequence, analysis results labeled with the medical entities are obtained, wherein the analysis results include the start / end time, absolute time and duration of each device and stage, and a visualization resource index; The step of obtaining the analysis results labeled with the medical entities based on the corrected image sequence includes: The preset intelligent analysis model records model information including model version, parameters, and anomaly rollback information, wherein the model information is used to ensure the traceability and quality control of the results.

[0009] In one embodiment, the video upload request includes user permissions, and the step of saving the target medical video locally for viewing by a user in the user interaction layer connected to the cloud server includes: When a user is detected to have a need for video learning, a user profile is obtained, which includes the user's major, permission level, and skill map. The permission level is determined based on role permissions, attribute permissions, and autonomous access permissions. Based on the user's expertise and the skill graph, determine the medical videos to be recommended within the scope of the permission level; If the user finishes watching the recommended medical video, the user's learning evaluation result is determined based on the user's behavioral indicators of watching the recommended medical video. The learning assessment results are fed back to the user's corresponding client, and the skill map is updated based on the learning assessment results.

[0010] Furthermore, to achieve the above objectives, this application also proposes a medical video sharing device, which includes: The receiving module is used to receive video upload requests initiated by the user interaction layer, wherein the video upload request includes the initial medical video; The analysis module is used to perform intelligent analysis on the initial medical video to obtain a target medical video labeled with the analysis results, wherein the analysis results include medical entity information, surgical stage and risk-related information; A storage module is used to save the target medical video locally so that users in the user interaction layer connected to the cloud server can browse the target medical video.

[0011] In addition, to achieve the above objectives, this application also proposes a medical video sharing device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the medical video sharing method as described above.

[0012] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the medical video sharing method described above.

[0013] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the medical video sharing method described above.

[0014] One or more technical solutions proposed in this application have at least the following technical effects: In contrast to related technologies that typically record, store, and transmit medical videos through recording and broadcasting systems, these methods cannot automatically identify surgical instruments, organs, and operational stages, requiring manual viewing and analysis, resulting in low efficiency and susceptibility to subjective bias. This application, however, receives a video upload request initiated by a user interaction layer. The video upload request includes an initial medical video and user permissions. The initial medical video undergoes intelligent analysis to obtain a target medical video labeled with analysis results, including medical entity information, surgical stage, and risk-related information. The target medical video is then saved locally for users in the user interaction layer connected to the cloud server to view. After obtaining the initial medical video from the video upload request, this application automatically performs intelligent analysis on the initial medical video to obtain a target medical video containing analyzed medical entity information, surgical stage, and risk-related information. The analyzed target medical video is then saved locally, allowing users in the user interaction layer connected to the cloud server to view the target medical video without requiring manual viewing and analysis, thus avoiding inefficiency and subjective bias. Attached Figure Description

[0015] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a flowchart illustrating an embodiment of the medical video sharing method of this application. Figure 2 This is a collaborative architecture diagram of the cloud service layer and user interaction layer of the medical video sharing method of this application; Figure 3 This is a schematic diagram of the permission management matrix for the medical video sharing method in this application; Figure 4 This is a flowchart illustrating Embodiment 2 of the medical video sharing method of this application; Figure 5 This is a schematic diagram of the module structure of the medical video sharing device according to an embodiment of this application; Figure 6 This is a schematic diagram of the device structure of the hardware operating environment involved in the medical video sharing method in this application embodiment.

[0018] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0019] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0020] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0021] The main solution of this application embodiment is: receiving a video upload request initiated by the user interaction layer, wherein the video upload request includes an initial medical video; performing intelligent analysis on the initial medical video to obtain a target medical video labeled with analysis results, wherein the analysis results include medical entity information, surgical stage, and risk-related information; and saving the target medical video locally for users in the user interaction layer connected to the cloud server to browse the target medical video.

[0022] In related technologies, medical videos are usually recorded, stored, and transmitted through recording and broadcasting systems. However, the above solutions cannot automatically identify surgical instruments, diseased organs, and operation stages, requiring manual viewing and analysis, which is inefficient and prone to subjective bias.

[0023] After obtaining the initial medical video from the video upload request, this application automatically performs intelligent analysis on the initial medical video to obtain the target medical video, which includes medical entity information, surgical stage, and risk-related information. The analyzed target medical video is then saved locally, allowing users in the user interaction layer connected to the cloud server to browse the target medical video without the need for manual viewing and analysis, thus avoiding inefficiency and subjective bias.

[0024] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device or medical video sharing device capable of performing the above functions. The following description uses a medical video sharing device as an example to illustrate this embodiment and the subsequent embodiments.

[0025] Based on this, embodiments of this application provide a method for sharing medical videos, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the medical video sharing method of this application.

[0026] In this embodiment, the medical video sharing method includes steps S10 to S30: Step S10: Receive a video upload request initiated by the user interaction layer, wherein the video upload request includes the initial medical video; It should be noted that the execution subject in this embodiment is a medical video sharing device. This medical video sharing device has a cloud service layer. Users (such as doctors or teaching administrators) select a recorded surgical video through the system interface (web or mini-program), set access permissions and processing priorities (normal or emergency), and construct a video upload request based on the surgical video, access permissions, and processing priorities. The constructed video upload request is encrypted using the national cryptographic algorithm (SM4) or AES-256. The encryption key is pre-negotiated and securely stored between the user interaction layer and the cloud service layer. The encrypted video upload request is sent to a designated interface of the cloud service layer. The cloud service layer listens to this designated interface, receives the video upload request, decrypts the video upload request using a pre-shared key or private key, and parses the video upload request to obtain the parsed initial medical video and the pre-set parameters, namely access permissions and processing priorities. (Refer to...) Figure 2 , Figure 2 It provides a collaborative architecture diagram of the cloud service layer and the user interaction layer.

[0027] Furthermore, all medical videos involved in this application were obtained with the user's permission or consent; that is, when this application is applied to specific products or technologies, user permission is required to acquire and process the relevant data, and the processing of the relevant data must comply with the relevant laws, regulations and regulatory standards of the relevant countries and regions.

[0028] Furthermore, after receiving the initial medical video, the medical video sharing device will sequentially verify file integrity, check format compatibility, and scan for viruses or malicious content. After the above verifications, the successfully verified initial medical video will be intelligently segmented and uploaded, and stored locally. The segment size will be adaptive, and the upload will be performed in parallel through multi-threading. The medical video sharing device will record the list of successfully uploaded segments. If the network is interrupted or the computer is in sleep mode, the upload can continue from the point of interruption without re-uploading.

[0029] Specifically, the medical video sharing device dynamically adjusts the size of the uploaded video segments to match the current network bandwidth and controls the number of concurrent upload threads to avoid CPU overload and network congestion.

[0030] Furthermore, the medical video sharing device offers fully private deployment for hospitals concerned about data security: 1. Localized deployment architecture: local NAS storage, local database cluster, intranet isolation, and optional VPN access.

[0031] 2. Data encryption scheme: end-to-end encryption, key management.

[0032] 3. Backup and disaster recovery.

[0033] Step S20: Perform intelligent analysis on the initial medical video to obtain a target medical video labeled with the analysis results, wherein the analysis results include medical entity information, surgical stage and risk-related information; Understandably, medical entity information includes instrument information and organ information. Surgical stages refer to different phases within the surgical procedure, such as preoperative preparation, surgical operation, and postoperative management, each with its specific procedures and precautions. Risk-related information refers to potential risks and precautions associated with the surgery or medical procedure; this information helps medical personnel identify and prevent potential problems in advance. Medical video sharing devices analyze initial medical videos, extracting and labeling medical entity information, surgical stages, and risk-related information to create target medical videos.

[0034] Step S30: Save the target medical video locally so that users in the user interaction layer connected to the cloud server can browse the target medical video.

[0035] It should be noted that the medical video sharing device stores the medical videos, after intelligent analysis and annotation, on a local storage device so that users can access and browse these videos through the user interaction layer.

[0036] In one feasible implementation, step S20 may include the following steps: The initial medical video is segmented frame by frame to obtain a sequence of segmented medical images; It should be noted that the medical video sharing device divides continuous video frames into medical image sequences.

[0037] The segmented medical image sequence is analyzed using a preset intelligent analysis model to obtain the analysis results to be labeled. Understandably, the pre-set intelligent analysis model is built on the Transformer architecture, using CLIP-ViT-bigG as the visual encoder, giving it powerful visual understanding and language generation capabilities. Combined with image segmentation preprocessing, medical knowledge base retrieval, and multimodal fusion technologies, it achieves deep semantic understanding of surgical videos. Compared to traditional CNN methods, it can better understand the contextual relationships and medical semantics of surgical scenes. The medical video sharing device uses the pre-set intelligent analysis model to perform detailed analysis on each frame of the initial medical video, extracting information that needs to be annotated frame by frame. This information will be used to generate the annotated content for the target medical video.

[0038] The analysis results to be labeled are then labeled on the corresponding video timeline in the initial medical video to obtain the target medical video labeled with the analysis results.

[0039] It's important to note that the video timeline represents the video's playback time. Timeline annotations help users quickly locate key information while watching the video. The medical video sharing device annotates the analysis results extracted by the intelligent analysis model (such as medical entity information, surgical stages, risk-related information, etc.) on the timeline of the initial medical video. After annotation, a new video version, the target medical video, is generated. This video includes the original video content as well as the annotated analysis results, making it easier for users to understand and utilize the video.

[0040] Understandably, medical video sharing devices generate analysis results through preset analysis models (such as classifiers, regressors, generators, etc.), and output the generated analysis results, which will be used for subsequent video annotation.

[0041] Optionally, the medical video sharing device, based on the intelligent analysis model, further introduces a structured medical knowledge graph as an external constraint mechanism in the model reasoning process to improve the medical rationality of the recognition results and the accuracy of the surgical stage judgment.

[0042] Specifically, firstly, a medical knowledge subgraph is constructed for specific surgical types. This subgraph extracts entities and their logical relationships related to the target surgery from authoritative medical knowledge bases (such as SNOMED CT, ICD-10-PCS, and hospital internal surgical operation guidelines). For example, for "laparoscopic cholecystectomy," the system automatically extracts the anatomical structures involved in the procedure (such as the gallbladder, liver, and common bile duct), permitted surgical instruments (such as electrocautery hooks, grasping forceps, and suction devices), standard surgical steps (such as establishing pneumoperitoneum, separating Calot's triangle, and clamping the cystic duct), and the reasonable combinations of each step with instruments and anatomical structures. Simultaneously, irrelevant or contraindicated instrument combinations (such as craniotomy drills, bone hammers, and other neurosurgical tools) are explicitly excluded. This knowledge subgraph is stored in a graph database format, supporting efficient querying and relation traversal.

[0043] Subsequently, during video analysis, the device first performs preliminary processing on the input medical video frames through a visual perception module, including object detection, semantic segmentation, and temporal feature extraction, thereby obtaining the surgical instruments identified at each time point, the exposed anatomical structures, and the preliminary surgical stage labels. These preliminary results may contain errors that do not conform to medical common sense, such as misdetecting a craniotomy drill during laparoscopic cholecystectomy.

[0044] Next, the device uses the current video context (including the identified surgical type, the instrument sequence within the current time window, and visible anatomical structures) as query conditions to access the previously constructed medical knowledge subgraph in real time. Specifically, the device performs two key operations: first, it verifies whether the currently identified instrument belongs to the set of instruments permitted for the surgery; second, it checks whether the current combination conforms to the known surgical procedure logic. If an instrument is found not to be in the list of legal instruments for the procedure, it is determined to be an identification anomaly, and a correction mechanism is triggered.

[0045] The correction mechanism has two levels: during the inference phase, all instrument categories incompatible with the current surgical type are directly blocked, preventing them from being output; during the training phase, predictions that violate medical logic are recorded, and an additional penalty term is added to the loss function to guide the model to reduce similar errors in the future. This combination of hard constraints and soft optimization ensures that the output always stays within the boundaries of medical rationality.

[0046] Building upon this, the device further utilizes a knowledge graph to support reasoning about implicit surgical stages. Because some surgical stages are visually highly similar (e.g., "gallbladder traction" and "gallbladder dissection"), accurate differentiation based solely on image classification is difficult. Therefore, the device matches consecutively identified instrument usage sequences within a time window (e.g., "continuous traction with forceps + initiation of cutting with an electric hook") with predefined "stage-instrument" mappings in the knowledge graph. The knowledge graph explicitly records that the "gallbladder dissection" stage is typically accompanied by a pattern of "coordinated operation of the electric hook and forceps." When the actual instrument sequence closely matches this pattern, even with blurred visual features, the device can still correct the current stage to "gallbladder dissection" based on knowledge-driven logical inference.

[0047] To enhance the temporal coherence of reasoning, the system also introduces a graph-based temporal reasoning module. This module treats the identification result (instrument, anatomical state) at each time step as a dynamic node in the graph and constructs connections between nodes based on prior relationships in the knowledge subgraph. By running graph neural network reasoning on this dynamic graph, the device can capture logical dependencies across time steps, such as "the cystic duct clamping stage can only proceed after Calot's triangle is fully exposed." This joint reasoning, combining temporal information and structured knowledge, significantly improves the robustness and clinical consistency of stage identification.

[0048] In one feasible implementation, the step of analyzing the segmented medical image sequence using a preset intelligent analysis model to obtain the analysis results to be labeled includes: The medical image sequence is preprocessed to obtain preprocessed image frames; It should be noted that the medical video sharing device uses the original medical image sequence to improve the performance of subsequent models.

[0049] The preprocessed image frame is semantically segmented using a preset intelligent analysis model to identify and segment medical entities in the image frame, and an analysis result labeled with the medical entities is obtained. The preset intelligent analysis model is a deep learning-based image segmentation model, which includes at least medical entities such as surgical instruments, organs and tissues, and lesion areas. Understandably, the medical video sharing device uses a pre-trained deep learning segmentation model to perform pixel-level classification of each frame of the image, identifying three key medical entities: medical entities such as surgical instruments, organs and tissues, and lesion areas. Each frame of the image generates a corresponding segmentation mask, and each pixel is assigned a category label (e.g., 0=background, 1=instrument, 2=organ, 3=lesion).

[0050] The analysis results are post-processed and optimized to obtain the optimized analysis results to be labeled.

[0051] After the deep learning model completes semantic segmentation, the medical video sharing device performs fine-tuning on the initial segmentation output to improve its accuracy, robustness, and clinical usability.

[0052] In one feasible implementation, the steps of using a preset intelligent analysis model to perform semantic segmentation on the preprocessed image frame, identifying and segmenting medical entities in the image frame, and obtaining analysis results labeled with the medical entities include: The pre-defined intelligent analysis model is used to output a set of candidate targets, including surgical stage and instrument category, for the preprocessed image frames, and the prior confidence of the candidate target set is given. It should be noted that the medical video sharing device uses a preset intelligent analysis model to simultaneously infer the surgical stage and instrument category from the pre-processed single-frame image, and provides a priori confidence for each inference result.

[0053] The candidate target set is transformed into region cues in the image space, and the model is driven by the region cues to perform pixel-level segmentation and instantiation separation. Morphological processing is used to improve the quality of the segmentation boundary, resulting in an image sequence with pixel-level masks. Understandably, the medical video sharing device transforms each candidate target in the candidate target set into a region cue in the image space. Specifically, this includes bounding box information. The machine uses the center point of the bounding box as a point cue, or directly uses the entire bounding box as a box cue. Subsequently, the machine invokes a pre-defined, cue-driven pixel-level segmentation model. After receiving the original pre-processed image frame and the aforementioned region cue, this model performs instance-aware segmentation inference. For each region cue, the model independently generates a corresponding pixel-level mask, ensuring that the entities corresponding to different cuees are separated into independent instances in the output. After obtaining the initial pixel-level masks, the machine sequentially performs a series of morphological post-processing operations on each mask to improve boundary quality. After completing the above processing, the machine organizes the optimized masks into an image sequence according to the temporal order of the original video frames.

[0054] Cross-frame tracking is performed on the pixel-level mask to obtain a mask sequence with trajectory sequence number and an image sequence with occurrence duration; It should be noted that the medical video sharing device starts by initializing the trajectory manager, and then performs instance parsing and feature extraction (including spatial location, geometric attributes, and appearance embedding) on ​​each frame mask. It then uses multimodal similarity matching (fusing IoU / distance and appearance cosine similarity) combined with the Hungarian algorithm to achieve optimal association between the current instance and historical trajectories. Based on the matching results, the machine dynamically performs trajectory updates, new trajectory creation, or miss markers, supplemented by a lifecycle management mechanism—the process ends when a trajectory has not appeared for a long time, calculating its actual occurrence duration. After each frame is processed, the original local instance ID is replaced with the global trajectory sequence number, generating a mask frame with trajectory identification.

[0055] Spatiotemporal alignment and conflict fusion correction are performed on the mask sequence with trajectory sequence number and the image sequence with occurrence duration to obtain the corrected image sequence; Understandably, medical video sharing devices perform spatiotemporal correction and conflict fusion on mask sequences with trajectory sequence numbers and image sequences with occurrence durations to resolve trajectory misalignment or information conflicts caused by tracking errors, target occlusion, and inter-frame inconsistencies.

[0056] Based on the corrected image sequence, analysis results labeled with the medical entities are obtained, wherein the analysis results include the start / end time, absolute time and duration of each device and stage, and a visualization resource index; It should be noted that the medical video sharing device is based on the corrected image sequence. The machine automatically generates structured medical entity analysis results by parsing the mask of each corrected trajectory. The results clearly mark the start time, end time, corresponding absolute time (such as specific hours, minutes, and seconds), and duration (accurate to seconds or milliseconds) of each type of surgical instrument and surgical stage, and associate each event with a visual resource index (such as keyframe screenshots and video clip paths with mask overlay).

[0057] The step of obtaining the analysis results labeled with the medical entities based on the corrected image sequence includes: The preset intelligent analysis model records model information including model version, parameters, and anomaly rollback information, wherein the model information is used to ensure the traceability and quality control of the results.

[0058] Understandably, after the analysis process is completed, the medical video sharing device automatically records key information of the preset intelligent analysis model used, including the model version number, core configuration parameters (such as input size, number of categories, confidence threshold, etc.), and execution logs of the anomaly handling and rollback mechanism (such as whether the backup model is triggered, and the reason for the downgrade, etc.).

[0059] In one feasible implementation, step S30 includes the following steps: When a user is detected to have a need for video learning, a user profile is obtained, which includes the user's major, permission level, and skill map. The permission level is determined based on role permissions, attribute permissions, and autonomous access permissions. It should be noted that "user specialty" refers to the user's professional field, such as surgery, internal medicine, radiology, etc. "Skill map" represents the user's professional skills and knowledge level, usually represented in a graph format, including the skills and knowledge nodes the user possesses. Role-based permissions are assigned based on the user's role (e.g., doctor, nurse, intern). Attribute-based permissions are assigned based on the user's attributes (e.g., specialty, title, department). Self-access permissions allow users to set and manage their own access permissions. When the medical video sharing device detects a user's need to watch or learn from medical videos, it will obtain the user's profile.

[0060] Based on the user's expertise and the skill graph, determine the medical videos to be recommended within the scope of the permission level; Understandably, medical video sharing devices filter out medical videos that are accessible within the user's permission level based on the user's professional and skill profile, and then select these videos as potential recommendations. Figure 3 , Figure 3 A diagram illustrating the permission management matrix is ​​provided.

[0061] If the user finishes watching the recommended medical video, the user's learning evaluation result is determined based on the user's behavioral indicators of watching the recommended medical video. It should be noted that the behavioral metrics include completion rate, test scores, and experiment improvements. When a user exits the viewing interface, the medical video sharing device collects a series of behavioral metrics that reflect the user's learning behavior and effectiveness. Based on the collected behavioral metrics, it generates learning evaluation results for the user. These evaluation results can be used to understand the user's learning progress and effectiveness in order to further optimize recommendations and learning paths.

[0062] The learning assessment results are fed back to the user's corresponding client, and the skill map is updated based on the learning assessment results.

[0063] Understandably, medical video sharing devices send users' learning assessment results to their devices (such as computers and mobile phones) so that users can view their learning progress and adjust and update their skill maps based on the assessment results to reflect their progress and new skill levels during the learning process.

[0064] Furthermore, the medical video sharing device is also equipped with an initial virtual classroom, which can be configured for student and teacher ends. The medical video sharing device uses WebSocket technology to grant the teacher end the permission to synchronize video playback.

[0065] Furthermore, the initial virtual classroom is also equipped with a case comparison and analysis function. The medical video sharing device extracts the feature vector (case_features) of each case. These features can be surgical steps, instruments used, patient information, etc. A specific algorithm (extract_key_differences) is used to extract key differences (Difference_highlights). These differences can be differences in surgical steps, instruments used, etc.

[0066] In this embodiment, visual-language fusion capabilities are used to simultaneously analyze surgical videos and medical text knowledge, thereby achieving a more accurate understanding of the scene.

[0067] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 4 The medical video sharing method further includes steps S01 to S06, whereby the initial medical video is analyzed frame by frame using a preset intelligent analysis model to obtain the analysis results to be labeled. Prior to the step where the preset intelligent analysis model is based on a preset knowledge graph encoding, the medical video sharing method also includes steps S01 to S06. Step S01: Obtain historical medical videos and perform preliminary analysis on the historical medical videos to obtain the de-identified historical medical videos and preliminary analysis results; It should be noted that the medical video sharing device retrieves previously recorded medical videos from the storage system and performs preliminary processing on the retrieved videos, including scene detection, privacy protection, and quality assessment. During the preliminary analysis and processing, the video content is de-identified to ensure that patient privacy is protected, while generating preliminary analysis results.

[0068] Step S02: The desensitized historical medical videos are divided into stages and risks are identified to obtain the surgical stages and risky actions. Understandably, the medical video sharing device performs two types of analysis on the anonymized historical medical videos: it obtains information on each stage of the surgery through stage division and identifies risky actions in the video through risk identification.

[0069] Step S03: Identify medical entity information in the de-identified historical medical videos; It should be noted that the medical video sharing device performs frame-by-frame image segmentation, using a deep learning-based segmentation model to identify and label medical entities such as surgical instruments, organs, tissues, and lesion areas. The labeling process includes accurately delineating entity boundaries and assigning category information, forming a high-quality training dataset for subsequent training and optimization of the image segmentation model.

[0070] Step S04: Generate operation improvement suggestions for the risky action, and determine risk-related information based on the operation improvement suggestions and the risky action; Understandably, medical video sharing devices generate specific improvement suggestions based on identified risky actions to help users avoid or reduce the occurrence of these risky actions. By combining the operational improvement suggestions with the specific content of the risky actions, detailed information related to these risks is determined.

[0071] Step S05: Based on the preliminary analysis results, the medical entity information, the surgical stage, and the risk-related information, the historical medical video is annotated to obtain the annotated historical medical video; It should be noted that the medical video sharing device uses preliminary analysis results, medical entity information, surgical stage and risk-related information to annotate historical medical videos, ultimately obtaining annotated videos.

[0072] Step S06: Input the labeled historical medical video into the initial intelligent analysis model for iterative training to obtain a preset intelligent analysis model that meets the accuracy requirements.

[0073] Understandably, the medical video sharing device uses labeled historical medical videos as training data to input into the initial intelligent analysis model. After iterative training, an intelligent analysis model with performance meeting the predetermined accuracy requirements is finally obtained, which can be used for subsequent video analysis tasks.

[0074] In this embodiment, professional knowledge such as medical device standards, anatomical knowledge, and surgical procedures are encoded into the model to improve the professionalism and accuracy of recognition.

[0075] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the medical video sharing method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.

[0076] This application also provides a medical video sharing device, please refer to... Figure 5 The medical video sharing device includes: The receiving module 10 is used to receive a video upload request initiated by the user interaction layer, wherein the video upload request includes an initial medical video; The analysis module 20 is used to perform intelligent analysis on the initial medical video to obtain a target medical video labeled with the analysis results, wherein the analysis results include medical entity information, surgical stage and risk-related information. The storage module 30 is used to save the target medical video to a local device so that users in the user interaction layer connected to the cloud server can browse the target medical video.

[0077] Optionally, the analysis module includes: The annotation submodule is used to analyze the segmented medical image sequence through a preset intelligent analysis model to obtain the analysis results to be annotated; and to annotate the analysis results to be annotated on the corresponding video timeline in the initial medical video to obtain the target medical video annotated with the analysis results.

[0078] Optionally, the annotation submodule includes: The extraction unit is used to perform image segmentation and feature extraction on the initial medical video frame by frame using a preset intelligent analysis model to obtain the segmented medical image sequence and dynamic video features; analyze the feature information and temporal relationship of medical entities based on the segmented medical image sequence; and output the analysis results to be labeled based on the feature information and temporal relationship.

[0079] The training unit is used to acquire historical medical videos and perform preliminary analysis on them to obtain desensitized historical medical videos and preliminary analysis results. It performs image segmentation and annotation on the desensitized historical medical videos, marking the boundaries and categories of medical entities such as surgical instruments, organs, tissues, and lesion areas. It further performs stage division and risk identification on the desensitized historical medical videos to obtain surgical stages and risky actions. It generates operation improvement suggestions for the risky actions and determines risk-related information based on these suggestions and actions. Based on the preliminary analysis results, the image segmentation and annotation results, the surgical stages, and the risk-related information, it annotates the historical medical videos to obtain annotated historical medical videos. Finally, it inputs the annotated historical medical videos into the initial intelligent analysis model for iterative training to obtain a preset intelligent analysis model that meets accuracy requirements.

[0080] Optionally, the training unit includes: A sub-unit is constructed to perform semantic segmentation on each frame of the desensitized historical medical video to obtain the segmented target image; key regions are extracted from the target image, and the key regions are accurately labeled to mark the boundaries and category information of medical entities, wherein the key regions include surgical instrument regions and organ lesion regions; an image segmentation model is trained based on the labeled image data, using an encoder-decoder architecture or a Transformer-based segmentation architecture.

[0081] Optionally, the storage module includes: The feedback submodule is used to, when a user's video learning needs are detected, obtain the user's user profile, wherein the user profile includes the user's profession, permission level, and skill graph, and the permission level is determined based on role permissions, attribute permissions, and autonomous access permissions; based on the user's profession and the skill graph, determine the medical videos to be recommended within the permission level range; if the user completes watching the recommended medical videos, determine the user's learning evaluation result based on the user's behavior indicators of watching the recommended medical videos; feed back the learning evaluation result to the user's corresponding user terminal, and update the skill graph based on the learning evaluation result.

[0082] The medical video sharing device provided in this application, employing the medical video sharing method in the above embodiments, can solve the technical problem of medical video sharing. Compared with the prior art, the beneficial effects of the medical video sharing device provided in this application are the same as those of the medical video sharing method provided in the above embodiments, and other technical features in the medical video sharing device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0083] This application provides a medical video sharing device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the medical video sharing method in Embodiment 1 above.

[0084] The following is for reference. Figure 6The diagram illustrates a structural schematic of a medical video sharing device suitable for implementing embodiments of this application. The medical video sharing device in these embodiments may include, but is not limited to, mobile terminals such as mobile phones, laptops, tablets, digital broadcast receivers, PDAs (Personal Digital Assistants), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The medical video sharing device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0085] like Figure 6 As shown, the medical video sharing device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.) that can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the medical video sharing device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. The communication device 1009 allows the medical video sharing device to communicate wirelessly or wiredly with other devices to exchange data. While the figures show medical video sharing devices with various systems, it should be understood that implementing or having all of the systems shown is not required. More or fewer systems may be implemented alternatively.

[0086] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0087] The medical video sharing device provided in this application, employing the medical video sharing method in the above embodiments, can solve the technical problems of medical video sharing. Compared with the prior art, the beneficial effects of the medical video sharing device provided in this application are the same as those of the medical video sharing method provided in the above embodiments, and other technical features in this medical video sharing device are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0088] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0089] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0090] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the medical video sharing method in the above embodiments.

[0091] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0092] The aforementioned computer-readable storage medium may be included in the medical video sharing device; or it may exist independently and not be assembled into the medical video sharing device.

[0093] The aforementioned computer-readable storage medium carries one or more programs that, when executed by the medical video sharing device, cause the medical video sharing device to: receive a video upload request initiated by the user interaction layer, wherein the video upload request includes an initial medical video; perform intelligent analysis on the initial medical video to obtain a target medical video labeled with analysis results, wherein the analysis results include medical entity information, surgical stage, and risk-related information; and save the target medical video locally for users in the user interaction layer connected to the cloud server to browse the target medical video.

[0094] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0095] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0096] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0097] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described medical video sharing method, thereby solving the technical problem of medical video sharing. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the medical video sharing method provided in the above embodiments, and will not be repeated here.

[0098] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the medical video sharing method described above.

[0099] The computer program product provided in this application can solve the technical problem of medical video sharing. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as the beneficial effects of the medical video sharing method provided in the above embodiments, and will not be repeated here.

[0100] The above description is only a part of the embodiments of this application and does not limit the scope of protection of this application. All equivalent structural transformations made under the technical concept of this application and using the content of this application specification and drawings, or direct / indirect applications in other related technical fields, are included in the scope of protection of this application.

Claims

1. A method for sharing medical videos, characterized in that, The medical video sharing method, applied to the cloud service layer, includes: Receive video upload requests initiated by the user interaction layer, where the video upload request includes the initial medical video; The initial medical video is subjected to intelligent analysis to obtain a target medical video labeled with the analysis results, wherein the analysis results include medical entity information, surgical stage and risk-related information; The target medical video is saved locally so that users in the user interaction layer connected to the cloud server can browse the target medical video.

2. The medical video sharing method as described in claim 1, characterized in that, The step of performing intelligent analysis on the initial medical video to obtain a target medical video labeled with the analysis results includes: The initial medical video is segmented frame by frame to obtain a sequence of segmented medical images; The segmented medical image sequence is analyzed using a preset intelligent analysis model to obtain the analysis results to be labeled. The analysis results to be labeled are then labeled on the corresponding video timeline in the initial medical video to obtain the target medical video labeled with the analysis results.

3. The medical video sharing method as described in claim 2, characterized in that, The step of analyzing the segmented medical image sequence using a preset intelligent analysis model to obtain the analysis results to be labeled includes: The medical image sequence is preprocessed to obtain preprocessed image frames; The preprocessed image frame is semantically segmented using a preset intelligent analysis model to identify and segment medical entities in the image frame, and an analysis result labeled with the medical entities is obtained. The preset intelligent analysis model is a deep learning-based image segmentation model, which includes at least medical entities such as surgical instruments, organs and tissues, and lesion areas. The analysis results are post-processed and optimized to obtain the optimized analysis results to be labeled.

4. The medical video sharing method as described in claim 2, characterized in that, Before the step of performing frame-by-frame analysis on the initial medical video after preliminary analysis using a preset intelligent analysis model to obtain the analysis results to be labeled, the following steps are included: Acquire historical medical videos and perform preliminary analysis on the historical medical videos to obtain the anonymized historical medical videos and preliminary analysis results; The desensitized historical medical videos were divided into stages and risks were identified, resulting in surgical stages and risky actions. Identify medical entity information in de-identified historical medical videos; For the risky actions, generate operation improvement suggestions, and based on the operation improvement suggestions and the risky actions, determine risk-related information; Based on the preliminary analysis results, the medical entity information, the surgical stage, and the risk-related information, the historical medical videos are annotated to obtain the annotated historical medical videos. The labeled historical medical videos are input into the initial intelligent analysis model for iterative training to obtain a preset intelligent analysis model that meets the accuracy requirements.

5. The medical video sharing method as described in claim 3, characterized in that, The steps of using a preset intelligent analysis model to perform semantic segmentation on the preprocessed image frame, identify and segment medical entities in the image frame, and obtain analysis results labeled with the medical entities include: The pre-defined intelligent analysis model is used to output a set of candidate targets, including surgical stage and instrument category, for the preprocessed image frames, and the prior confidence of the candidate target set is given. The candidate target set is transformed into region cues in the image space, and the model is driven by the region cues to perform pixel-level segmentation and instantiation separation. Morphological processing is used to improve the quality of the segmentation boundary, resulting in an image sequence with pixel-level masks. Cross-frame tracking is performed on the pixel-level mask to obtain a mask sequence with trajectory sequence number and an image sequence with occurrence duration; Spatiotemporal alignment and conflict fusion correction are performed on the mask sequence with trajectory sequence number and the image sequence with occurrence duration to obtain the corrected image sequence; Based on the corrected image sequence, analysis results labeled with the medical entities are obtained, wherein the analysis results include the start / end time, absolute time and duration of each device and stage, and a visualization resource index; The step of obtaining the analysis results labeled with the medical entities based on the corrected image sequence includes: The preset intelligent analysis model records model information including model version, parameters, and anomaly rollback information, wherein the model information is used to ensure the traceability and quality control of the results.

6. The medical video sharing method as described in claim 1, characterized in that, The video upload request includes user permissions. Following the step of saving the target medical video locally for viewing by a user in the user interaction layer connected to the cloud server, the process includes: When a user is detected to have a need for video learning, a user profile is obtained, which includes the user's major, permission level, and skill map. The permission level is determined based on role permissions, attribute permissions, and autonomous access permissions. Based on the user's expertise and the skill graph, determine the medical videos to be recommended within the scope of the permission level; If the user finishes watching the recommended medical video, the user's learning evaluation result is determined based on the user's behavioral indicators of watching the recommended medical video. The learning assessment results are fed back to the user's corresponding client, and the skill map is updated based on the learning assessment results.

7. A medical video sharing device, characterized in that, The device includes: The receiving module is used to receive video upload requests initiated by the user interaction layer, wherein the video upload request includes the initial medical video; The analysis module is used to perform intelligent analysis on the initial medical video to obtain a target medical video labeled with the analysis results, wherein the analysis results include medical entity information, surgical stage and risk-related information; The storage module is used to save the target medical video locally so that users in the user interaction layer connected to the cloud server can browse the target medical video.

8. A medical video sharing device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the medical video sharing method as described in any one of claims 1 to 6.

9. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the medical video sharing method as described in any one of claims 1 to 6.

10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the medical video sharing method as described in any one of claims 1 to 6.