A time slice intelligent labeling method for medical teaching video knowledge tags

CN122332606BActive Publication Date: 2026-09-11SHANGHAI LINGLI HEALTH MANAGEMENT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610789091.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-03
Publication Date
2026-09-11
Estimated Expiration
2046-06-03

AI Technical Summary

Technical Problem

同理,由于当前时间片中未直接出现“ST段抬高”这一术语,但明显的,从前文讲解、当前画面指示区域以及指代性语言可以确定,该时间片仍然属于“ST段抬高”的知识讲解内容;若此时系统不能识别该类隐含标签,会导致学生检索“ST段抬高”时,依旧无法跳转到真正具有判读价值的圈画片段

Benefits of technology

[0027]本发明分时间片标注结果不仅包含最终标签,还包含标签来源、画面依据和置信度信息,从而便于后续用于知识点检索、教学视频片段跳转、课程目录生成和人工复核。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122332606B_ABST
    Figure CN122332606B_ABST
Patent Text Reader

Abstract

The application provides a knowledge label time slice intelligent labeling method for medical teaching videos, and relates to the field of video data processing and labeling. The method comprises the following steps: obtaining medical teaching video analysis information, wherein the video analysis information comprises voice text information, picture content information and teacher instruction action information; determining a current basic time slice as a to-be-backfilled time slice according to the voice text information, wherein the to-be-backfilled time slice is a time slice in which there is referential explanation and no complete explicit medical knowledge label is identified; determining a picture indication area in the picture of the to-be-backfilled time slice according to the teacher instruction action information, and extracting regional medical clues; associating and matching the regional medical clues with previous medical knowledge labels to determine an implied medical knowledge label corresponding to the to-be-backfilled time slice; and binding the implied medical knowledge label with the to-be-backfilled time slice to generate a time slice labeling result, so that the missing labeling of medical knowledge labels caused by the referential explanation of teachers in medical teaching videos can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of video data processing and medical teaching resource annotation technology, and in particular to a method for intelligent time-slice annotation of knowledge tags for medical teaching videos. Background Technology

[0002] With the development of digital medical teaching resources, courses such as anatomy, imaging, pathology, clinical diagnostics, electrocardiogram interpretation, ultrasound examination, and operational demonstrations are increasingly stored and disseminated in video format. To facilitate students' retrieval, review, and navigation by knowledge point, existing technologies typically perform speech recognition, subtitle recognition, courseware text recognition, or keyframe analysis on medical teaching videos. Based on the recognized medical terminology, corresponding video tags are generated, and these tags are then associated with the corresponding video time slices.

[0003] Current methods for annotating medical teaching videos typically rely on the direct appearance of medical terms within the current time slice as the basis for tag generation. For example, after recognizing medical terms such as "hydronephrosis," "renal pelvis dilation," "ST segment elevation," and "gallbladder wall thickening" in speech-to-text, subtitles, or courseware text, the system uses these terms as the knowledge tag for the current time slice. This method is applicable to teaching videos to some extent; however, in actual medical teaching, teachers do not always repeatedly pronounce complete medical terms. Especially when explaining medical images, anatomical structures, pathological slides, electrocardiogram waveforms, ultrasound images, or operational demonstrations, teachers often first explain the medical term in the previous time slice, and then use descriptive language such as "here," "this area," "this sign," "this location," or "this manifestation" in the current time slice, along with mouse hovering, drawing paths, annotation pen pointing, local zooming, or highlighting areas for explanation. In this case, although the complete medical term does not reappear in the current time slice, the teacher is actually explaining the specific manifestation or interpretation location of the previous medical knowledge point in the image.

[0004] For example, in an ultrasound teaching video, the teacher initially says that "hydronephrosis is characterized by dilation of the renal pelvis and calyces," then circles the anechoic dilation area on the ultrasound image in the current time slice and says, "You can see obvious dilation here." Since "hydronephrosis" or "renal pelvis dilation" does not appear again in the current audio text, if the system generates labels based only on the explicit terms in the current time slice, this key image explanation segment may not be labeled as a "hydronephrosis" related segment.

[0005] For example, in an electrocardiogram (ECG) instructional video, the teacher first explains, "Let's look at the typical manifestations of ST segment elevation," then circles the corresponding waveform on the ECG with a marker, saying, "This position has exceeded the baseline." Similarly, although the term "ST segment elevation" does not appear directly in the current time slice, it is clear from the preceding explanation, the current screen indicator area, and the referential language that this time slice still belongs to the knowledge explanation of "ST segment elevation." If the system cannot recognize this type of implicit label, students searching for "ST segment elevation" will still not be able to jump to the circled segment that is truly valuable for interpretation.

[0006] Therefore, existing technologies have the following shortcomings: On the one hand, when the current time slice lacks direct medical terminology, the system is prone to missing medical knowledge content expressed by teachers through descriptive explanations and on-screen cues; on the other hand, simply mechanically extending the labels of the previous time slice to subsequent segments may lead to over-expansion of labels to irrelevant time slices, resulting in mislabeling. In other words, existing technologies lack an intelligent time-slice labeling method that can, based on the identification of descriptive explanation segments, combine prior medical knowledge labels, the current on-screen cues, and regional medical clues to fill in implicit medical labels for the current time slice. Summary of the Invention

[0007] This application provides a time-slice intelligent annotation method for knowledge tags in medical teaching videos. When the complete medical terminology does not appear directly in the current time slice of a medical teaching video, the method combines the preceding medical knowledge tags and the current screen indication area to annotate the medical knowledge expressed by the teacher through descriptive explanations in time slices.

[0008] Firstly, this application provides a method for intelligent time-segmented annotation of knowledge tags for medical teaching videos, which can be executed by an annotation device. In this method, the annotation device acquires video parsing information corresponding to the medical teaching video, including audio-text information, screen content information, and teacher instruction / action information. Based on the audio-text information, the annotation device determines the current base time segment as a time segment to be filled, which is a time segment with referential explanations but without a fully identified explicit medical knowledge tag. Based on the teacher instruction / action information, the annotation device determines a screen instruction area in the screen of the time segment to be filled and extracts regional medical clues based on the screen instruction area. The annotation device associates and matches the regional medical clues with the preceding medical knowledge tags before the time segment to be filled, obtaining a backfilling association result for candidate medical knowledge tags, and determines the implicit medical knowledge tag corresponding to the time segment to be filled based on the backfilling association result. The annotation device binds the implicit medical knowledge tag to the time segment to be filled, generating the time-segmented annotation result of the medical teaching video.

[0009] In the above method, the annotation device does not simply search for the direct appearance of medical terms in the current time slice. Instead, when there is a referential explanation in the current time slice but a complete explicit medical knowledge label is missing, it further combines the preceding medical knowledge labels and the current screen indication area to determine the implicit medical knowledge labels. This method allows for the back-filling of annotations for medical knowledge explanations delivered by teachers using expressions such as "here," "this area," and "this sign," along with on-screen instructions, reducing the chances of key medical teaching segments being missed due to the lack of repeated medical terminology.

[0010] In one possible design, obtaining the video parsing information corresponding to the medical teaching video includes: dividing the medical teaching video into time slices to obtain multiple basic time slices; extracting speech-to-text, subtitle text, courseware page images, on-screen text information, and teacher instruction action information from the multiple basic time slices, wherein the teacher instruction action information includes at least one of teacher instruction action trajectory, local magnified area information, and on-screen highlight information; using the speech-to-text and subtitle text as the speech text information, and the courseware page images and on-screen text information as the on-screen content information.

[0011] Through this design, the annotation device can uniformly convert the audio, subtitles, courseware pages, on-screen text, and teacher instructions in medical teaching videos into video analysis information that can be used for time-slice annotation, so that the subsequent filling of implicit medical knowledge tags can rely on both text clues and on-screen instructions.

[0012] In one possible design, determining the current basic time slice as the time slice to be filled based on the voice-text information includes: recognizing referential expressions in the voice-text information of the current basic time slice; when a referential expression is recognized in the voice-text information and no complete explicit medical knowledge tag that can be directly mapped to a preset medical knowledge node is recognized in the current basic time slice, determining whether there is teacher instruction action information in the current basic time slice; when there is teacher instruction action information in the current basic time slice and a preceding medical knowledge tag exists before the current basic time slice, determining the current basic time slice as the time slice to be filled.

[0013] With this design, the labeling device can first filter out the time slices that truly need to be backfilled, avoiding implicit label inference for all time slices; for time slices that already have complete explicit medical knowledge labels, the labels can be directly bound; for time slices that have referential expressions but lack complete medical terms, the backfilling process can then begin.

[0014] In one possible design, the preceding medical knowledge tags are determined as follows: medical entity identification is performed on at least one preceding time slice located before the time slice to be backfilled to obtain explicit medical entities in the at least one preceding time slice; the explicit medical entities are standardized to obtain preceding medical knowledge tags; and a preceding medical knowledge tag cache queue is established according to the occurrence time, source modality, and standardized name of the preceding medical knowledge tags.

[0015] This design allows the annotation device to use medical knowledge tags that have appeared previously in time slices as candidate sources for backfilling. For example, if a teacher mentions "hydronephrosis is characterized by dilation of the renal pelvis and calyces" in a previous time slice, and only says "obvious dilation can be seen here" in a later time slice, the cached queue of previous medical knowledge tags can provide candidate backfill tags for the later time slice.

[0016] In one possible design, a cache queue of preceding medical knowledge tags is established according to their occurrence time, source modality, and standardized name. Specifically, this includes: adding the preceding medical knowledge tags to the cache queue within a preset backtracking time range; setting a time decay value for the preceding medical knowledge tags based on the time interval between the preceding medical knowledge tags and the time slice to be backfilled; wherein, the larger the time interval, the lower the retention weight corresponding to the time decay value.

[0017] This design allows the labeling device to limit the time range of candidate labels for backfilling, preventing medical knowledge labels that are too far removed from the current time slice from being incorrectly carried over to the current time slice. Simultaneously, the time decay value allows medical knowledge labels appearing in more recent time slices to have higher reference value in backfilling judgments.

[0018] In one possible design, determining the screen indication area in the screen of the time slice to be filled based on the teacher's instruction action information includes: detecting the teacher's instruction action information to determine the type of teacher's instruction action in the time slice to be filled; determining the screen indication area in the screen of the time slice to be filled based on the type of teacher's instruction action; wherein, the type of teacher's instruction action includes at least one of mouse hover, mouse drawing, annotation pen trajectory, local zoom, and screen highlighting.

[0019] Through this design, the annotation device can determine the actual area of ​​the screen that the teacher is pointing to or emphasizing in the current time slice. This makes the backfilling of implicit medical knowledge labels no longer solely dependent on the preceding text, but further constrained by the current screen indication area, thereby reducing mislabeling caused by the mechanical continuation of preceding labels.

[0020] In one possible design, the step of extracting regional medical clues based on the screen indication area includes: performing at least one of the following processing on the screen indication area: local text recognition, image annotation recognition, and extraction of adjacent courseware information to obtain the regional medical clues; wherein, the regional medical clues include at least one of the following: medical text within the screen indication area, medical text within the adjacent area of ​​the screen indication area, image annotation text, arrow pointing text, legend text, courseware title, and chapter prompt information.

[0021] This design allows the annotation device to associate medical clues extracted from the current screen's indicated area with preceding medical knowledge tags, rather than directly selecting the most recently appearing tag. This enables the device to determine the implicit medical knowledge tag that best suits the current subject of explanation when multiple preceding medical knowledge tags exist simultaneously, by combining the content of the screen area.

[0022] In one possible design, the step of associating and matching the regional medical clues with the preceding medical knowledge tags before the time slice to be filled, and determining the implicit medical knowledge tags corresponding to the time slice to be filled, includes: matching the regional medical clues with candidate medical knowledge tags in the preceding medical knowledge tag cache queue; calculating the filling relevance degree of each candidate medical knowledge tag relative to the time slice to be filled based on regional text matching degree, courseware topic relevance, preceding occurrence intensity, time decay value, and referential explanation text relevance; and determining the candidate medical knowledge tags whose filling relevance degree meets a preset threshold condition as the implicit medical knowledge tags.

[0023] Through this design, the annotation device can determine implicit medical knowledge tags by comprehensively considering regional content, courseware theme, strength of preceding tags, time distance, and semantics of referential text, making the backfilling process have a calculable basis for judgment and avoiding reliance solely on whether keywords are the same.

[0024] In one possible design, determining the candidate medical knowledge tag whose backfill correlation meets a preset threshold condition as the implicit medical knowledge tag includes: when the backfill correlation of a candidate medical knowledge tag is greater than or equal to a first threshold, determining the candidate medical knowledge tag as the implicit medical knowledge tag; when the backfill correlation of multiple candidate medical knowledge tags is greater than or equal to the first threshold, and the difference in backfill correlation between the multiple candidate medical knowledge tags is less than a second threshold, determining the multiple candidate medical knowledge tags as candidate backfill tags, and generating a review mark for the time slice to be backfilled.

[0025] With this design, the annotation device can automatically bind implicit medical knowledge tags when the backfill results are clear, and retain candidate backfill tags and generate verification marks when multiple candidate tags are difficult to distinguish, thus balancing automatic annotation efficiency and annotation reliability.

[0026] In one possible design, generating the time-slice annotation results of the medical teaching video includes: outputting the start and end times of the time slice to be backfilled, explicit medical knowledge tags, implicit medical knowledge tags, the preceding source time slice corresponding to the implicit medical knowledge tags, the location of the screen indicator area, and the backfill confidence; wherein, the backfill confidence is determined based on the backfill correlation degree corresponding to the implicit medical knowledge tags.

[0027] The time-slice annotation results of this invention not only include the final labels, but also the label source, the basis of the image, and the confidence level information, which facilitates subsequent use for knowledge point retrieval, teaching video clip jumping, course catalog generation, and manual review. Attached Figure Description

[0028] Figure 1 A schematic diagram of a time-slice annotation system for medical teaching videos provided in this application embodiment; Figure 2 A flowchart illustrating a time-slice intelligent annotation method for knowledge tags in medical teaching videos, provided as an embodiment of this application; Figure 3 This application provides a schematic diagram of a process for determining a time slice to be backfilled, as illustrated in an embodiment of the present application. Figure 4 A schematic diagram of a pre-processor medical knowledge tag cache queue provided in an embodiment of this application; Figure 5 A schematic diagram illustrating the mapping between a teacher's instruction action and a screen instruction area, provided in an embodiment of this application; Figure 6 A schematic diagram illustrating the calculation of the relevance of implicit medical knowledge tag backfilling in an embodiment of this application; Figure 7 This is a schematic diagram of the output of time slice annotation results provided in an embodiment of this application. Detailed Implementation

[0029] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings. In the description of the embodiments of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance. The term "at least one" refers to one or more, and "more than one" refers to two or more.

[0030] The intelligent time-slice annotation method for knowledge tags in medical teaching videos provided in this application can be executed by an annotation device. The annotation device can be a server, a teaching resource management platform, a video annotation platform, an edge computing device, or an electronic device with video annotation software installed. The annotation device can acquire medical teaching videos, perform video parsing, and output annotation results with the start and end times of the time slices and medical knowledge tags.

[0031] In this application embodiment, medical teaching videos refer to videos used for medical teaching, medical training, or medical course resource management, including at least one of medical imaging teaching videos, anatomical structure teaching videos, pathological slide teaching videos, electrocardiogram teaching videos, clinical case teaching videos, and operation demonstration teaching videos.

[0032] A base time slice is the smallest unit of analysis obtained by initially dividing a medical teaching video. Base time slices can be obtained based on speech sentence boundaries, subtitle time boundaries, or preset short windows. Base time slices are used for text recognition, image recognition, and gesture recognition, and are not necessarily directly equivalent to the final output teaching segment. The current base time slice refers to a base time slice that is currently being analyzed by the annotation device.

[0033] The time slice to be filled refers to the time slice in which there is a referential explanation, but no complete explicit medical knowledge label is identified in the current basic time slice, and the current basic time slice contains teacher instruction action information and has a candidate medical knowledge label in the preceding time slice.

[0034] A medical entity refers to the original medical object identified from audio text, subtitle text, on-screen text, or image annotations, including at least one of the following: disease name, anatomical structure, symptoms and signs, examination items, imaging signs, operational procedures, treatment methods, and medication information. A medical knowledge tag refers to a knowledge node obtained after standardizing a medical entity. Standardization processing may include synonym unification, abbreviation unification, hierarchical mapping, and course knowledge node mapping.

[0035] A complete explicit medical knowledge tag refers to a tag in the speech-to-text, subtitle text, or on-screen text information in the current base time slice that can be directly mapped to a preset medical knowledge node after medical entity recognition and standardization. For example, "hydronephrosis," "renal pelvis dilation," "ST segment elevation," "gallbladder wall thickening," and "acute myocardial infarction" are complete explicit medical knowledge tags. If the current text only contains pronouns, degree words, directional words, location words, or symptom words that do not contain medical objects, such as "dilation here," "this position exceeds the baseline," or "this manifestation is obvious," then it does not constitute a complete explicit medical knowledge tag.

[0036] Implicit medical knowledge tags refer to medical knowledge tags that do not appear directly in the current base time slice but can be identified through previous medical knowledge tags and the current screen indicator area; For example, if a teacher mentions in a previous segment that "hydronephrosis is characterized by dilation of the renal pelvis and calyces," and in a later segment only says "obvious dilation can be seen here" and circles the dilated area, then "hydronephrosis" in the later segment can be used as an implicit medical knowledge label.

[0037] Preceding medical knowledge tags refer to medical knowledge tags identified in time slices preceding the time slice to be backfilled. Candidate medical knowledge tags refer to medical knowledge tags in the preceding medical knowledge tag cache queue that participate in the current backfilling matching. The preceding medical knowledge tag cache queue is a set of medical knowledge tags cached by the annotation device in chronological order, preceding the time slice to be backfilled. Each tag in the cache queue may include tag name, normalized name, occurrence time, source modality, preceding time slice, and occurrence count.

[0038] The indicated area on the screen refers to the area within the frame of the time slice to be filled in, determined based on the teacher's instruction gestures. Teacher instruction gestures refer to teaching emphasis information that forms a locatable area on the video screen, including at least one of the following: mouse hover, mouse drawing, pen stroke marking, local zoom, and screen highlighting.

[0039] Regional medical cues refer to text, image annotations, legends, arrow directions, structural identifiers, courseware titles, or chapter prompts extracted from the indicated area on the screen and its adjacent areas, which can be used to determine medical knowledge tags.

[0040] Backfill relevance refers to the calculated value used to select implicit medical knowledge tags from candidate medical knowledge tags. Backfill confidence refers to the output level or numerical value mapped from the backfill relevance of the finally selected tag.

[0041] Figure 1 This is a schematic diagram of the structure of a medical teaching video time-slice annotation system provided in an embodiment of this application; the system may include a video acquisition module, a video parsing module, a time-slice identification module to be filled, a preceding label caching module, an indication area determination module, an area clue extraction module, a hidden label filling module, and a result output module.

[0042] The video acquisition module acquires medical teaching videos. The video parsing module extracts audio-visual information, video content, and teacher gestures. The time-slice identification module determines whether the current base time-slice is a time-slice to be filled based on the audio-visual information. The preceding label caching module establishes a preceding medical knowledge label cache queue. The instruction region determination module determines the on-screen instruction region based on the teacher gestures. The region cue extraction module extracts regional medical cues from the on-screen instruction regions. The implicit label backfilling module associates and matches regional medical cues with preceding medical knowledge labels to determine implicit medical knowledge labels. The result output module outputs the time-slice annotation results.

[0043] The above modules can be implemented in software, hardware, or a combination of both. This application does not limit the specific deployment form of each module.

[0044] Figure 2 The flowchart illustrates a method for intelligent time-slice annotation of knowledge tags for medical teaching videos, as provided in this application embodiment. The method may include the following steps.

[0045] S101: The annotation device acquires video analysis information corresponding to the medical teaching video; The video analysis information includes audio and text information, video content information, and teacher instruction and action information; Voice-to-text information can include speech-transcribed text and subtitle text; The content information on the screen can include images on the courseware page and text information on the screen; Teacher instruction actions may include at least one of the following: mouse hover trajectory, mouse drawing trajectory, pen annotation trajectory, zoomed-in area, and highlighted area on the screen.

[0046] S102: The annotation device determines the current baseline time slice as a time slice to be filled based on the voice and text information. In some embodiments, the annotation device can identify whether there is a referential expression in the current baseline time slice and determine whether a complete explicit medical knowledge tag has not been identified in the current baseline time slice. When the current baseline time slice contains a referential expression, does not identify a complete explicit medical knowledge tag, contains teacher instruction action information, and has a preceding medical knowledge tag, the annotation device can determine the current baseline time slice as a time slice to be filled.

[0047] S103: The annotation device determines the screen indication area in the screen of the time slice to be filled based on the teacher's instruction action information.

[0048] For example, when a mouse hover is detected, the annotation device can determine a preset range centered on the hover point as the screen indication area; When a mouse circle or pen stroke is detected, the annotation device can use the area enclosed by the circle or pen stroke as the display indication area. When a local magnification or highlight is detected, the annotation device can use the magnified or highlighted area as the image indication area.

[0049] S104: The annotation device extracts regional medical clues based on the indicated area on the screen; the regional medical clues may include at least one of the following: medical text in the indicated area on the screen, medical text in the adjacent area, image annotation text, arrow pointing text, legend text, courseware title, and chapter prompt information.

[0050] S105: The annotation device will associate and match regional medical clues with preceding medical knowledge tags to obtain the backfill association results of candidate medical knowledge tags, and determine the implicit medical knowledge tags corresponding to the time slice to be backfilled based on the backfill association results. In some embodiments, the annotation device can calculate the backfilling relevance of each candidate medical knowledge tag relative to the time slice to be backfilled based on the region text matching degree, courseware topic relevance, preceding occurrence intensity, time decay value, and referential explanation text relevance, and determine the candidate medical knowledge tags whose backfilling relevance meets the preset threshold condition as implicit medical knowledge tags.

[0051] S106: The annotation device binds implicit medical knowledge tags to the time slices to be filled, generating time slice annotation results for medical teaching videos; The time slice annotation results can include the start and end times of the time slice to be backfilled, explicit medical knowledge tags, implicit medical knowledge tags, the preceding source time slices corresponding to the implicit medical knowledge tags, the location of the on-screen indicator area, and the backfill confidence level.

[0052] In some embodiments, the annotation device can divide the medical teaching video into time slices to obtain multiple basic time slices; the basic time slices can be determined according to the boundaries of speech sentences, the boundaries of subtitle time, or a preset short window; optionally, the preset short window can be one second, two seconds, five seconds, or other time lengths suitable for parsing medical teaching videos.

[0053] For each basic time slice, the annotation device can extract speech-to-text, subtitle text, courseware page images, on-screen text information, and teacher instruction gesture information. Specifically, the speech-to-text can be obtained by recognizing the teacher's audio using a speech recognition model; the subtitle text can be derived from embedded subtitles in the video or automatically generated subtitles; the courseware page images can be obtained through keyframe extraction; the on-screen text information can be obtained through text recognition of the courseware page images; and the teacher instruction gesture information can include at least one of the following: the trajectory of the teacher's instruction gesture, information about magnified local areas, and on-screen highlight information.

[0054] The annotation device can use speech-to-text and subtitle text as speech-text information, courseware page images and on-screen text information as on-screen content information, and teacher gesture information as information for determining on-screen instruction areas. Through these methods, the annotation device can provide a unified data foundation for subsequent time-slice recognition and implicit medical knowledge tag backfilling.

[0055] In some embodiments, teacher instruction action information may come from video frame detection results, or from interaction logs recorded by courseware playback software, recording software, or teaching whiteboard system. If the teacher instruction action information comes from video frame detection results, the annotation device can detect it through changes in cursor position in consecutive frames, newly added annotation lines on the screen, changes in local magnified boxes or highlighted areas; if the teacher instruction action information comes from interaction logs, the annotation device can read the action type, action time, and area coordinates.

[0056] Figure 3 This is a schematic diagram of a process for determining a time slice to be filled, provided for an embodiment of this application; in some embodiments, the annotation device can perform referential expression recognition on the voice and text information of the current basic time slice; Referential expressions can include pronouns that indicate areas of the screen, objects in the image, medical signs, or locations of operations, such as "here," "this area," "this location," "this sign," "this manifestation," "this part," "this line," "look here," "look at this place," etc.

[0057] The annotation device can also determine whether a complete explicit medical knowledge tag is identified in the current basic time slice; a complete explicit medical knowledge tag refers to a medical term that can directly represent a disease, anatomical structure, examination item, imaging sign, operation procedure or treatment method, and can be directly mapped to a preset medical knowledge node; For example, terms like "hydronephrosis," "renal pelvis dilation," "ST segment elevation," "gallbladder wall thickening," and "acute myocardial infarction" can be used as complete explicit medical knowledge tags.

[0058] In some embodiments, the annotation device may identify a current base time slice as a time slice to be backfilled when the following conditions are met: First, the voice text information contains at least one pronoun from a pre-defined list of pronouns; Second, the current basic time slice does not identify complete explicit medical knowledge tags that can be directly mapped to medical knowledge nodes; Third, the current basic time slice contains information about teacher instructions and actions; Fourth, there is at least one candidate medical knowledge tag in the preceding medical knowledge tag cache queue.

[0059] For example, if the audio text of the current baseline time slice is "A clear expansion can be seen here", which includes the referential expression "here", but does not directly contain complete medical terms such as "hydronephrosis" or "renal pelvis dilation", and there is a mouse circle trajectory in the current baseline time slice, and the "hydronephrosis" label appeared in the previous time slice, then the annotation device can determine that the current baseline time slice is a time slice to be filled.

[0060] When the current baseline time slice contains only general connecting phrases, or while it contains referential expressions, it lacks teacher instruction information, or the preceding medical knowledge tag cache queue is empty, the annotation device will not identify the current baseline time slice as a time slice to be backfilled. When a complete explicit medical knowledge tag has been identified for the current baseline time slice, the annotation device can directly bind the complete explicit medical knowledge tag to the current baseline time slice without executing the implicit medical knowledge tag backfilling process.

[0061] Figure 4 This is a schematic diagram of a preceding medical knowledge tag cache queue provided for an embodiment of this application; in some embodiments, the tagging device can perform medical entity identification on at least one preceding time slice located before the time slice to be backfilled, and obtain explicit medical entities in the preceding time slice; the medical entity may include at least one of disease name, anatomical structure, symptoms and signs, examination items, imaging signs, operation steps, treatment methods and medication information.

[0062] In some embodiments, the annotation device may pre-configure a medical knowledge tag library. The medical knowledge tag library includes standard medical knowledge nodes, synonyms, abbreviations, hierarchical relationships, and course chapter associations. Medical entity recognition results can first be precisely matched against standard medical knowledge nodes; if no precise match is found, then matching with synonyms and abbreviations is performed; if still no match is found, candidate standardized names are determined based on the course chapter to which the entity belongs and the contextual semantics. Text fragments that cannot be mapped to standard medical knowledge nodes in the medical knowledge tag library are not considered complete explicit medical knowledge tags.

[0063] The annotation device can standardize explicit medical entities to obtain preceding medical knowledge tags; specifically, the standardization process can include synonym unification, abbreviation unification, hierarchical relationship mapping, and course knowledge node mapping. For example, "ST elevation" and "ST segment elevation" can be grouped under the same standardized label, and "renal pelvis dilation" can be associated with "hydronephrosis".

[0064] The annotation device can establish a cache queue of preceding medical knowledge tags based on their occurrence time, source modality, and standardized name. The source modality can include audio sources, subtitle sources, courseware text sources, and video annotation sources. Each item in the preceding medical knowledge tag cache queue can include the tag name, standardized name, occurrence time, source modality, preceding time slice, and number of occurrences.

[0065] In some embodiments, the labeling device can add preceding medical knowledge tags to the preceding medical knowledge tag cache queue within a preset backtracking time range. The preset backtracking time range can be determined according to the average pacing of the medical teaching video, for example, it can be set to ten seconds, twenty seconds, thirty seconds or a complete group of explanatory sentences before the current time slice.

[0066] The annotation device can set a time decay value for the preceding medical knowledge tag based on the time interval between the preceding medical knowledge tag and the time slice to be filled. The closer the preceding medical knowledge tag is to the time slice to be filled, the higher the retention weight corresponding to the time decay value; the farther the preceding medical knowledge tag is from the time slice to be filled, the lower the retention weight corresponding to the time decay value.

[0067] For example, if the preceding medical knowledge tag "hydronephrosis" appears three seconds before the time slice to be filled, and the preceding medical knowledge tag "urinary tract stones" appears twenty-five seconds before the time slice to be filled, then, all other things being equal, "hydronephrosis" has a higher retention weight than "urinary tract stones." This method can reduce the possibility of outdated medical knowledge tags being incorrectly carried over to the current time slice to be filled.

[0068] In one example, the time decay value D can be determined as follows: D equals 1 divided by 1 and the sum of the normalized value of the time interval; where the time interval is the difference between the most recent occurrence time of the candidate medical knowledge tag and the start time of the time slice to be filled; the larger the time interval, the smaller D is. It should be noted that the above calculation method is only an example. In practice, exponential decay, linear decay, or piecewise decay methods can also be used to determine the time decay value.

[0069] Figure 5 This is a schematic diagram illustrating the mapping between teacher instruction actions and on-screen instruction areas, provided as an embodiment of this application. In some embodiments, the annotation device can detect teacher instruction action information to determine the type of teacher instruction action in the time slice to be filled. The type of teacher instruction action may include at least one of mouse hover, mouse drawing, annotation pen trajectory, local zoom, and on-screen highlighting.

[0070] When the mouse pointer is detected to have lingered in a certain area for more than a preset time, the annotation device can use that lingering point as the center and take a preset radius or preset rectangular area as the screen indication area corresponding to the mouse lingering. When the mouse trajectory or annotation pen trajectory is detected to form a closed or semi-closed trajectory, the annotation device can determine the area enclosed by the trajectory or the rectangular area circumscribed by the trajectory as the screen indication area to be drawn.

[0071] When the annotation pen trajectory does not form a closed region but covers a linear or boundary structure, the annotation device can take the outer region of the trajectory-covered point set and extend it outward by a preset margin as the screen indication area. When a local magnified area is detected in the screen, the annotation device can determine the original image area or the magnified display area corresponding to the magnified window as the screen indication area.

[0072] When a highlighted box, arrow, or color-emphasized area is detected in the courseware screen, the annotation device can identify the area covered by the highlighted box, the area pointed to by the arrow, or the area emphasized by the color as the screen indication area. If multiple teacher instruction actions occur simultaneously in the same time slice to be filled in, the annotation device can determine the main screen indication area according to the priority of local zoom, mouse circle, annotation pen trajectory, screen highlight, and mouse hover; or it can merge multiple areas into a composite screen indication area.

[0073] For example, in an ultrasound teaching video, the teacher uses the mouse to circle an anechoic expansion area in the ultrasound image and says, "You can see obvious expansion here." The annotation device can then determine the expansion area as the screen indication area based on the mouse circle and use this area as the screen reference for subsequent filling in of implicit medical knowledge labels.

[0074] In some embodiments, the annotation device may perform at least one of the following processing on the indicated area of ​​the screen: local text recognition, image annotation recognition, and extraction of information from adjacent courseware, to obtain regional medical clues.

[0075] Local text recognition can be used to identify medical text within or near the indicated area on the screen. Image annotation recognition is not medical image diagnosis recognition, but rather the identification of arrows, annotation lines, brackets, boundary lines, legend labels, and their corresponding text in teaching screens. Adjacent courseware information extraction can be used to extract the title, image title, legend description, adjacent text, and chapter prompts of the page containing the indicated area on the screen.

[0076] Regional medical cues may include at least one of the following: medical text within the indicated area on the screen, medical text in the adjacent area, image annotation text, arrow pointing text, legend text, courseware title, and chapter prompt information.

[0077] In some embodiments, the adjacent area can be the area obtained by extending the screen indicator area outward by a preset pixel range, or it can be a text area that is connected to the screen indicator area by an arrow, annotation line, or bracket.

[0078] For example, if the text near the indicated area on the screen includes "renal pelvis dilation" and the preceding medical knowledge tag cache queue includes "hydronephrosis", "ureteral stones" and "urinary tract obstruction", then the annotation device can determine that "hydronephrosis" has a high regional correlation with the indicated area on the screen.

[0079] Figure 6 This is a schematic diagram illustrating the calculation of the relevance of implicit medical knowledge tag backfilling, provided in an embodiment of this application. In some embodiments, the annotation device can match regional medical clues with candidate medical knowledge tags in the preceding medical knowledge tag cache queue, and calculate the backfilling relevance of each candidate medical knowledge tag relative to the time slice to be backfilled based on regional text matching degree, courseware topic relevance, preceding occurrence intensity, time decay value, and referential explanation text relevance.

[0080] In some embodiments, the association matching between regional medical cues and preceding medical knowledge tags can be achieved using rule matching, knowledge graph matching, similarity calculation, or a trained multimodal matching model; wherein: Rule matching can determine the matching result based on whether there are identical words, synonyms, abbreviations, or hierarchical relationships between regional medical clues and candidate medical knowledge tags; Knowledge graph matching can determine the matching result based on the distance between the nodes of the two entities in the medical knowledge graph. Similarity calculation can determine the matching result based on text similarity, semantic similarity, or the medical relevance between region clues and candidate labels; The multimodal matching model can take regional medical clues, current referential explanatory text and candidate medical knowledge tags as input, and output the backfilling association results of candidate medical knowledge tags; The methods described above are all existing conventional techniques and will not be elaborated upon here.

[0081] In one optional example, the backfill correlation degree R can be determined according to the following formula: R = aM + bT + cF + dD + eS; Wherein, R is the backfill relevance; M is the regional text matching degree, representing the similarity between the candidate medical knowledge tag and the text in the on-screen indicator area, adjacent text, or image annotation text; T is the courseware topic relevance, representing the degree of relevance between the candidate medical knowledge tag and the current courseware title, chapter title, or course topic; F is the preceding occurrence intensity, representing the number of occurrences, duration, and number of source modalities of the candidate medical knowledge tag within the preset backtracking time range; D is the time decay value, representing the time distance influence of the candidate medical knowledge tag from the time slice to be backfilled; S is the referential explanation text relevance, representing the semantic relevance between the candidate medical knowledge tag and the current referential explanation statement; a, b, c, d, and e are weights, and the sum of a, b, c, d, and e is 1.

[0082] In some embodiments, M, T, F, D, and S can all be normalized to between 0 and 1; weights a, b, c, d, and e can be pre-configured or obtained statistically based on manually annotated samples; the first threshold and the second threshold can be set according to the course type, annotation accuracy requirements, or manual verification results. For example: The region text matching degree M can be determined as follows: M is set to 1 when the standardized name of the candidate medical knowledge tag is exactly the same as the medical text in the regional medical clue; When candidate medical knowledge tags and regional medical clues have a synonym relationship or a hierarchical relationship, M is taken as 0.7 to 0.9; When there is only partial overlap in keywords between the two, M is between 0.3 and 0.7; When there is no medical semantic connection between the two, M is set to 0.

[0083] The relevance T of the courseware topic can be determined based on the distance in the medical knowledge graph between candidate medical knowledge tags and courseware titles, chapter titles, or course topics: When the two are the same knowledge node, T is 1; when they are direct hierarchical nodes, T is 0.8; when they are related nodes under the same chapter, T is 0.5 to 0.7; when they are unrelated, T is 0.

[0084] The preceding intensity F can be determined based on the number of occurrences, the number of source modes, and the duration, i.e.: In one example, F = pC + qL + rU, where C is the normalized value of the number of occurrences of the candidate medical knowledge tag within the preset backtracking time range, L is the normalized value of the duration of the candidate medical knowledge tag in the preceding time slice, U is the normalized value of the number of source modalities of the candidate medical knowledge tag, and p, q, and r are weights, and the sum of p, q, and r is 1.

[0085] The relevance S of the referential explanatory text can be determined based on the association between medical sign words in the referential explanatory sentences and candidate medical knowledge tags: for example, "dilation" is highly associated with "hydronephrosis" and "renal pelvis dilation", and "above the baseline" is highly associated with "ST segment elevation".

[0086] The labeling device can identify candidate medical knowledge tags that meet the preset threshold conditions for backfill relevance as implicit medical knowledge tags.

[0087] Furthermore, the calculation process of this application embodiment is illustrated using a urinary system ultrasound teaching video as an example: Assume the preceding medical knowledge tag cache queue obtained by the annotation device includes the following candidate medical knowledge tags: The first candidate tag is "hydronephrosis," which appears between 48 and 52 seconds in the video and originates from the modalities of audio and subtitles. The second candidate tag was "urinary tract obstruction," which appeared between 42 and 45 seconds into the video and originated from the audio modality. The third candidate tag is "ureteral stone", which appears between the 30th and 35th seconds of the video and is sourced from the audio modality.

[0088] The current base time slice is from 56 seconds to 61 seconds of the video, and the current audio text is "A clear expansion can be seen here"; the audio text contains the referential expression "here" and does not contain a complete explicit medical knowledge label; when the teacher circles the anechoic expansion area on the right side of the ultrasound image with the mouse in the current screen, the annotation device identifies the circled area as the screen indication area; and after identifying the screen indication area and its adjacent areas, the regional medical clue "renal pelvis dilation" is obtained.

[0089] The annotation device further calculates the relevance of each candidate medical knowledge tag; let a, b, c, d, and e be 0.25, 0.20, 0.20, 0.20, and 0.15 respectively; for "hydronephrosis", the regional text matching degree M is 0.85, the courseware topic relevance T is 0.90, the preceding occurrence intensity F is 0.80, the time decay value D is 0.92, and the referential explanation text relevance S is 0.75, then the relevance of the relevance R is 0.8465; For "urinary tract obstruction", the regional text matching degree M is 0.45, the courseware topic relevance T is 0.60, the preceding occurrence intensity F is 0.55, the time decay value D is 0.70, the referential explanation text relevance S is 0.50, and the backfill relevance R is 0.5575. For "ureteral stones", the regional text matching degree M is 0.25, the courseware topic relevance T is 0.40, the preceding occurrence intensity F is 0.40, the time decay value D is 0.35, the referential explanation text relevance S is 0.30, and the backfill relevance R is 0.3375. If the first threshold is 0.70, the annotation device determines "hydronephrosis" as the implicit medical knowledge label for the current time slice to be backfilled. The annotation device can output the following time slice annotation results: the start and end times of the time slice are from 56 seconds to 61 seconds, the explicit medical knowledge label is empty, the implicit medical knowledge label is "hydronephrosis", the preceding source time slice is from 48 seconds to 52 seconds, the image indication area is the anechoic expansion area on the right side of the ultrasound image, and the backfill confidence is high.

[0090] Using the above method, when students search for "hydronephrosis," they can not only jump to the segment where the teacher first mentioned "hydronephrosis," but also to the segment where the teacher circles and explains the key scene of the dilated area.

[0091] In some embodiments, when the backfilling relevance of a candidate medical knowledge tag is greater than or equal to a first threshold, the annotation device can determine the candidate medical knowledge tag as the implicit medical knowledge tag for the time slice to be backfilled. For example, if the backfilling relevance of the candidate medical knowledge tag "ST segment elevation" is significantly higher than that of other candidate tags and is greater than the first threshold, the annotation device can use "ST segment elevation" as the implicit medical knowledge tag for the current time slice.

[0092] In other embodiments, when the backfill correlation of multiple candidate medical knowledge tags is greater than or equal to the first threshold, and the difference in backfill correlation between multiple candidate medical knowledge tags is less than the second threshold, it indicates that there is a strong correlation between the current screen indication area and multiple candidate tags. The annotation device can determine multiple candidate medical knowledge tags as candidate backfill tags and generate a review mark for the time slice to be backfilled.

[0093] For example, the preceding medical knowledge tag cache queue includes "renal pelvis dilation" and "hydronephrosis". The current screen indicates the dilation area, and the correlation between the two is high and the difference is small. At this time, the annotation device can use both "renal pelvis dilation" and "hydronephrosis" as candidate tags for backfilling and generate a mark to be reviewed so that the main tag or the merged tag can be manually confirmed later.

[0094] Figure 7 This is a schematic diagram of the output of time-slice annotation results provided in an embodiment of this application; in some embodiments, the annotation device can output the time-slice annotation results of medical teaching videos; for time slices to be backfilled, the time-slice annotation results may include the start and end times of the time slices to be backfilled, explicit medical knowledge tags, implicit medical knowledge tags, the preceding source time slices corresponding to the implicit medical knowledge tags, the location of the screen indicator area, and the backfill confidence level.

[0095] Among them, explicit medical knowledge tags are the tags corresponding to medical terms directly identified in the current time slice; implicit medical knowledge tags are tags determined by associating the preceding medical knowledge tag cache queue with the screen indication area; the preceding source time slice is used to indicate the explicit appearance position of the implicit medical knowledge tag before the current time slice; the screen indication area position is used to indicate the screen area corresponding to the teacher's instruction action; the backfill confidence can be determined based on the backfill correlation degree corresponding to the implicit medical knowledge tag.

[0096] In some embodiments, the backfill confidence level can be divided into three levels: high, medium, and low. For example, when the backfill relevance is greater than or equal to 0.85, the backfill confidence level is high; when the backfill relevance is greater than or equal to 0.70 and less than 0.85, the backfill confidence level is medium; when the backfill relevance is less than 0.70, implicit medical knowledge tags are not automatically determined, or they are output as low-confidence candidate results. The first threshold is used to determine whether automatic backfilling is necessary, and the confidence level threshold is used to determine the output level.

[0097] As an example, in an ECG instructional video, the teacher could first explain that "ST segment elevation is an important ECG manifestation of acute myocardial infarction," and then circle a certain waveform area on the ECG in the next time slice with a marker and say, "This position has exceeded the baseline."

[0098] The phrase "ST segment elevation" did not reappear in the current base time-slice audio text, but it did contain the referential expression "this location." The annotation device can recognize the trajectory of the annotation pen and determine that the circled area in the ECG waveform is the area indicated on the screen. The annotation device can extract candidate medical knowledge tags such as "ST segment elevation," "acute myocardial infarction," and "ECG" from the cache queue of preceding medical knowledge tags, and calculate the backfilling relevance based on regional medical clues, courseware titles, and the time of preceding occurrence.

[0099] When the backfill correlation of "ST segment elevation" meets the threshold condition, the annotation device identifies "ST segment elevation" as the implicit medical knowledge label of the current time slice and binds it to the current time slice; thus, when students search for "ST segment elevation", they can locate the interpretation segment where the teacher draws the waveform and explains that it "exceeds the baseline".

[0100] As an alternative example, in an anatomical teaching video, the teacher can first explain that "the portal vein is located in the porta hepatis region and is an important channel for blood flow into the liver," and then in the next segment, point to a tubular structure in the anatomical diagram and say, "This is an important channel for blood to enter the liver."

[0101] The term "portal vein" does not appear directly in the current base time-slice audio text, but it contains the referential expression "this one." The annotation device can determine the indicated area on the screen based on the mouse hover position or the arrow pointing position, and extract the adjacent annotation text or courseware title in that area. If the refill relevance of "portal vein" in the previous medical knowledge tag cache queue meets the threshold condition, the annotation device can identify "portal vein" as the implicit medical knowledge tag for the current time slice.

[0102] As an alternative example, in a demonstration video, the teacher can first explain that "the puncture point should be selected in the area outside the line connecting the anterior superior iliac spine and the umbilicus," and then highlight the operation location in the next time frame, saying, "This is the recommended needle insertion position."

[0103] The current base time-slice audio text does not contain complete medical knowledge tags that repeatedly appear for "puncture point" or "needle insertion location," but it does contain the referential expression "here," and there are highlighted areas in the image. The annotation device can determine the indicated area in the image based on the highlighted area, and then match the candidate tags such as "puncture point" and "needle insertion location" in the previous medical knowledge tag cache queue with the medical clues in this area to determine the implicit medical knowledge tags for the current time slice.

[0104] As can be seen from this embodiment, the embodiments of this application are not only applicable to medical imaging teaching videos, but also to anatomy teaching videos and operation demonstrations and other related teaching videos.

[0105] As an example, in a pathology slide teaching video, the teacher can first explain that "glandular structural dysplasia is characterized by disordered glandular arrangement and abnormal cell morphology." Then, in the next time slice, the teacher can circle a certain area in the pathology slide with a marker and say, "Here you can see that the arrangement is obviously disordered." Since the phrase "glandular structural dysplasia" does not reappear in the audio text of the current base time slice, but it contains the referential expression "here," and there is a circled area in the image, the annotation device can identify the circled area as the image indication area, extract the image annotation text or courseware title in the adjacent area, and match it with candidate medical knowledge tags such as "glandular structural dysplasia" and "disordered glandular arrangement" in the previous medical knowledge tag cache queue.

[0106] When the backfill correlation of "glandular structural dysplasia" meets the threshold condition, the annotation device can identify it as the implicit medical knowledge label of the current baseline time slice.

[0107] This application also provides a labeling device. The labeling device may include at least one processor and at least one memory, wherein the at least one memory stores computer program instructions, and when the labeling device is running, the at least one processor executes the method described in any of the above embodiments.

[0108] In some embodiments, the annotation device may further include a communication interface for acquiring medical teaching videos or sending time-slice annotation results to a teaching resource management platform, a course retrieval platform, or a manual review platform.

[0109] This application also provides a computer-readable storage medium storing a computer program that, when executed by a computer, causes the computer to perform the methods described in any of the above embodiments.

[0110] This application also provides a computer program product containing instructions that, when run on a computer, cause the computer to perform the methods described in any of the above embodiments.

Claims

1. A method for intelligent time-slice annotation of knowledge tags for medical teaching videos, characterized in that, include: Obtain video analysis information corresponding to medical teaching videos, including audio and text information, screen content information, and teacher instruction and action information; Based on the voice and text information, the current basic time slice is determined to be a time slice to be filled back. The time slice to be filled back is a time slice that contains referential explanations but has not been identified as a complete explicit medical knowledge tag. Based on the teacher's instruction action information, the screen instruction area is determined in the screen of the time slice to be filled, and regional medical clues are extracted based on the screen instruction area; The regional medical clues are associated and matched with the preceding medical knowledge tags before the time slice to be backfilled to obtain the backfilling association results of the candidate medical knowledge tags, and the implicit medical knowledge tags corresponding to the time slice to be backfilled are determined based on the backfilling association results. The implicit medical knowledge tags are bound to the time slices to be filled, and the time slice annotation results of the medical teaching video are generated. The step of determining the current base time slice as the time slice to be filled based on the voice and text information includes: Referential expression recognition is performed on the speech and text information of the current basic time slice; If a referential expression is identified in the voice text information, and no complete explicit medical knowledge tag that can be directly mapped to a preset medical knowledge node is identified in the current basic time slice, it is determined whether there is teacher instruction action information in the current basic time slice; When the current basic time slice contains teacher instruction action information and there is a preceding medical knowledge tag before the current basic time slice, the current basic time slice is determined as the time slice to be filled back. The referential expressions include pronouns that indicate screen areas, image objects, medical signs, or operation locations; The preceding medical knowledge tags are determined in the following way: Medical entity identification is performed on at least one preceding time slice located before the time slice to be backfilled to obtain explicit medical entities in the at least one preceding time slice; The explicit medical entities are standardized to obtain preceding medical knowledge tags; Establish a cache queue of preceding medical knowledge tags according to their occurrence time, source modality, and standardized name; The step of establishing a cache queue of preceding medical knowledge tags according to their occurrence time, source modality, and standardized name includes: Within a preset backtracking time range, the preceding medical knowledge tags are added to the preceding medical knowledge tag cache queue; Based on the time interval between the preceding medical knowledge tag and the time slice to be filled, a time decay value is set for the preceding medical knowledge tag; The larger the time interval, the lower the retention weight corresponding to the time decay value; The step of associating and matching the regional medical clues with the preceding medical knowledge tags before the time slice to be backfilled, obtaining the backfilling association results of candidate medical knowledge tags, and determining the implicit medical knowledge tags corresponding to the time slice to be backfilled based on the backfilling association results includes: The regional medical clues are matched with the candidate medical knowledge tags in the preceding medical knowledge tag cache queue; Based on the regional text matching degree, courseware theme relevance, preceding occurrence intensity, time decay value, and referential explanation text relevance, the backfilling relevance of each candidate medical knowledge tag to the time slice to be backfilled is calculated. Candidate medical knowledge tags whose backfill correlation meets the preset threshold condition are identified as implicit medical knowledge tags.

2. The intelligent time-slice annotation method for knowledge tags in medical teaching videos according to claim 1, characterized in that, The acquisition of video parsing information corresponding to the medical teaching video includes: The medical teaching video was divided into time slices to obtain multiple basic time slices; For the multiple basic time slices, extract speech-to-text, subtitle text, courseware page images, on-screen text information and teacher instruction action information respectively. The teacher instruction action information includes at least one of the following: teacher instruction action trajectory, local magnified area information and on-screen highlight information. The speech-to-text and the subtitle text are used as the speech-text information, and the courseware page image and the on-screen text information are used as the on-screen content information.

3. The method according to claim 1, characterized in that, The step of determining the screen instruction area in the screen of the time slice to be filled based on the teacher's instruction action information includes: The teacher's instruction action information is detected to determine the type of teacher's instruction action in the time slice to be filled; The indicated area of ​​the screen is determined in the screen of the time slice to be filled based on the type of teacher instruction action; The teacher's instruction action types include at least one of the following: mouse hover, mouse circle, pen stroke annotation, local zoom, and screen highlighting.

4. The method according to claim 3, characterized in that, The step of extracting regional medical clues based on the indicated area on the screen includes: The indicated area on the screen is processed by at least one of local text recognition, image annotation recognition, and extraction of information from adjacent courseware to obtain medical clues for the area. The regional medical clues include at least one of the following: medical text within the indicated area of ​​the screen, medical text in the adjacent area of ​​the indicated area of ​​the screen, image annotation text, arrow pointing text, legend text, courseware title, and chapter prompt information.

5. The method according to claim 1, characterized in that, The step of determining the candidate medical knowledge tags whose backfill relevance meets the preset threshold condition as the implicit medical knowledge tags includes: When there is a candidate medical knowledge tag whose backfill relevance is greater than or equal to the first threshold, the candidate medical knowledge tag is determined as the implicit medical knowledge tag. When the re-filling correlation of multiple candidate medical knowledge tags is greater than or equal to the first threshold, and the difference in re-filling correlation between the multiple candidate medical knowledge tags is less than the second threshold, the multiple candidate medical knowledge tags are determined as candidate re-filling tags, and a review mark is generated for the time slice to be re-filled.

6. The method according to any one of claims 1 to 5, characterized in that, The time-slice annotation results for generating the medical teaching video include: Output the start and end times of the time slice to be backfilled, explicit medical knowledge tags, implicit medical knowledge tags, the preceding source time slice corresponding to the implicit medical knowledge tags, the location of the screen indicator area, and the backfill confidence level; The confidence level of the backfill is determined based on the backfill correlation degree corresponding to the implicit medical knowledge tag.

Citation Information

Patent Citations

  • Method for carrying out anaphora resolution on teaching video on the basis of unsupervised mode

    CN106997346A

  • Practical training teaching video intelligent analysis and knowledge point automatic marking method and system based on multi-modal fusion

    CN121743805A