Video abstract generation method and device, equipment and storage medium

Through scene segmentation and multimodal information extraction methods, the generated video summary solves the problems of narrative inconsistency and factual errors in the medical and health insurance scenario, and realizes high-quality video summary generation to meet the needs of rapid claims review.

CN120236231APending Publication Date: 2025-07-01PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510354410.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The existing video summary generation technology has problems such as incoherence of narratives, insufficient emotional information capture and factual errors in the medical and health insurance scenarios, which is difficult to meet the needs of insurance practitioners and medical personnel for quick browsing and obtaining key information.

Method used

The video is divided into sub-target videos through scene segmentation, and multimodal information, including visual, audio and text information, is extracted, and a preset model is input to generate a video summary. This method clearly demonstrates the logical relationship between scenes and integrates multimodal information to improve the accuracy and logic of the summary.

Benefits of technology

The generated video summary is more organized, covering comprehensive content, improving the quality of the summary, helping users fully understand the video content, making it easier to quickly grasp key information and make accurate judgments, and meeting the needs of rapid claims review.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236231A_ABST
    Figure CN120236231A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence and the field of medical health, and discloses a video abstract generation method and device, equipment and a storage medium, and the method comprises the steps: obtaining a target video; analyzing different scenes in the target video; according to different scenes of the target video, the target video is segmented into sub-target videos, and the sub-target videos are in one-to-one correspondence with the different scenes in the target video; multi-mode information of the sub target videos is extracted; and inputting the multi-modal information of each sub-target video into a preset model to obtain a video abstract corresponding to the target video. The invention provides a video abstract generation method and device, equipment and a storage medium, and solves the problems existing in a video abstract generation method in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of artificial intelligence and healthcare, and in particular, to a method, apparatus, device, and storage medium for generating video summaries. Background Art

[0002] Video summary generation technology compresses video content into a concise summary. In the field of healthcare insurance, video summary technology can detect key frames, scenes, and important moments in medical training videos and claims-related videos with the help of various algorithms and artificial intelligence, thereby creating a concise summary. For example, during claims review, video summaries can assist insurance reviewers in quickly understanding the core situation of claims events. However, in this specific scenario, the technology still faces many challenges.

[0003] Maintaining narrative fluency is crucial. In claims review videos, if the narrative is not fluent, reviewers may have difficulty accurately grasping the overall picture of the event, thus affecting claims decisions. Capturing emotional tones is also essential. For example, in a video of a patient describing their condition, the patient's emotional changes may imply the severity or particularity of the condition. If the video summary fails to capture this emotional information, it may lead to an incomplete insurance claims assessment. Ensuring factual accuracy is of utmost importance. Whether it is a record of a medical procedure or a description of a claims event, any errors may cause serious consequences, such as an unreasonable claims determination.

[0004] Although the technology has made some progress in automatically generating video summaries, in the healthcare insurance scenario, there is still a need to significantly improve accuracy and intelligence to meet the needs of insurance practitioners and medical staff for quickly browsing content and obtaining key information.

[0005] Currently, in the healthcare insurance scenario, the main methods of video summary technology are as follows:

[0006] Based on key frame extraction: In medical training videos, video summaries are generated by identifying key frames, aiming to capture the main content of key operations and important turning points. As discrete pictures, key frames are difficult to present continuous dynamic information of surgical operations, disrupting the narrative fluency of the surgical process. When processing a large number of claims-related videos, as the video duration increases, the number of key frames also increases. This not only makes the generated summary too long for reviewers to quickly browse, but also makes it difficult to analyze the correlation between key frames, easily resulting in a situation where key frames are piled up without logical connection, affecting the accurate judgment of claims events.

[0007] Based on content summary: Text generation technology is used to convert medical-related video content into text descriptions, and then generate video summaries. This method combines natural language processing and computer vision technology. However, in medical scenarios, complex medical equipment, blurred lesion images, etc. may lead to understanding deviations when computer vision and natural language processing technology are combined. For example, when interpreting medical imaging videos, normal tissues may be misjudged as lesions. Video summaries generated based on incorrect text descriptions will inevitably contain factual errors, affecting medical decisions and insurance claims judgments. In addition, it is difficult for text to fully express the non-verbal information in the video, such as the emotions conveyed by the patient's tone and expression when describing the condition. These emotional information may be of great significance in assessing the severity of the disease and insurance claims, and text-based summaries cannot accurately capture these emotional tones. In addition, the text generation process is computationally intensive. When a large number of medical claims videos need to be processed quickly, the processing speed is slow, and it cannot meet the needs of real-time summary generation for rapid claims review.

[0008] Summary generation based on visual features: Computer vision technology is used to extract visual features in medical videos, such as the color of medical equipment and the movement of surgical instruments, to generate video summaries. However, for videos that are mainly based on doctor-patient dialogues and where the diagnosis of the disease and the basis for claims rely on language information, it is difficult to cover the core medical information and key claims content by relying solely on visual features such as color and movement, causing the summary to deviate from the main purpose. Moreover, changes in light and occlusion of surgical instruments in medical environments are relatively common, which will lead to inaccurate visual feature extraction and affect the quality of video summaries. Furthermore, this method focuses on the extraction of underlying visual features and lacks an understanding of the semantic relationship of medical content. It is unable to grasp the communication logic between doctors and patients and the causal relationship of the development of the disease, which makes the generated summary lack internal logic, which is not conducive to the effective implementation of medical training and insurance claims review. Summary of the invention

[0009] The present invention provides a method, device, equipment and storage medium for generating a video summary, so as to solve the problems existing in the video summary generating method in the prior art.

[0010] In a first aspect, the present invention provides a method for generating a video summary, comprising:

[0011] Get the target video;

[0012] Analyzing different scenes in the target video;

[0013] According to different scenes of the target video, the target video is divided into sub-target videos, and the sub-target videos correspond to different scenes in the target video one by one;

[0014] Extracting multimodal information of the sub-target video;

[0015] Input the multimodal information of each of the sub-target videos into a preset model to obtain a video abstract corresponding to the target video.

[0016] In a second aspect, the present invention provides a device for generating a video abstract, including:

[0017] An acquisition module, configured to acquire a target video;

[0018] An analysis module, configured to analyze different scenes in the target video;

[0019] A segmentation module, configured to segment the target video into sub-target videos according to different scenes in the target video, where the sub-target videos correspond one-to-one to different scenes in the target video;

[0020] An extraction module, configured to extract multimodal information of the sub-target videos;

[0021] A prediction module, configured to input the multimodal information of each of the sub-target videos into a preset model to obtain a video abstract corresponding to the target video.

[0022] In a third aspect, the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, where the processor implements the steps of the above-mentioned method for generating a video abstract when executing the computer program.

[0023] In a fourth aspect, the present invention provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and the computer program implements the steps of the above-mentioned method for generating a video abstract when executed by a processor.

[0024] The solution implemented by the above-mentioned method for generating a video abstract divides the video into sub-target videos according to different scenes through scene segmentation, clearly showing the sequence and logical relationship between each scene, and effectively solving the problem of discontinuous narration. For example, in a medical training surgery video, instead of isolated key frames, each surgical stage scene is presented completely, making the operation process coherent. In the solution, through scene segmentation and multimodal information extraction, key events are identified at the scene level, and combining multimodal information can understand key events from different abstraction levels. For example, in a surgical scene, key surgical steps can be understood from multiple modalities such as vision and language. The abstract generated by this solution according to scenes and multimodal information is more organized, covers comprehensive content, improves the quality of the abstract, helps users comprehensively understand the video content, and facilitates quickly grasping key information to make accurate judgments.

[0025] The solution uses multi-modal information fusion to avoid the deviation in understanding of a single modality. Different modality information complements each other, enabling a better understanding of the content structure relationship. For example, when interpreting a medical imaging video, combining visual and other modality information can prevent misjudging normal tissues as lesions, making the abstract more logical. The multi-modal information extraction of the solution includes non-verbal information, which can identify roles and themes at different abstraction levels. For example, when a patient describes their condition, the patient's emotions and the theme of the condition can be understood from multi-modal information such as tone of voice and facial expressions. Through multi-modal information fusion and a preset model, the solution reduces factual errors, captures the emotional tone, optimizes the calculation process to improve the processing speed, enhances the quality of the abstract, enables users to comprehensively and accurately understand the video content, and meets the needs of rapid claims settlement review.

[0026] The solution extracts multi-modal information, not limited to visual features, which can better include core information and reduce the influence of factors such as light. By fusing multi-modal information and using a preset model to understand semantic relationships, the content structure is sorted out. For example, in a doctor-patient dialogue video, combining multi-modal information such as language can cover core medical information. Through multi-modal information and a preset model, the solution can understand the logic of doctor-patient communication and the causal relationship of the development of the condition at different abstraction levels. For example, the diagnostic logic of the condition can be understood from multi-modal information such as the content of the dialogue and the doctor's examination actions. The abstract generated by the solution covers the core content, has an internal logic, improves the quality of the abstract, helps users comprehensively understand the video content, and is conducive to the development of medical training and insurance claims settlement review work.

[0027] In summary, the solution implemented by the above method for generating a video abstract can solve the problems existing in the prior art methods for generating video abstracts. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention. Obviously, the following described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0029] Figure 1 is a flowchart of a method for generating a video abstract according to an embodiment of the present invention;

[0030] Figure 2 is Figure 1 a flowchart of step S120 in

[0031] Figure 3 is Figure 2 a flowchart of step S121 in

[0032] Figure 4 is Figure 1 a flowchart of step S140 in

[0033] Figure 5 is Figure 1 A schematic flowchart of step S150 in

[0034] Figure 6 is Figure 5 A schematic flowchart of step S152 in

[0035] Figure 7 A schematic structural diagram of a video summary generation device in an embodiment of the present invention;

[0036] Figure 8 A schematic structural diagram of a computer device in an embodiment of the present invention;

[0037] Figure 9 Another schematic structural diagram of a computer device in an embodiment of the present invention. Detailed implementation manners

[0038] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0039] Please refer to Figure 1 As shown, the embodiments of the present invention provide a flowchart of a method for generating a video summary, including the following steps.

[0040] Step S110, obtain a target video.

[0041] Specifically, the target video can be any video. For example, when a patient applies for medical insurance claims, relevant videos may need to be submitted as supporting materials. For instance, a video of a patient who was accidentally injured during rehabilitation training in a rehabilitation center, which shows the patient's rehabilitation training process, the rehabilitation equipment used, and the guidance of the rehabilitation therapist, etc. The insurance company obtains this video to verify the authenticity and rationality of the patient's rehabilitation situation and the claims application.

[0042] Step S120, analyze different scenes in the target video.

[0043] It should be noted that the manner of analyzing different scenes in the target video can be any achievable manner.

[0044] For example, audio information can be used to assist in scene analysis. For a patient's claim settlement declaration video, if there are groans of pain from the patient in the audio, combined with the patient's expression and state in the video footage, it can be identified as a scene where the patient is injured and uncomfortable. For videos containing a large amount of language information, semantic analysis can also be used to determine the specific scene. For example, in the process of medical insurance review, accurately identifying these scenes helps the reviewers comprehensively understand the medical service process and provides a basis for reasonably evaluating medical insurance costs.

[0045] In some embodiments of the present invention, as Figure 2 shown, step S120 includes the following steps:

[0046] Step S121, compare each video frame of the target video to obtain video frame landmark points that produce visual differences;

[0047] Step S122, determine the scenes corresponding to the video frames before the video frame landmark points and the scenes corresponding to the video frames after the video frame landmark points as different scenes in the target video.

[0048] Specifically, in step S121, the method of comparing the video frames of the target video to obtain video frame landmark points that produce visual differences can be any achievable method. For example, image feature extraction algorithms and matching algorithms can be used to complete the comparison of video frames. Image feature extraction algorithms can include Scale-Invariant Feature Transform (SIFT) feature extraction algorithms and Speeded-Up Robust Features (SURF) extraction methods. Taking the Scale-Invariant Feature Transform (SIFT) feature extraction algorithm as an example, it first constructs a scale space for the video frame and detects key points through a Difference of Gaussian (DoG) pyramid. Then, it calculates the gradient direction histogram of the neighborhood of each key point to generate a feature descriptor. When comparing video frames, the Euclidean distance between the key point feature descriptors in different video frames is calculated to find matching key point pairs. When the number of matching key point pairs is lower than a certain threshold, it can be considered that there is a large visual difference between these two frames, and this frame is the video frame landmark point that produces visual differences.

[0049] After determining the video frame landmark points, the video is divided into the front and the back based on this landmark point. During the division process, for a series of consecutive video frames before the video frame landmark point, the visual elements, scene layouts, etc. they present are similar, so they can be grouped into one scene. And for the video frames after the video frame landmark point, since the visual features have changed significantly, presenting different scene elements and layouts from before, they are divided into another scene. This division method can intuitively distinguish the scenes in different stages of the video, facilitating subsequent separate analysis and processing of each scene.

[0050] It can be understood that determining scene transition points by precisely comparing the visual features of video frames can capture the actual changes in the scenes in the video more accurately than some simple scene division methods based on time or preset rules. In the medical insurance health scenario, this is particularly important because accurately dividing the different stage scenes during surgery or the different link scenes of patients' rehabilitation training helps medical insurance reviewers and medical professionals better understand the specific content of medical services, providing a reliable basis for medical insurance cost assessment, medical quality monitoring, etc. This scene analysis method based on visual differences has good adaptability to various types of medical insurance health-related videos. Whether it is a surgical video, a rehabilitation training video, or a health education video, etc., as long as there are scene changes in the video, the corresponding video frame landmark points can be found through image feature comparison, and then the scenes can be accurately divided. Different from some scene analysis methods in specific fields, it is only applicable to videos of specific formats or content types. After clearly dividing different scenes, it is more convenient for subsequent operations such as extracting multi-modal information of sub-target videos and generating video summaries. For example, when generating a video summary, key information can be extracted separately for specific scenes, etc., to generate a more targeted and logical summary, improving the quality and practicality of the video summary, and better meeting the needs of all parties in the medical insurance health scenario for quick access and understanding of video information.

[0051] In some embodiments of the present invention, as Figure 3 shown, step S121 includes:

[0052] Step S1211, comparing the pixel differences of the pixels at the same position in any two adjacent video frames in the target video. When the pixel difference is greater than a preset threshold, taking the middle position of the two adjacent video frames as the video frame landmark point where visual differences occur.

[0053] Specifically, in the target video, each frame of the image is composed of numerous pixels. For any two adjacent video frames, in order to determine whether there is a visual difference between them, we can analyze the pixels at the same position one by one. For example, we can first obtain the color values of the pixels at the same position in two adjacent frames, such as red, green, blue, hue, saturation, and lightness. Taking the RGB color space as an example, each pixel has corresponding red, green, and blue component values. Calculate the absolute values of the differences of these two pixels in the three components of red, green, and blue, and then add up these three absolute differences of the differences to obtain the total pixel difference. For example, using R, G, and B to represent red, green, and blue respectively, for the pixels A(R1, G1, B1) and pixel B(R2, G2, B2) at the same position in two adjacent frames, the pixel difference = |R1 - R2| + |G1 - G2| + |B1 - B2|. When the calculated pixel difference is greater than the pre-set threshold, it means that the visual features at this position of these two adjacent video frames have changed significantly. At this time, the middle position between these two adjacent video frames is determined as the video frame flag point indicating the visual difference. The way to determine the middle position can be simply taking the average of the two frame numbers (assuming the video frames are numbered in sequence). For example, for the adjacent nth frame and the (n + 1)th frame, the frame corresponding to the middle position can be considered as the (n + 0.5)th frame (in actual applications, rounding and other processing can be performed according to specific situations).

[0054] As a specific example, in the medical insurance health scenario, there are often videos of the insured's life scenarios submitted to assist in proving their health status or the reasonableness of the claim application. For example, an insured person applied for a claim due to accidental injury and submitted a video of their daily activities at home for a period of time before the injury. In the video, at first, the insured person was moving normally in the living room, the picture was stable, and the pixel difference of the pixels at the same position between adjacent video frames was small. However, when the insured person got up and walked towards the kitchen, due to the differences in light conditions, background layouts, etc. between the living room and the kitchen, the pixel difference of the pixels at the same position in adjacent video frames increased. Assuming the pre-set threshold is 50 (this threshold can be determined through experiments or experience according to the actual situation), when the calculated pixel difference reaches 60, it exceeds the pre-set threshold. At this time, the middle position between these two adjacent video frames is determined as the video frame flag point. This flag point represents the key node of the transition from the living room scene to the kitchen scene. Another example is that when evaluating the health management situation of the insured, there will be videos of their fitness activities. In the video, from the insured person running on the treadmill to switching to doing stretching exercises on the yoga mat, during this process, due to the differences in exercise scenes, body postures, and surrounding fitness equipment, the pixel difference of the pixels at the same position between adjacent video frames will exceed the pre-set threshold, thus determining the corresponding video frame flag point to identify the transition from the running scene to the stretching scene.

[0055] It is understandable that directly comparing the pixel differences of pixels at the same position in adjacent video frames has a simple and easy-to-understand calculation method, without the need for complex algorithm models and a large amount of computing resources. In the insurance assessment scenario, it is often necessary to quickly process a large number of video materials submitted by insured persons. This simple and efficient method can determine the video frame landmark points in a relatively short time, improving the work efficiency of insurance assessment. Moreover, whether it is the change of indoor and outdoor environments, or the changes in the activity status of people, background layout, etc., as long as such changes can cause obvious differences in the pixel values of adjacent video frames, they can be accurately detected by this method. In the insurance assessment scenario, the video content submitted by insured persons is rich and diverse, covering various life scenarios. This method can well adapt to these different types of scenario changes and accurately find the landmark points of scene conversion. Furthermore, by determining the video frame landmark points, different life scene segments can be clearly divided. Insurance assessors can analyze the behavior performance, health status, etc. of insured persons in different scenarios more targeted based on these divisions, providing strong evidence for accurately assessing the rationality of claim applications, health management situations, etc., and enhancing the accuracy and reliability of insurance assessment.

[0056] Step S130, according to different scenarios of the target video, divide the target video into sub-target videos, and the sub-target videos correspond one by one to different scenarios in the target video.

[0057] Specifically, after determining the different scenarios in the target video, the target video can be segmented based on these scenarios to generate sub-target videos corresponding to each scenario. In specific implementation, the starting and ending video frames corresponding to each scenario can be first determined. This can be completed based on the previously determined video frame landmark points. For example, if the video frame landmark point for the transition from the living room scene to the kitchen scene has been determined, then the starting frame of the sub-target video corresponding to the living room scene is the beginning of the video, and the ending frame is the frame before this landmark point; the starting frame of the sub-target video corresponding to the kitchen scene is this landmark point frame, and the ending frame is determined according to the landmark point of the subsequent scene or the end of the video.

[0058] As a specific example, in the medical health insurance scenario, taking the case of a patient applying for a claim for hospitalization due to an accidental fracture, the patient submitted a video of the entire process from the injury scene to the hospital treatment. After the previous scene analysis, the video frame landmarks of different scenes such as the injury scene, ambulance transfer, hospital emergency room, operating room, and ward rehabilitation were determined. Then, based on these landmarks, the video corresponding to the injury scene is segmented from the beginning to the frame before the arrival of the ambulance as a sub-target video; the video corresponding to the ambulance transfer scene is segmented from the ambulance arrival frame to the frame before entering the hospital emergency room, generating another sub-target video, and so on, the operating room, ward rehabilitation and other scenes are segmented into independent sub-target videos. In this way, insurance auditors can conduct detailed analysis of the sub-target videos of each scene, such as checking the environment and cause of the accident in the injury scene sub-target video, and confirming whether the surgical process meets the specifications in the operating room sub-target video, etc., to provide a clearer and more accurate basis for claim review.

[0059] Step S140: extracting multimodal information of the sub-target video.

[0060] Specifically, the number of multimodal information of the sub-target video extracted in this step can be set according to the specific needs of the application. For example, the multimodal information may include visual modality, audio modality, and text modality. Among them, the visual modality can use computer vision technology to extract image features in the video, and also use edge detection algorithms to obtain the contour information of objects in the video, and use target detection algorithms to identify key objects in the video; it can also extract color features, texture features, etc. of the video, which can reflect the overall style and detail information of the video picture. The audio modality is to analyze the audio in the video and extract audio features. For example, the voice content in the audio can be converted into text through speech recognition technology to obtain the dialogue information between doctors and patients, the doctor's diagnosis instructions, etc.; the frequency features and volume features of the audio can also be extracted. The frequency change of the audio can reflect the pitch of the sound, and the volume can reflect the strength of the sound. These features can assist in judging the atmosphere of the scene or the emotional state of the characters in some cases. The text modality may include the text obtained by audio conversion and the text information with subtitles or annotations in the video itself. These text information can directly convey the key information in the video.

[0061] As a specific example, taking the above-mentioned patient fracture claim video as an example, in the injury scene sub-target video, the image features of the injured part, such as the shape of the wound, bleeding, etc., are extracted from the visual modality; the shouts of the surrounding people and the instructions of the emergency personnel are extracted from the audio modality, and the description of the accident by the on-site personnel is obtained through speech recognition; the text information on the possible on-site signboards is extracted from the text modality, such as the name of the place where the accident occurred. In the ward rehabilitation sub-target video, the visual modality extracts the action posture features of the patient's rehabilitation training, such as the extension angle and range of motion of the limbs; the audio modality extracts the conversation content between the rehabilitation therapist and the patient, and the plan and progress of the rehabilitation training are understood through speech recognition; the text modality extracts the text information such as the rehabilitation precautions posted on the wall of the ward. By extracting these multi-modal information, insurance auditors can have a more comprehensive understanding of the patient's treatment process and rehabilitation status, so as to more accurately evaluate the rationality of the claim application.

[0062] In some embodiments of the present invention, Figure 4 As shown, step S140 includes the following steps:

[0063] Step S141, extracting text information from the sub-target video;

[0064] Step S142, extracting key frame image information in the sub-target video;

[0065] Step S143: input the text information and the key frame image information into a preset feature encoder to obtain a target feature vector for characterizing comprehensive information of the text information and the key frame image information.

[0066] Specifically, for step S141, extracting text information from the sub-target video mainly involves two situations: the hard subtitle text of the video itself and the text obtained by audio-to-text conversion. When extracting hard subtitle text, if the video contains hard subtitles, the position of the subtitles in the video frame can be identified first. This can be achieved with the help of optical character recognition technology. For example, in some medical health insurance promotion videos, there may be subtitles about insurance terms, claims procedures, etc., and these hard subtitle texts can be accurately extracted through optical character recognition technology. When obtaining text through audio-to-text conversion, speech recognition technology can be used to convert the audio in the video into text. For example, a speech recognition tool can be used to sample and frame the audio, and then extract audio features, such as Mel frequency cepstral coefficients. Then, the extracted features are input into a pre-trained speech recognition model, and the model converts the audio signal into a corresponding text sequence through a comprehensive analysis of the acoustic model and the language model. In the medical health insurance scenario, the audio of the conversation between the doctor and the patient, the audio of the insurance customer service explanation, etc., can all be converted into text information through such speech recognition technology.

[0067] As a specific example, in the scenario of medical and health insurance, a video of a medical treatment process submitted by a patient contains a conversation between a doctor and the patient. In the video, the doctor explains the treatment plan to the patient in the office, and at the same time, a brief written description (hard subtitles) of the treatment plan is displayed on the computer screen. Through speech recognition technology, the audio of the conversation between the doctor and the patient is converted into text, from which key information such as the doctor's diagnosis of the condition and treatment suggestions can be obtained. At the same time, optical character recognition technology is used to extract the text content of the hard subtitles on the computer screen, such as the specific name of the treatment plan, the dosage of the medicine, and other information. These extracted text information is crucial for insurance reviewers to understand the patient's medical treatment situation and treatment details and to judge the reasonableness of the claim application.

[0068] Specifically, for step S142, when selecting key frames, any achievable method can be used. For example, key frames can be selected based on image features, such as calculating the image feature differences between video frames. Feature descriptors are extracted for each frame of the video, and then the similarity between adjacent frame feature descriptors is calculated. When the similarity is lower than a certain threshold, it indicates that this frame has a large difference from the adjacent frames and may be a key frame. Key frames can also be selected based on changes in video content. For example, key frames can be selected according to semantic changes in the video content. For example, in a medical and health insurance video, if the video content changes from the doctor explaining the condition to showing the treatment equipment, the frame corresponding to this content change point is very likely to be a key frame. The target detection algorithm can be used to identify key objects in the video. When a significant change in the object type is detected, this frame is used as a key frame candidate. For example, in a surgical video, when it is detected that the surgical instrument changes from a scalpel to a suture needle, the corresponding video frame can be used as a key frame.

[0069] Specifically, after determining the key frames, the image information of the key frames can be extracted, which can include, for example, the pixel values of the image, color histograms, edge features, etc. These image information can reflect important scenes, object appearances, etc. in the video and provide a basis for subsequent analysis.

[0070] As a specific example, in the scenario of medical and health insurance, the surgical process can be divided into multiple stages, such as preoperative preparation, surgical operation, and postoperative treatment. Through the method of selecting key frames based on image features and content changes, in the preoperative preparation stage, key frames of the nurse arranging surgical instruments and the patient being pushed into the operating room are selected. These key frame images contain information such as the types of surgical instruments and the patient's status. In the surgical operation stage, key frames of the doctor performing key surgical steps (such as organ resection, blood vessel suture) are selected. From these key frame images, information such as the usage method of surgical instruments and the details of the surgical site can be extracted. These key frame image information is of important reference value for insurance reviewers to evaluate the complexity of the surgery and the standardization of medical services.

[0071] Specifically, for step S143, the text information can be preprocessed. For example, the word embedding technique in natural language processing can be used to convert each word in the text into a vector representation of a fixed length. Additionally, techniques such as recurrent neural networks or their variants, long short-term memory networks, and gated recurrent units can be utilized to process the word vectors corresponding to the text information, capture the context information of the text, and generate a vector (i.e., the target feature vector) that can represent the entire text content.

[0072] Specifically, for the key frame image information, a convolutional neural network can be used for encoding. The convolutional neural network extracts features from the image through a series of convolutional layers, pooling layers, and fully connected layers. The convolutional kernels in the convolutional layers slide over the image to extract local features of the image, and the pooling layers downsample the feature maps to reduce the feature dimensions. After multiple layers of processing, the fully connected layer converts the extracted features into a vector of a fixed length.

[0073] More specifically, the vector obtained by encoding the text and the vector obtained by encoding the image can be fused. For example, the fusion method can be concatenation, weighted summation, etc. Then, the fused vector can be input into a fully connected layer for further processing to obtain the target feature vector that finally represents the comprehensive information of the text information and the key frame image information. This target feature vector contains various information such as the text semantics and image vision in the video, providing a more comprehensive and rich feature representation for subsequent video analysis tasks.

[0074] As a specific example, in the medical health insurance scenario, for a sub-target video containing a doctor's diagnosis description text and key frame images of a patient's treatment process. First, the doctor's diagnosis description text can be encoded to obtain a vector representing the text semantics, which contains information such as the doctor's analysis of the condition and treatment suggestions. Then, the key frame images of the patient's treatment process are input into the model to obtain a vector representing the visual features of the images, which contains information such as the patient's treatment environment and treatment equipment. Next, these two vectors are concatenated and then processed through a fully connected layer to obtain the target feature vector. This target feature vector combines the text and image information, and insurance assessors can more accurately evaluate the risk level of the claim application based on this vector and determine whether there are unreasonable medical expenses or fraud risks, etc.

[0075] It is understandable that by separately extracting text information and key-frame image information and encoding and fusing them into a target feature vector, the effective integration of video multi-modal information is achieved. In the medical health insurance scenario, both text information and image information contain important values, such as disease diagnosis and treatment plan description in the text, and medical operations and patient status in the image. Fusing these multi-modal information can provide a more comprehensive and rich video content representation, which helps insurance reviewers and evaluators understand the information conveyed by the video more accurately and make more reasonable decisions. Moreover, by separately encoding and processing the text information and key-frame image information, mature technologies and algorithms in the fields of natural language processing and computer vision are utilized, enabling efficient extraction and processing of information in different modalities. Compared with directly performing complex analysis on the entire video, this modular processing method reduces the computational complexity and improves the speed of information processing, especially suitable for processing a large amount of medical health insurance-related video data. By fusing multi-modal information into a unified target feature vector, it provides a more representative input for subsequent analysis models (such as video summary generation models, risk assessment models, etc.). This multi-modal feature representation can better capture various information and relationships in the video, enabling the model to have stronger generalization ability when facing different types of medical health insurance videos, being able to analyze and predict more accurately, and improving the accuracy and reliability of insurance business processing.

[0076] Step S150: Input the multi-modal information of each of the sub-target videos into a preset model to obtain the video summary corresponding to the target video.

[0077] Specifically, in this step, the preset model can be any model that can obtain the video summary corresponding to the target video when inputting the multi-modal information of the sub-target video. For example, the preset model can be a sequence-to-sequence model based on deep learning, a recurrent neural network with an attention mechanism, and a long short-term memory network, etc.

[0078] Specifically, before using the preset model, it can be trained with a large amount of data. The training data can include various medical health insurance-related videos and their corresponding manually annotated high-quality summaries. During the training process, the model can continuously adjust its own parameters to learn the mapping relationship between the input multi-modal information and the output summary. For example, taking videos of the medical treatment processes of a large number of different diseases and their corresponding accurate and detailed text summaries as training data and inputting them into the long short-term memory network model, the model will optimize its own weights based on these data to improve the accuracy of generating summaries.

[0079] Specifically, before inputting the multi-modal information into the preset model, the multi-modal information can be fused to make it in a unified format acceptable to the model. For example, the vectors encoded from the text information and the vectors encoded from the key-frame image information can be concatenated or weighted and summed to form a comprehensive feature vector.

[0080] More specifically, the processed sub-goal video multi-modal information is sequentially input into the preset model. Taking the long short-term memory network model as an example, the model can gradually generate each part of the video summary according to the input feature vector, following the patterns and rules learned during training. At each time step, the model calculates a new hidden state based on the current input feature vector and the hidden state of the previous time step, and generates the corresponding summary text segment through the output layer. To enable the model to pay more attention to the important parts of the input information and improve the accuracy and relevance of the generated summary, an attention mechanism can be introduced into the model. The attention mechanism can calculate the correlation weights of each part of the input information with the current part of the generated summary, enabling the model to focus on the multi-modal information most relevant to the current summary content when generating the summary. For example, when generating a video summary of a surgical procedure, the model can pay more attention to the surgical operation details in the key-frame images of the surgery and the key step descriptions in the corresponding text through the attention mechanism, thereby generating a more accurate summary.

[0081] As a specific example, in the medical health insurance scenario, assume that a patient is hospitalized for a heart attack and applies for insurance claims, submitting a video containing the entire process from the scene of the onset to the hospital treatment. After the previous steps, the video has been segmented into multiple sub-goal videos, such as the sub-goal video of the scene of the onset, the sub-goal video of the ambulance transfer, the sub-goal video of the hospital emergency room, the sub-goal video of the operating room, the sub-goal video of the ward rehabilitation, etc., and the multi-modal information of each sub-goal video has been extracted. During the training phase of the preset model, a large number of similar videos of the heart disease treatment process and their corresponding accurate summaries can be used for training. These training data cover information such as the onset situations, treatment methods, and rehabilitation processes of different patients, enabling the model to learn various patterns and key information in the heart disease treatment process. The processed multi-modal information of each sub-goal video is sequentially input into the trained preset model. After receiving the input data of the sub-goal video of the scene of the onset, the model uses the attention mechanism to focus on the description of the patient's symptoms in the text and the patient's state in the key-frame images, generating the beginning part of the summary. When the multi-modal information of the sub-goal video of the operating room is input, the model can focus on the key operation steps and related text descriptions during the surgery through the attention mechanism. After processing the multi-modal information of all sub-goal videos, the model can generate a series of summary text segments, which are then sorted and combined to finally obtain the complete video summary.

[0082] In some embodiments of the present invention, the preset model is the encoder part of the Transformer model.

[0083] Specifically, when the preset model is the encoder part of the Transformer model, for text information, the text can first be split into small units. Then, a unique vector is assigned to each unit. This vector not only contains the meaning of the unit itself but also adds position information. For key-frame images, the image can be divided into many small blocks, and each small block is similar to a unit in the text. Specific image processing methods are used to extract the features of each small block to obtain corresponding feature vectors. Similarly, these vectors also need to add position information to indicate the position of the small block in the image. The text and image information after the above embedding process will be input into the multi-head self-attention layer of the Transformer encoder. The multi-head self-attention mechanism allows the model to observe this information from multiple different perspectives. It will transform the input information through different "perspectives" to obtain things like queries, keys, and values. Then, it calculates the similarity between each "query" and all "keys", weights the "values" according to the similarity to obtain the attention weights for each position, and combines these weights with the corresponding values. This process is performed in parallel multiple times, each time calculating from a different "perspective", and finally, the results of multiple calculations are combined, so that the complex relationships between multi-modal information can be captured. Further, the output of the multi-head self-attention will enter the feed-forward neural network layer. This layer has two fully connected layers, with an activation function in the middle to increase the processing ability of the model. Its role is to further process the information output by the multi-head self-attention, enabling the model to better understand and process this information. In the multi-head self-attention layer and the feed-forward neural network layer, layer normalization and residual connection techniques are used. Layer normalization processes the feature dimensions of each sample to make the data more regular, making it easier for the model to learn. Residual connection is to directly add the input to the output of the layer, which can avoid the problem of gradient disappearance during the training process of the model, enabling the model to be trained deeper and learn more complex information. The final output of the encoder part is a feature representation that integrates multi-modal information. Based on this output, a linear layer is added to transform it into a vector of a fixed dimension. Then, through a decoder, this vector is converted into a video summary in natural language form. The decoder will generate the summary word by word according to this vector until the end flag is generated to obtain the complete video summary.

[0084] It is understandable that, different from traditional recurrent neural networks, the encoder part of the Transformer can process a lot of information simultaneously without processing one by one in sequence. When processing long videos containing a large amount of information in medical health insurance, it can greatly speed up the calculation speed and save the time for training and generating summaries. Moreover, the self-attention mechanism of the Transformer can well capture the relationships between different parts in a long sequence. In medical health insurance videos, from the causes of diseases to the treatment processes and then to the recovery situations, these information have a large span and a long time. The encoder part of the Transformer can better associate this information and generate video summaries more accurately and comprehensively than other models. The multi-head self-attention mechanism can process various types of information such as text and images simultaneously. Through different "perspectives", it can better fuse multi-modal information. For example, in a medical scenario, the condition described in the text and the symptoms, treatment processes, etc. shown in the key-frame images can be deeply fused through the self-attention mechanism, making the generated video summary more accurately reflect the content in the video. In addition, the structure of the Transformer is simple and easy to expand and modify. According to the special requirements of the medical health insurance scenario, the model can be adjusted, such as increasing the number of layers, changing the number of heads, etc., to make the model more suitable for this scenario and improve the quality of video summary generation.

[0085] In some embodiments of the present invention, as Figure 5 shown, step S150 includes the following steps:

[0086] Step S151, extracting keyword information from the text information;

[0087] Step S152, inputting the target feature vector of the sub-target video and the keyword information into a preset model to obtain the video summary corresponding to the target video.

[0088] Specifically, for step S151, the way to extract keyword information from text information can be any implementable way. For example, the term frequency-inverse document frequency algorithm can be used to calculate the frequency of each word in the text and its rarity in the entire document set. Words with high frequency and high rarity can be considered as keywords; the TextRank algorithm can also be used. The words in the text are regarded as nodes in a graph, and a graph is constructed according to the co-occurrence relationship between the words. Then, the importance scores of the nodes are calculated iteratively, and the words with high scores are determined as keywords.

[0089] Specifically, the target feature vector is obtained by integrating the sub-goal video text information and key-frame image information in the previous steps, and it contains the multi-modal features of the video. The keyword information highlights the core points of the text. After inputting these two into the preset model, the model will perform a fusion process on them. The internal structure of the model will learn the association between the target feature vector and the keywords, and explore the potential relationship between the multi-modal information and the core content of the text. Then, based on these learned relationships, the model generates a video summary that generalizes the content of the target video. This summary concisely presents the key information and important plots in the video.

[0090] As a specific example, in the field of medical insurance and health, assume that the text information of a sub-goal video is about a patient being hospitalized for pneumonia treatment. The text content is as follows: "The patient was admitted to the hospital due to coughing and fever for three days and was diagnosed with pneumonia. The doctor administered antibiotic treatment and performed a chest CT scan to determine the severity of the condition. After a week of treatment, the patient's symptoms improved significantly and was ready to be discharged." When extracting keywords, it will be found that words such as "pneumonia", "antibiotic treatment", "chest CT scan", "symptom improvement", and "discharge" appear relatively frequently in this text and are relatively unique in the text collection related to pneumonia treatment. Therefore, these words will be extracted as keywords. The target feature vector of this sub-goal video has been obtained previously, and this vector integrates information such as the patient's status in the key-frame images of the video, the use of medical equipment, and the description of the condition and treatment in the text. Now, input this target feature vector and the extracted keywords "pneumonia", "antibiotic treatment", "chest CT scan", "symptom improvement", and "discharge" into the encoder part of the preset Transformer model. The model will analyze the comprehensive information of the image and text contained in the target feature vector and combine the core points represented by the keywords. For example, based on the patient's status shown in the image in the target feature vector, the symptoms described in the text, and the keyword "symptom improvement", the model can judge the patient's recovery situation during the treatment process. Then, based on this information, the model generates a video summary such as "The patient was admitted to the hospital due to pneumonia, received antibiotic treatment and a chest CT scan, and was ready to be discharged after a week with symptom improvement."

[0091] It is understandable that extracting keyword information can focus on the core content of the text and avoid interference from irrelevant information. By inputting keywords and target feature vectors into the model together, the model can more accurately capture the key information and important plots in the video, thereby generating a more accurate video summary. In the healthcare insurance scenario, an accurate summary helps insurance auditors quickly understand the patient's condition, treatment process and results, and make more reasonable claims decisions. In addition, the target feature vector contains multimodal information of the video, while the keyword information highlights the core of the text. Combining the two and inputting them into the model can better achieve the interaction and fusion of multimodal information. The model can find relevant image and text information in the target feature vector based on the guidance of keywords, so that the generated summary can more comprehensively and deeply reflect the content of the video. In addition, the keyword information simplifies and refines the text, reducing the redundant information input to the model. The model can focus on key content more efficiently during processing, speed up the generation of summaries, and improve the efficiency of processing a large number of healthcare insurance-related videos.

[0092] In some embodiments of the present invention, Figure 6 As shown, step S152 includes the following steps.

[0093] Step S1521, inputting the target feature vector of the sub-target video and the keyword information into a preset model to obtain a predicted video summary of the target video;

[0094] Step S1522, inputting the predicted video summary and the actual video summary of the target video into a cross entropy loss function, and calculating a loss value of the predicted video summary of the target video relative to the actual video summary;

[0095] Step S1523: when the loss value is less than a preset threshold, the predicted video summary is determined as the video summary corresponding to the target video.

[0096] Specifically, for step S1521, after the target feature vector and keyword information are input into the preset model, the model will perform a series of processing on these inputs. The neural network layer inside the model will perform feature transformation and learning on the input, and mine the semantic and logical relationships in the input information. Through these learning processes, the model will generate a series of word units, which are finally combined into a predicted video summary. This predicted video summary attempts to summarize the key information and core content in the target video.

[0097] Regarding step S1522, the cross-entropy loss function is a commonly used function to measure the difference between the predicted result and the true result. The actual video summary is usually an accurate and comprehensive summary manually annotated or written by experts based on the content of the target video, which represents the true core information of the video. After inputting the predicted video summary and the actual video summary into the cross-entropy loss function, the function will compare them word by word. For each word in the predicted video summary, the difference between its predicted probability distribution and the true probability distribution of the corresponding word in the actual video summary is calculated. This difference is quantified as a loss value through the calculation method of cross-entropy. The larger the loss value, the greater the difference between the predicted video summary and the actual video summary, that is, the worse the prediction effect of the model; on the contrary, the smaller the loss value, the closer the predicted result is to the true situation.

[0098] Specifically, regarding step S1523, the preset threshold can be a pre-set critical value used to determine whether the quality of the predicted video summary meets the acceptable standard. When the calculated loss value is less than the preset threshold, it means that the difference between the predicted video summary and the actual video summary is small, and the prediction effect of the model is good. In this case, the predicted video summary can be determined as the final video summary corresponding to the target video. If the loss value is greater than or equal to the preset threshold, it indicates that the prediction result is inaccurate, and the model may need to be adjusted or retrained to improve the prediction accuracy.

[0099] As a specific example, assume that in the field of medical and health insurance, there is a target video about the treatment process of a patient's heart disease. The sub-goal videos contain information about various scenarios such as the patient's examinations, diagnoses, and treatments in the hospital. The keywords extracted from the relevant text information are "heart disease", "coronary angiography", "stent implantation", and "rehabilitation treatment". The target feature vector integrates the features of key frame images in the video (such as the pictures of examination equipment, surgical scenes, etc.) and the text information. The target feature vector and the keywords are input into a preset Transformer model. After processing, the model generates a predicted video summary: "The patient was admitted to the hospital due to heart disease, underwent coronary angiography, then received stent implantation, and underwent rehabilitation treatment after the operation." The actual video summary written by medical experts based on the content of the target video is: "The patient was diagnosed with heart disease. After being admitted to the hospital, coronary angiography was first performed. After clarifying the condition, stent implantation was carried out, and comprehensive rehabilitation treatment was carried out after the operation." The predicted video summary and the actual video summary are input into the cross-entropy loss function. The function will compare the vocabulary and semantics of the two. For example, there are certain differences in the expressions between "admitted to the hospital due to heart disease" in the predicted summary and "diagnosed with heart disease, after being admitted to the hospital" in the actual summary. The cross-entropy loss function will calculate the corresponding loss value according to this difference. Suppose that after calculation, the obtained loss value is 0.3. Suppose the preset threshold is 0.5. Since the calculated loss value of 0.3 is less than the preset threshold of 0.5, it indicates that the difference between the predicted video summary and the actual video summary is small, and the prediction effect is good. Therefore, the predicted video summary "The patient was admitted to the hospital due to heart disease, underwent coronary angiography, then received stent implantation, and underwent rehabilitation treatment after the operation" is determined as the video summary corresponding to the target video.

[0100] It can be understood that by using the cross-entropy loss function to measure the difference between the predicted video summary and the actual video summary, and judging the quality of the prediction result according to the loss value, it can be ensured that the finally determined video summary can reflect the core content of the target video as accurately as possible. In the field of medical health insurance, accurate video summaries are crucial for tasks such as insurance claim review and risk assessment. For example, an accurate summary can help insurance reviewers quickly understand the patient's condition and treatment process and determine whether the claim application is reasonable. When the loss value is greater than the preset threshold, it can be found that the prediction effect of the model is not good. At this time, the model can be adjusted and optimized. The performance of the model can be improved by increasing training data, adjusting model parameters, etc. This feedback mechanism enables the model to continuously learn and adapt to different medical health insurance video contents, improving the accuracy and stability of the prediction. The setting of the preset threshold provides an objective standard for judging the quality of the predicted video summary. Only when the prediction result meets a certain accuracy requirement is it determined as the final video summary. This ensures that the generated video summary has high reliability and reduces decision-making errors caused by inaccurate summaries, such as incorrect claim approvals.

[0101] In this way, the solution of the embodiment of the present invention divides the video into sub-target videos according to different scenes through scene segmentation, clearly showing the sequence and logical relationship between each scene, and effectively solving the problem of discontinuous narration. For example, in a medical training surgical video, instead of isolated key frames, each surgical stage scene is presented completely, making the operation process coherent. In the solution, key events are identified at the scene level through scene segmentation and multi-modal information extraction. Combining multi-modal information can understand key events from different abstraction levels. For example, in a surgical scene, key surgical steps can be understood from multiple modalities such as vision and language. The summary generated by this solution according to the scene and multi-modal information is more organized, covers comprehensive content, improves the quality of the summary, helps users comprehensively understand the video content, and facilitates quickly grasping key information to make accurate judgments.

[0102] The solution avoids the deviation of single-modal understanding through multi-modal information fusion. Different modal information complements each other, and can better sort out the content structure relationship. For example, when interpreting a medical imaging video, combining visual and other modal information can avoid misjudging normal tissues as lesions, making the summary more logical. The multi-modal information extraction of the solution includes non-verbal information, and can identify roles and themes at different abstraction levels. For example, when a patient describes their condition, the patient's emotions and the theme of the condition can be understood from multi-modal information such as tone and expression. The solution reduces factual errors, captures the emotional tone, optimizes the calculation process to improve the processing speed, and improves the quality of the summary through multi-modal information fusion and a preset model, enabling users to comprehensively and accurately understand the video content and meet the needs of rapid claim review.

[0103] The solution extracts multimodal information, not limited to visual features, can better include core information, reduce the influence of factors such as light, understand semantic relationships through multimodal information fusion and a preset model, and sort out the content structure. For example, in a doctor-patient dialogue video, combining multimodal information such as language can cover core medical information. Through multimodal information and a preset model, the solution can understand the logic of doctor-patient communication and the causal relationship of the development of the disease from different abstraction levels. For example, it can understand the disease diagnosis logic from multimodal information such as the dialogue content and the doctor's examination actions. The abstract generated by the solution covers the core content, has an internal logic, improves the quality of the abstract, helps users comprehensively understand the video content, and is beneficial to the development of medical training and insurance claim review work.

[0104] In summary, the solution implemented in the embodiments of the present invention can solve the problems existing in the video abstract generation method in the prior art.

[0105] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention. The non-company software tools or components that appear in the embodiments of this application are only for illustrative introduction and do not represent actual use.

[0106] In one embodiment, a video abstract generation device is provided. The video abstract generation device corresponds one-to-one with the video abstract generation method in the above embodiment. As Figure 7 shown, the video abstract generation device includes an acquisition module 710, an analysis module 720, a segmentation module 730, an extraction module 740, and a prediction module 750. The detailed description of each functional module is as follows:

[0107] The acquisition module 710 is used to acquire a target video;

[0108] The analysis module 720 is used to analyze different scenes in the target video;

[0109] The segmentation module 730 is used to segment the target video into sub-target videos according to different scenes of the target video, and the sub-target videos correspond one-to-one with different scenes in the target video;

[0110] The extraction module 740 is used to extract multimodal information of the sub-target videos;

[0111] The prediction module 750 is used to input the multimodal information of each sub-target video into a preset model to obtain a video abstract corresponding to the target video.

[0112] In one embodiment, the analysis module 720 is specifically used for:

[0113] Compare each video frame of the target video to obtain video frame landmark points that produce visual differences;

[0114] Determine the scenes corresponding to the video frames before the video frame landmark points and the scenes corresponding to the video frames after the video frame landmark points as different scenes in the target video.

[0115] In one embodiment, the analysis module 720 is further configured to:

[0116] Compare the pixel differences of pixels at the same positions in any two adjacent video frames in the target video. When the pixel difference is greater than a preset threshold, use the middle position in the two adjacent video frames as the video frame landmark point where visual differences occur.

[0117] In one embodiment, the extraction module 740 is specifically configured to:

[0118] Extract the text information in the sub-target video;

[0119] Extract the key frame image information in the sub-target video;

[0120] Input the text information and the key frame image information into a preset feature encoder to obtain a target feature vector for characterizing the comprehensive information of the text information and the key frame image information.

[0121] In one embodiment, the prediction module 750 is specifically configured to:

[0122] Extract the keyword information in the text information;

[0123] Input the target feature vector of the sub-target video and the keyword information into a preset model to obtain a video summary corresponding to the target video.

[0124] In one embodiment, the preset model is the encoder part of the Tranformer model.

[0125] In one embodiment, the prediction module 750 is further configured to:

[0126] Input the target feature vector of the sub-target video and the keyword information into a preset model to obtain a predicted video summary of the target video;

[0127] Input the predicted video summary of the target video and the actual video summary into a cross-entropy loss function, and calculate the loss value of the predicted video summary of the target video relative to the actual video summary;

[0128] When the loss value is less than a preset threshold, determine the predicted video summary as the video summary corresponding to the target video.

[0129] The solution provided by the video abstract generation device of the present invention divides the video into sub-target videos according to different scenes through scene segmentation, clearly showing the sequence and logical relationship between scenes, and effectively solving the problem of incoherent narration. For example, in a medical training surgery video, instead of isolated key frames, each surgical stage scene is presented completely, making the operation process coherent. In the solution, through scene segmentation and multi-modal information extraction, key events are identified at the scene level. Combining multi-modal information can understand key events from different abstraction levels. For example, in a surgical scene, key surgical steps can be understood from multiple modalities such as vision and language. The abstract generated according to scenes and multi-modal information is more organized, covers comprehensive content, improves the quality of the abstract, helps users comprehensively understand the video content, and facilitates quickly grasping key information to make accurate judgments.

[0130] The solution avoids the understanding deviation of a single modality through multi-modal information fusion. Different modality information complements each other, and can better sort out the structural relationship of the content. For example, when interpreting a medical imaging video, combining visual and other modality information can avoid misjudging normal tissues as lesions, making the abstract more logical. The multi-modal information extraction of the solution includes non-verbal information, and can identify roles and themes at different abstraction levels. For example, when a patient describes their condition, the patient's emotions and the theme of the condition can be understood from multi-modal information such as tone and expression. The solution reduces factual errors, captures the emotional tone, optimizes the calculation process to improve the processing speed, and improves the quality of the abstract through multi-modal information fusion and a preset model, enabling users to comprehensively and accurately understand the video content and meeting the needs of rapid claims settlement review.

[0131] The solution extracts multi-modal information, not limited to visual features, can better include core information, reduces the influence of factors such as light, and understands semantic relationships and sorts out the content structure through multi-modal information fusion and a preset model. For example, in a doctor-patient dialogue video, core medical information can be covered by combining multi-modal information such as language. The solution can understand the logic of doctor-patient communication and the causal relationship of the development of the condition from different abstraction levels through multi-modal information and a preset model. For example, understand the diagnostic logic of the condition from multi-modal information such as the dialogue content and the doctor's examination actions. The abstract generated by the solution covers the core content, has an internal logic, improves the quality of the abstract, helps users comprehensively understand the video content, and is conducive to the development of medical training and insurance claims settlement review work.

[0132] In summary, the solution implemented by the above video abstract generation method can solve the problems existing in the video abstract generation method in the prior art.

[0133] For the specific limitations of the video summary generation device, reference may be made to the limitations of the video summary generation method in the foregoing text, which will not be elaborated herein. Each module in the above video summary generation device can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.

[0134] Based on the above video summary generation method, as Figure 8 shown, an embodiment of the present invention further provides a structural schematic diagram of a video summary generation device, which includes a processor 81 and a memory 82 coupled to the processor 81. The memory 82 stores a computer program, and when the computer program is executed by the processor 81, the processor 81 is caused to execute the steps of the video summary generation method in the above embodiment.

[0135] For other details of the processor 81 in the above video summary generation device to implement the above technical solution, reference may be made to the description in the video summary generation method provided in the above embodiment of the present invention, which will not be elaborated herein.

[0136] Among them, the processor 81 can also be referred to as a CPU (Central Processing Unit), and the processor 81 may be an integrated circuit chip with signal processing capabilities; the processor 81 can also be a general-purpose processor, a DSP (Digital Signal Process), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, where the general-purpose processor can be a microprocessor or the processor 81 can also be any conventional processor, etc.

[0137] As Figure 9As shown in the figure, an embodiment of the present invention further provides a schematic structural diagram of a computer-readable storage medium, on which a readable computer program 91 is stored; wherein, the computer program 91 can be stored in the above storage medium in the form of a software product, including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, external hard drives, magnetic disks or optical discs, ROM (Read-Only Memory), RAM (Random Access Memory), etc. that can store program codes, or terminal devices such as computers, servers, mobile phones, tablets, etc.

[0138] In several embodiments provided by the present invention, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or modules can be in electrical, mechanical, or other forms.

[0139] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they can be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0140] In addition, in various embodiments of the present invention, the functional modules can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium.

[0141] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in the form of a computer program product in whole or in part.

[0142] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be stored by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium, or a semiconductor medium (such as an SSD (solid state disk)).

[0143] The technical solutions provided by the present invention have been introduced in detail above. Specific examples are used in the present invention to illustrate the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

[0144] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, optical storage, etc.) containing computer-usable program code.

[0145] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for implementing in the process Figure 1 one process or multiple processes and / or blocksFigure 1 means for the functions specified in one or more boxes.

[0146] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction means that implements in the process Figure 1 one process or more processes and / or boxes Figure 1 the functions specified in one or more boxes.

[0147] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to generate a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide for implementing in the process Figure 1 one process or more processes and / or boxes Figure 1 the steps of the functions specified in one or more boxes.

[0148] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and modifications. The non-company software tools or components that appear in the embodiments of the present invention are only introduced by way of example and do not represent actual use.

Claims

1. A method for generating a video summary, characterized in that: include: Get the target video; Analyzing different scenes in the target video; According to different scenes of the target video, the target video is divided into sub-target videos, and the sub-target videos correspond to different scenes in the target video one by one; Extracting multimodal information of the sub-target video; The multimodal information of each sub-target video is input into a preset model to obtain a video summary corresponding to the target video.

2. The method for generating a video summary according to claim 1, characterized in that: The analyzing different scenes in the target video includes: Comparing the video frames of the target video to obtain the video frame landmarks that produce visual differences; The scene corresponding to the video frame before the video frame marker point and the scene corresponding to the video frame after the video frame marker point are determined as different scenes in the target video.

3. The method for generating a video summary according to claim 2, characterized in that: The comparing the video frames of the target video to obtain the video frame landmarks that produce visual differences includes: The pixel difference between pixels at the same position in any two adjacent video frames in the target video is compared. When the pixel difference is greater than a preset threshold, the middle position in the two adjacent video frames is used as a video frame marker point where a visual difference occurs.

4. The method for generating a video summary according to claim 1, characterized in that: The extracting multimodal information of the sub-target video includes: Extracting text information from the sub-target video; Extracting key frame image information from the sub-target video; The text information and the key frame image information are input into a preset feature encoder to obtain a target feature vector for characterizing comprehensive information of the text information and the key frame image information.

5. The method for generating a video summary according to claim 4, characterized in that: The step of inputting the multimodal information of each sub-target video into a preset model to obtain a video summary corresponding to the target video includes: Extracting keyword information from the text information; The target feature vector of the sub-target video and the keyword information are input into a preset model to obtain a video summary corresponding to the target video.

6. The method for generating a video summary according to claim 5, characterized in that: The preset model is the encoder part of the Tranformer model.

7. The method for generating a video summary according to claim 5, characterized in that: The step of inputting the target feature vector of the sub-target video and the keyword information into a preset model to obtain a video summary corresponding to the target video includes: Inputting the target feature vector of the sub-target video and the keyword information into a preset model to obtain a predicted video summary of the target video; Inputting the predicted video summary and the actual video summary of the target video into a cross entropy loss function, and calculating a loss value of the predicted video summary of the target video relative to the actual video summary; When the loss value is less than a preset threshold, the predicted video summary is determined as the video summary corresponding to the target video.

8. A device for generating a video summary, characterized in that: include: An acquisition module, used to acquire a target video; An analysis module, used for analyzing different scenes in the target video; A segmentation module, used for segmenting the target video into sub-target videos according to different scenes of the target video, wherein the sub-target videos correspond to different scenes in the target video one by one; An extraction module, used for extracting multimodal information of the sub-target video; The prediction module is used to input the multimodal information of each sub-target video into a preset model to obtain a video summary corresponding to the target video.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method for generating a video summary according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method for generating a video summary according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Video data coding method and device, electronic equipment and storage medium

    CN121125992A

  • Method and system for automatically generating movie script and movie abstract

    CN121561139A