Medical video caching method and device, electronic equipment and storage medium
By gridding and associating medical video frames, the problem of large storage space occupied by high-resolution video data is solved, and video caching with efficient storage and fast access is achieved.
Patent Information
- Application Number
- CN202510856514.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-10-10
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing video caching technology takes up a lot of storage space when processing high-resolution video data, resulting in low efficiency. In particular, spatiotemporal redundant frames in medical videos are not effectively utilized, affecting the inference speed of medical models.
The original video frames are gridded and divided into background and lesion grid video frames, and associated tags are given. The background frames are compressed and the lesion frames are directly cached. The storage space is associated with the associated tags to reduce the storage amount.
It achieves efficient storage of medical video data, saves storage space, shortens storage time, improves video storage efficiency, and meets real-time diagnosis needs.
Smart Images

Figure CN120765449A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology and is applicable to medical technology scenarios, and in particular to a medical video caching method and device, electronic equipment, and storage medium. Background Art
[0002] Video caching is a technology that stores video data locally or on edge nodes to facilitate subsequent video processing. Video storage technology can be applied in a variety of scenarios, such as storing collected endoscopic images and ultrasound dynamic videos in medical technology scenarios.
[0003] Currently, video caching technology mainly uses a frame-by-frame storage method to store video data. However, when processing high-resolution video data, due to the large amount of video data, the frame-by-frame storage method will take up a lot of storage space, thus affecting the efficiency of video caching.
[0004] Therefore, how to improve the efficiency of video caching has become a technical problem that needs to be solved urgently. Summary of the Invention
[0005] The main purpose of the embodiments of the present application is to propose a medical video caching method and device, an electronic device and a storage medium, aiming to improve the efficiency of video caching.
[0006] To achieve the above-mentioned purpose, a first aspect of an embodiment of the present application proposes a medical video caching method, the method comprising:
[0007] Acquiring original medical video data, and performing video frame extraction on the original medical video data to obtain original video frames;
[0008] Performing gridding processing on the original video frame to obtain an original grid video frame;
[0009] Performing region division on the original grid video frame to obtain a background grid video frame and a lesion grid video frame; wherein the background grid video frame and the lesion grid video frame have associated labels;
[0010] Performing video compression on the background grid video frame to obtain a target background video frame;
[0011] Creating a lesion storage space based on the lesion grid video frame, and caching the lesion grid video frame into the lesion storage space;
[0012] The target background video frame is cached in a preset background storage space, and the lesion storage space is associated with the background storage space based on the association tag.
[0013] In some embodiments, the creating a lesion storage space based on the lesion grid video frame comprises:
[0014] performing type identification on the lesion grid video frame to obtain a lesion type;
[0015] determining a lesion heat parameter based on the lesion type;
[0016] performing heat calculation based on the lesion heat parameter to obtain a target access heat factor;
[0017] creating a storage space based on the target access heat factor to obtain the lesion storage space.
[0018] In some embodiments, the background grid video frame comprises a plurality of background space grid units, each of the background space grid units containing a background sub-video frame; and the video compression on the background grid video frame to obtain a target background video frame comprises:
[0019] grouping the background sub-video frames for each of the background space grid units to obtain a plurality of groups of continuous video frames;
[0020] performing similarity calculation on the continuous video frames for each group of the continuous video frames to obtain inter-frame similarity data;
[0021] performing video compression on the background grid video frame according to the inter-frame similarity data to obtain the target background video frame.
[0022] In some embodiments, the video compression on the background grid video frame according to the inter-frame similarity data to obtain the target background video frame comprises:
[0023] dividing a compression region of the background grid video frame according to the inter-frame similarity data to obtain a compressed grid video frame and a non-compressed grid video frame;
[0024] performing video frame compression on the compressed grid video frame to obtain a compressed video frame;
[0025] fusing the compressed video frame and the non-compressed grid video frame to obtain the target background video frame.
[0026] In some embodiments, the compressed grid video frame comprises a grid position parameter and a grid time parameter; and the video frame compression on the compressed grid video frame to obtain a compressed video frame comprises:
[0027] obtaining an original feature value of each of the compressed grid video frames;
[0028] performing pooling processing on the original feature value to obtain a pooled feature value;
[0029] Performing singular value decomposition on the pooled eigenvalues to obtain the compressed video frame.
[0030] In some embodiments, dividing the original grid video frame into regions to obtain a background grid video frame and a lesion grid video frame includes:
[0031] Extracting motion features from the original grid video frame to obtain grid motion features;
[0032] The original grid video frame is segmented into regions based on the grid motion features to obtain the background grid video frame and the lesion grid video frame.
[0033] In some embodiments, performing gridding processing on the original video frame to obtain the original grid video frame includes:
[0034] Obtaining a resolution parameter of the original video frame;
[0035] Performing grid calculation based on preset grid size parameters and the resolution parameters to obtain grid division ratio data;
[0036] The original video frame is segmented based on the grid division ratio data to obtain the original grid video frame.
[0037] To achieve the above-mentioned purpose, a second aspect of an embodiment of the present application provides a medical video caching device, the device comprising:
[0038] A video data acquisition module is used to acquire original medical video data and perform video frame extraction on the original medical video data to obtain original video frames;
[0039] A video gridding processing module, configured to perform gridding processing on the original video frame to obtain an original grid video frame;
[0040] A video region division module is used to divide the original grid video frame into regions to obtain a background grid video frame and a lesion grid video frame; wherein the background grid video frame and the lesion grid video frame have associated labels;
[0041] A video compression module, configured to compress the background grid video frame to obtain a target background video frame;
[0042] a lesion video caching module, configured to create a lesion storage space based on the lesion grid video frame, and cache the lesion grid video frame into the lesion storage space;
[0043] The compressed video cache module is used to cache the target background video frame to a preset background storage space, and associate the lesion storage space with the background storage space based on the association tag.
[0044] To achieve the above-mentioned purpose, the third aspect of an embodiment of the present application proposes an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the method described in the first aspect when executing the computer program.
[0045] To achieve the above-mentioned purpose, the fourth aspect of the embodiments of the present application proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the method described in the first aspect.
[0046] The medical video caching method and device, electronic device, and storage medium proposed in this application obtain raw medical video data and extract video frames from it to generate raw video frames, laying the foundation for subsequent video caching. Next, the raw video frames are gridded, subdividing them into multiple spatial grids, facilitating more refined analysis and storage of video content. The raw grid video frames are divided into regions, splitting them into background grid video frames and lesion grid video frames. These frames are then assigned association tags, clearly distinguishing between critical and non-critical information and facilitating targeted processing of video frames in different regions. Furthermore, background grid video frames are compressed, effectively reducing data storage and conserving storage space. A lesion storage space is created based on the lesion grid video frames, and the lesion grid video frames are cached in the lesion storage space without compression, ensuring the clarity of the lesion video frames. The target background video frames are cached in a pre-set background storage space, conserving storage space. Finally, the lesion storage space is associated with the background storage space via association tags. By reducing the storage volume of medical video data, the storage time of data can be shortened, efficient storage of medical video data is achieved, and the efficiency of video storage is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 is a flowchart of the medical video caching method provided in an embodiment of the present application;
[0048] Figure 2 yes Figure 1 Flowchart of step S102 in FIG.
[0049] Figure 3 yes Figure 1 Flowchart of step S103 in FIG.
[0050] Figure 4 yes Figure 1 Flowchart of step S104 in FIG.
[0051] Figure 5 yes Figure 4 Flowchart of step S403 in FIG.
[0052] Figure 6 yes Figure 5 Flowchart of step S502 in FIG.
[0053] Figure 7 is a flowchart of a medical video caching method provided by another embodiment of the present application;
[0054] Figure 8 Schematic diagram of the structure of the medical video caching device provided in an embodiment of the present application;
[0055] Figure 9 This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0056] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0057] It should be noted that although the device schematics illustrate functional module divisions and the flowcharts illustrate logical sequences, in certain circumstances, the steps shown or described may be performed in a sequence that differs from the module divisions in the device or the sequence in the flowcharts. The terms "first," "second," and so on, in the specification, claims, and drawings, are used to distinguish similar items and are not necessarily used to describe a specific sequence or precedence.
[0058] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0059] First, let’s analyze some of the terms used in this application:
[0060] Artificial intelligence (AI) is a new technical discipline that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. A branch of computer science, AI seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thinking. It also encompasses the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results.
[0061] Video caching refers to the technology that temporarily stores video data in a local device (such as memory, hard drive, or edge server) in advance or in real time to optimize playback smoothness, reduce network load, and facilitate subsequent video processing. The core principles of video caching technology include segmented caching (such as the HLS / DASH protocol, which splits videos into small files for on-demand loading), preloading (buffering subsequent content in advance), and edge caching (distributing content locally through CDN nodes). It can effectively alleviate network fluctuations, reduce latency, and reduce pressure on source servers, making it particularly suitable for high-concurrency streaming scenarios.
[0062] Video storage technology can be applied to a variety of application scenarios. For example, in medical technology scenarios, it can store collected endoscopic images, ultrasound dynamic videos, etc.
[0063] Currently, video caching technology mainly uses a frame-by-frame storage method to store video data. However, when processing high-resolution video data, due to the large amount of video data, the frame-by-frame storage method will take up a lot of storage space, thus affecting the efficiency of video caching.
[0064] In addition, medical videos usually contain a large number of temporally and spatially repeated or similar frames (such as static backgrounds and periodic movements of organs), but existing caching technologies do not effectively utilize such characteristics, resulting in storage redundancy and occupying a large amount of storage space, which in turn affects the inference speed of medical models and makes it difficult to meet real-time diagnosis needs.
[0065] Based on this, the embodiments of the present application provide a medical video caching method and device, an electronic device, and a storage medium, aiming to improve the efficiency of video caching.
[0066] The medical video caching method and device, electronic device, and storage medium provided in the embodiments of the present application are specifically described through the following embodiments. First, the medical video caching method in the embodiments of the present application is described.
[0067] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.
[0068] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0069] The medical video caching method provided in the embodiment of the present application relates to the field of artificial intelligence technology. The medical video caching method provided in the embodiment of the present application can be applied to a terminal, can be applied to a server side, or can be software running in a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or as a server cluster or distributed system composed of multiple physical servers, or as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the medical video caching method, etc., but is not limited to the above forms.
[0070] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.
[0071] It should be noted that in each specific embodiment of the present application, when it comes to the need to perform relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first, and the collection, use, and processing of such data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present application needs to obtain the user's sensitive personal information, the user's separate permission or consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the user's separate permission or consent, the necessary user-related data for the normal operation of the embodiment of the present application will be obtained.
[0072] Figure 1 This is an optional flowchart of the medical video caching method provided in the embodiment of the present application. Figure 1 The method may include but is not limited to steps S101 to S106.
[0073] Step S101, obtaining original medical video data, and performing video frame extraction on the original medical video data to obtain original video frames;
[0074] Step S102, performing gridding processing on the original video frame to obtain an original grid video frame;
[0075] Step S103, dividing the original grid video frame into regions to obtain a background grid video frame and a lesion grid video frame; wherein the background grid video frame and the lesion grid video frame have associated labels;
[0076] Step S104, performing video compression on the background grid video frame to obtain a target background video frame;
[0077] Step S105, creating a lesion storage space based on the lesion grid video frame, and caching the lesion grid video frame into the lesion storage space;
[0078] Step S106 , buffering the target background video frame into a preset background storage space, and associating the lesion storage space with the background storage space based on the association tag.
[0079] The steps S101 to S106 shown in the embodiments of the present application obtain the original video frame by obtaining the original medical video data and performing video frame extraction on the original medical video data, thereby laying a foundation for subsequent video caching. Then, the original video frame is subjected to grid processing, and the video frame is subdivided into a plurality of spatial grids, which helps to more finely analyze and store the video content. The original grid video frame is divided into a background grid video frame and a lesion grid video frame by region, and is assigned an associated label, thereby clearly distinguishing between key and non-key information and facilitating targeted processing of video frames in different regions. Further, the background grid video frame is subjected to video compression, thereby effectively reducing the data storage amount and saving storage space. The lesion storage space is created based on the lesion grid video frame, and the lesion grid video frame is cached to the lesion storage space without compression, thereby ensuring the clarity of the lesion video frame. The target background video frame is cached to the preset background storage space, thereby saving storage space. Finally, the lesion storage space and the background storage space are associated through the associated label. By reducing the storage amount of the medical video data, the storage time of the data can be shortened, thereby realizing efficient storage of the medical video data and improving the efficiency of video storage.
[0080] In step S101 of some embodiments, the original medical video data is video data obtained in advance or in real time in a medical scene, including but not limited to endoscopic images, ultrasonic dynamic videos, CT / MRI images, cell observation videos, etc.
[0081] In some embodiments, the original medical video data can be subjected to video frame extraction by means such as a deep learning model, a machine learning model, and a video processing tool, etc., to obtain the original video frame, without being limited thereto.
[0082] Please refer to Figure 2 In some embodiments, step S102 can include but is not limited to steps S201 to S203:
[0083] Step S201, obtaining a resolution parameter of the original video frame;
[0084] Step S202, performing grid calculation based on the preset grid size parameter and the resolution parameter to obtain grid division ratio data;
[0085] Step S203, performing picture cutting on the original video frame based on the grid division ratio data to obtain the original grid video frame.
[0086] Steps S201 to S203, as shown in the embodiment of the present application, flexibly adapt to videos of different resolutions by obtaining the resolution parameters of the original video frame and calculating grid division ratio data based on preset grid size parameters and resolution parameters. Subsequently, the original video frame is segmented based on the grid division ratio data to obtain the original grid video frame. This ensures the standardization and consistency of the segmentation, facilitating more accurate and effective subsequent processing and analysis of the video frame.
[0087] In step S201 of some embodiments, the resolution parameter is used to describe the pixel size of the video frame, including two values of width and length, such as 1920×1080, 1280×720, etc.
[0088] In step S202 of some embodiments, the preset grid size parameter is the pixel size of the grid unit, for example, a grid unit divided into 14×28 pixels.
[0089] It should be noted that the spatial grid size parameters need to be determined according to the video resolution and scene requirements, so that they can flexibly adapt to the processing requirements of different medical video data.
[0090] For example, for video data with a resolution of 1920×1080, the corresponding preset grid size parameter may be set to 96×54; for video data with a resolution of 1280×720, the corresponding preset grid size parameter may be set to 64×36.
[0091] Furthermore, grid calculation is performed according to the grid size parameter and the resolution parameter to obtain grid division ratio data, and then the original video frame is segmented based on the grid division ratio data to obtain the original grid video frame.
[0092] For example, if the resolution parameter is 1920×1080 and the grid size parameter is 96×54, the grid division ratio is: 20 equal-spaced horizontal divisions and 20 equal-spaced vertical divisions. After the screen is cut, 400 original grid video frames are obtained.
[0093] Specifically, the original grid video frame obtained after screen segmentation can be numbered in a manner similar to that of matrix elements, ie, x[i,j], representing the jth original grid video frame in the i-th row.
[0094] It should be noted that each video frame is divided using the same grid division ratio, which facilitates subsequent comparison and analysis of the same original grid video frames in consecutive video frames.
[0095] See also Figure 3 In some embodiments, step S103 may include but is not limited to steps S301 to S302:
[0096] Step S301, extracting motion features from the original grid video frame to obtain grid motion features;
[0097] Step S302 : performing region segmentation on the original grid video frame based on the grid motion feature to obtain a background grid video frame and a lesion grid video frame.
[0098] In steps S301 and S302, illustrated in the present embodiment, motion features are extracted from the original grid video frames to obtain grid motion features, accurately capturing the dynamic changes in different areas of the video. Next, the original grid video frames are segmented based on the grid motion features to obtain background grid video frames and lesion grid video frames. This allows for precise distinction between the dynamically changing lesion area and the relatively stable background area, providing clearer data for subsequent processing and analysis.
[0099] In step S301 of some embodiments, an optical flow algorithm or feature matching technology can be used to identify the motion vectors of the original grid video frames with the same number under continuous original video frames to obtain grid motion features, which are used to reflect the displacement direction, amplitude, speed and other information of the pixel points in the original grid video frames. This can effectively depict subtle movements in the video frames and detect dynamic changes in the lesion area.
[0100] In step S302 of some embodiments, a clustering algorithm (such as K-means) or a threshold method can be used to distinguish grid cells with high and low motion intensity based on grid motion characteristics, classifying grid cells with high motion intensity as lesion areas and grid cells with low motion intensity as background areas, thereby obtaining background grid video frames for the background area and lesion grid video frames for the lesion area. This method is applicable to different modal medical videos (such as ultrasound and endoscopy) and does not rely on prior anatomical knowledge.
[0101] Specifically, background grid video frames are grid areas with weak or regular motion features, usually corresponding to normal tissue or fixed anatomical structures. Lesion grid video frames are grid areas with abnormal motion features (such as irregular jitter or significant displacement), which may correspond to pathological tissue.
[0102] It should be noted that each original video frame is divided into a background grid video frame and a lesion grid video frame. A background grid video frame includes multiple background spatial grid cells, and the image information contained in each background spatial grid cell is the background sub-video frame. Similarly, a lesion grid video frame includes multiple lesion spatial grid cells, and the image information contained in each lesion spatial grid cell is the lesion sub-video frame.
[0103] In some embodiments, the associated labels of the background grid video frame and the lesion grid video frame are used to characterize: in the nth video frame, the jth grid video frame in the i-th row belongs to which spatial grid unit, whether it is a background spatial grid unit or a lesion spatial grid unit, so that the corresponding background sub-video frame or lesion sub-video frame can be obtained according to the associated labels to form a complete video frame.
[0104] See also Figure 4 In some embodiments, step S104 may include but is not limited to steps S401 to S403:
[0105] Step S401: for each background space grid unit, grouping background sub-video frames to obtain multiple groups of continuous video frames;
[0106] Step S402 , for each group of continuous video frames, performing similarity calculation on the continuous video frames to obtain inter-frame similarity data;
[0107] Step S403 : compressing the background grid video frame according to the inter-frame similarity data to obtain a target background video frame.
[0108] In steps S401 to S403, as shown in the embodiment of the present application, for each background spatial grid unit, background sub-video frames are grouped to obtain multiple groups of continuous video frames. Next, inter-frame similarity data is calculated for each group of continuous video frames, accurately capturing inter-frame correlations. Finally, video compression is performed on the background grid video frames based on the inter-frame similarity data to obtain the target background video frame. This effectively removes redundant information, significantly reducing the amount of video data while ensuring that key video information is not lost, achieving efficient video compression and saving storage and transmission costs.
[0109] In step S401 of some embodiments, background sub-video frames are grouped based on a preset time window or a fixed frame number to obtain multiple groups of continuous video frames. The preset time window can be set to 1 second or 0.5 seconds, and the fixed frame number can be set to 5 frames per group or 10 frames per group, but is not limited thereto.
[0110] Furthermore, by grouping the background sub-video frames for each background space grid unit, multiple groups of continuous video frames are obtained, and similarity calculation is performed on the continuous video frames to obtain inter-frame similarity data; thereby, the similarity between the background sub-video frames of the same background space grid unit in the continuous video frames can be measured to determine whether they are similar areas.
[0111] Specifically, for the same background spatial grid unit, the similarity calculation is performed on the feature vectors of the background sub-video frames of consecutive video frames. The Euclidean distance or other distance metrics can be used to represent the difference between the two to obtain the inter-frame similarity data, see formula (1):
[0112] d(R)=||x t -x t-1 || (1);
[0113] Among them, d(R) is the inter-frame similarity data, R represents the Rth background space grid unit, x t Indicates the background sub-video frame of the current frame, x t-1 Indicates the background sub-video frame of the previous frame of the current frame.
[0114] In step S403 of some embodiments, the inter-frame similarity data is compared with a preset similarity threshold. If the inter-frame similarity data exceeds the preset similarity threshold, indicating that the background sub-video frames of the same background spatial grid unit in consecutive video frames are similar, compression processing can be performed.
[0115] If the inter-frame similarity data is less than a preset similarity threshold, indicating that background sub-video frames of the same background spatial grid unit in consecutive video frames are dissimilar, they can be cached with complete key-value features.
[0116] See also Figure 5 In some embodiments, step S403 may also include but is not limited to steps S501 to S503:
[0117] Step S501, performing compression region division on the background grid video frame according to inter-frame similarity data to obtain compressed grid video frames and uncompressed grid video frames;
[0118] Step S502, compressing the compressed grid video frame to obtain a compressed video frame;
[0119] Step S503: Fusing the compressed video frame and the uncompressed grid video frame to obtain a target background video frame.
[0120] In steps S501 to S503, as shown in the embodiment of the present application, the background grid video frames are divided into compressed regions based on inter-frame similarity data, and compressed grid video frames and uncompressed grid video frames are accurately identified, thereby improving compression efficiency. Next, the compressed grid video frames are compressed to obtain compressed video frames, and the compressed video frames and uncompressed grid video frames are fused to obtain the target background video frame. This not only preserves the clarity of the important uncompressed regions, but also reduces the overall data volume, significantly improving the efficiency of video storage and transmission.
[0121] In step S501 of some embodiments, the inter-frame similarity data is compared with a preset similarity threshold to obtain a comparison result. According to the comparison result, the background grid video frame is divided into a compressed grid video frame and a non-compressed grid video frame.
[0122] If the inter-frame similarity data exceeds the preset similarity threshold, it indicates that the background sub-video frame of the same background space grid unit in the continuous video frame is a compressed grid video frame, which can be compressed.
[0123] If the inter-frame similarity data is less than the preset similarity threshold, it indicates that the background sub-video frame of the same background space grid unit in the continuous video frame is a non-compressed grid video frame, which is cached with complete key-value features.
[0124] In step S502 of some embodiments, for the region identified as similar, i.e., the compressed grid video frame, the key-value features of the compressed grid video frame can be compressed using the average pooling and low-rank approximation method to reduce redundant storage.
[0125] Referring to Figure 6 In some embodiments, step S502 includes but is not limited to steps S601 to S603:
[0126] Step S601: obtaining the original feature value of each compressed grid video frame;
[0127] Step S602: performing pooling processing on the original feature value to obtain a pooled feature value;
[0128] Step S603: performing singular value decomposition on the pooled feature value to obtain a compressed video frame.
[0129] The steps S601 to S603 shown in the embodiments of the present application obtain the original feature value of each compressed grid video frame, which completely retains the initial feature information of the video frame. Then, the original feature value is pooled to obtain a pooled feature value, which can effectively reduce the feature dimension, reduce data redundancy, and retain key features. Finally, the pooled feature value is singularly decomposed to obtain a compressed video frame, which can further compress data and extract core features, reducing the amount of data while efficiently extracting and retaining key features of the video frame, which helps to improve the efficiency and quality of video caching.
[0130] In step S601 of some embodiments, the original feature value of the compressed grid video frame is the key-value feature of the compressed grid video frame, which contains the underlying features extracted from the video frame, such as pixel intensity, gradient, frequency domain coefficient, etc., reflecting the local or global information of the video frame.
[0131] Specifically, a feature extraction algorithm (such as convolution operation, optical flow analysis, etc.) can be used to extract original feature values from the compressed grid video frames.
[0132] In step S602 of some embodiments, the average pooling method is used to perform pooling processing on the original eigenvalues, which can aggregate local feature areas and output pooled eigenvalues, helping to reduce feature dimensions and improve computing efficiency.
[0133] Finally, the pooled eigenvalues are subjected to singular value decomposition, the first k largest singular values are retained (low-rank approximation), the minor components are eliminated, and the features after dimensionality reduction are reconstructed to obtain compressed video frames, which helps to reduce the data volume of the video frames, save storage space, and thus improve the efficiency of video caching.
[0134] For example, the key features of the same background space grid unit in N adjacent frames are recorded as {K1, K2, ..., K N};
[0135] The average pooling method is used to compress the video frame. The compression formula is shown in formula (2):
[0136]
[0137] in, To compress the background video frame, K i is the key-value feature of the compressed grid video frame of the i-th frame.
[0138] Then, the singular value decomposition technique is used to compress the background video frame. Low-rank approximation is performed to obtain compressed video frames, which constitute the key feature description after compression, thereby ensuring the expression of effective data information and reducing the storage burden.
[0139] It is understandable that compressed video frames can be reused, which is equivalent to multiple video frames being represented by one compressed video frame. Therefore, they only need to be stored once to meet the storage needs of multiple frames, save storage space, and improve the efficiency of video caching.
[0140] It should be noted that the compressed grid video frame includes a grid position parameter (ie, a grid number) and a grid time parameter.
[0141] In step S503 of some embodiments, after the compressed grid video frame is compressed to obtain a compressed video frame, the compressed video frame and the uncompressed grid video frame are fused based on the numbered grid position parameters and grid time parameters of the video frame to obtain a target background video frame. This facilitates combining the compressed video frame data into a complete video frame when subsequently reading and using the video frame data, thereby ensuring the accuracy of the information.
[0142] See also Figure 7In some embodiments, the step of "creating a lesion storage space based on the lesion grid video frame" in step S105 may include but is not limited to steps S701 to S704:
[0143] Step S701, performing type recognition on the lesion grid video frame to obtain the lesion type;
[0144] Step S702, determining a lesion heat parameter based on the lesion type;
[0145] Step S703, performing heat calculation based on the lesion heat parameter to obtain a target access heat factor;
[0146] Step S704: creating a storage space based on the target access heat factor to obtain a lesion storage space.
[0147] In steps S701 to S704, as shown in the embodiment of the present application, the lesion grid video frame is identified to obtain the lesion type, and the lesion heat parameter is determined based on the lesion type, which can reflect the search attention level of the lesion. Then, the heat calculation is performed based on the lesion heat parameter to obtain the target access heat factor, and storage space is created based on the target access heat factor to obtain the lesion storage space. This can be combined with the access characteristics of medical data and adopt a dynamic cache allocation strategy to ensure that high-access areas can obtain more cache resources, ensuring the clarity and accuracy of the cached video frames, thereby improving the quality of medical video cache.
[0148] In step S701 of some embodiments, an image classification model (such as a convolutional neural network) or an expert rule base may be used to determine the type of lesion, such as tumor, inflammation, calcification, etc. At the same time, the lesion type may be supplemented by combining corresponding items of the medical video data.
[0149] Furthermore, lesion heat parameters are determined based on the lesion type, wherein the lesion heat parameters include the historical occurrence frequency of the lesion and the vivid heat parameters of the lesion, so that heat calculation can be performed based on the lesion heat parameters to obtain the target access heat factor.
[0150] It should be noted that the historical occurrence frequency of the lesion and the vivid heat parameters of the lesion are related to the lesion type and are set according to the actual scene requirements.
[0151] Among them, the heat calculation can be achieved through the following formula:
[0152] h(R) = λ1*historical frequency of lesions + λ2*brightness heat parameter (3);
[0153] Where R is the lesion grid video frame, h(R) is the target access heat factor, λ1 and λ2 are preset hyperparameters, and λ1+λ2=1.
[0154] It should be noted that the settings of hyperparameters λ1 and λ2 need to be set according to the actual scenario requirements. For example, if you focus on real-time frequency, set λ1 to a larger value; if you focus on storage quality, set λ1 to a smaller value.
[0155] Furthermore, a storage space is created based on the calculated target access heat factor to obtain the lesion storage space. Specifically, the calculation process of the storage space size can be implemented by the following formula:
[0156] C(R)=C base +α·A(h(R)) (4);
[0157] Among them, C(R) is the lesion storage space, C base is the basic cache amount, α is the adjustment parameter, which is a constant, and A(h(R)) is the storage resource allocation function.
[0158] A(h(R))=log(h(R)+1) (5);
[0159] It should be noted that A(h(R)) is positively correlated with the target access heat factor of the lesion grid video frame R. The higher the target access heat factor, the larger the allocated cache space. Therefore, the storage resource allocation function introduces a logarithmic function to process the increment of high-heat lesion grid video frames more smoothly.
[0160] In step S105 of some embodiments, after the lesion storage space is created, the key-value features of the lesion grid video frame are cached to the lesion storage space, and a mapping relationship between the addresses of the lesion grid video frame and the lesion storage space is established to obtain a first mapping relationship.
[0161] Furthermore, in step S106 of some embodiments, the target background video frame is cached in a preset background storage space, and a mapping relationship is established between the target background video frame and the address of the background storage space to obtain a second mapping relationship. The preset background storage space is pre-created and has an initial size. When the background storage space is full, the size can be dynamically increased.
[0162] Finally, the first mapping relationship, the second mapping relationship and the associated label are fused to obtain the storage label. Based on the storage label, the lesion storage space is associated with the background storage space, so that when the video frame data is subsequently read and used, the key value data corresponding to each video frame can be accurately obtained based on the storage label, thereby ensuring the integrity of the picture.
[0163] For example:
[0164] Suppose there is medical video data containing 4 video frames, which has been stored using the video caching method provided in the embodiment of the present application. The video frames are divided into 3 parts with equal spacing horizontally and 3 parts with equal spacing vertically, totaling 9 grid units.
[0165] The first video frame is recorded as The second video frame is recorded as The third video frame is recorded as The fourth video frame is recorded as
[0166] Among them, element 0 represents the background grid video frame, and element 1 represents the lesion grid video frame.
[0167] The video frame in row 1 and column 1 of each video frame belongs to the background grid video frame, is similar, and has been compressed. For row 1 and column 1 of each video frame, the address in the background storage space is determined based on the second mapping relationship in the storage tag and the associated tag. Based on the same address, the same video frame is retrieved as the video frame in row 1 and column 1.
[0168] The video frame in the 1st row and 3rd column of each video frame belongs to the background grid video frame and is similar and has been compressed. The acquisition process is the same as that of the video frame in the 1st row and 1st column, so it will not be repeated.
[0169] As for the remaining locations, which are not compressed, the specific addresses in the corresponding storage space can be determined according to the storage tags and directly obtained.
[0170] In some embodiments, the medical video caching method provided in the embodiments of the present application can also be used for real-time caching, and after caching, the semantic information of the video frame is captured through a preset hierarchical model;
[0171] Specifically, the pre-qualified hierarchical model includes a top-level network, a bottom-level network, and a converged network.
[0172] The top-level network extracts video information from medical video data that has not been stored (i.e., the video data that has just been acquired) to obtain first video information, and simultaneously caches the medical video data.
[0173] The underlying network term extracts video information from the stored medical video data to obtain second video information;
[0174] Finally, the first video information and the second video information are fused through a fusion network, which can ensure information accuracy while improving computing efficiency.
[0175] The core of the medical video caching method provided in the embodiments of this application is to utilize a video key-value caching mechanism, which is different from the traditional frame-by-frame storage method. This application combines the spatiotemporal continuity of video content to compress similar frames or non-critical areas (background areas), thereby reducing memory usage and improving video caching efficiency.
[0176] In medical technology scenarios, the medical video caching method provided by the embodiments of this application can adapt to devices with limited computing power. By dynamically allocating resources, it solves problems such as storage redundancy, lack of real-time performance, and wasted computing power, while retaining key diagnostic information, meeting the needs of clinical real-time analysis and improving inference efficiency. Furthermore, it can take into account existing VLM architectures (such as the VideoBERT model and the Flamingo model), eliminating the need to modify the model backbone network and making it easy to integrate into medical model platforms.
[0177] See also Figure 8 The embodiment of the present application further provides a medical video caching device, which can implement the above-mentioned medical video caching method, and the device includes:
[0178] The video data acquisition module 801 is used to acquire original medical video data and perform video frame extraction on the original medical video data to obtain original video frames;
[0179] The video gridding processing module 802 is used to perform gridding processing on the original video frame to obtain the original grid video frame;
[0180] The video region division module 803 is used to divide the original grid video frame into regions to obtain a background grid video frame and a lesion grid video frame; wherein the background grid video frame and the lesion grid video frame have associated labels;
[0181] The video compression module 804 is used to compress the background grid video frame to obtain a target background video frame;
[0182] A lesion video caching module 805 is configured to create a lesion storage space based on the lesion grid video frame and cache the lesion grid video frame into the lesion storage space;
[0183] The compressed video cache module 806 is configured to cache the target background video frame into a preset background storage space, and associate the lesion storage space with the background storage space based on the association tag.
[0184] The specific implementation of the medical video caching device is basically the same as the specific embodiment of the above-mentioned medical video caching method, and will not be repeated here.
[0185] The present application also provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the above-described medical video caching method when executing the computer program. The electronic device can be any smart terminal, such as a tablet computer or an in-vehicle computer.
[0186] See also Figure 9 , Figure 9 The hardware structure of an electronic device according to another embodiment is shown. The electronic device includes:
[0187] The processor 901 can be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;
[0188] The memory 902 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program codes are stored in the memory 902 and are called by the processor 901 to execute the medical video caching method of the embodiments of this application.
[0189] Input / output interface 903, used to implement information input and output;
[0190] Communication interface 904, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);
[0191] Bus 905 , which transmits information between various components of the device (e.g., processor 901 , memory 902 , input / output interface 903 , and communication interface 904 );
[0192] The processor 901 , the memory 902 , the input / output interface 903 and the communication interface 904 are connected to each other in communication within the device via a bus 905 .
[0193] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned medical video caching method is implemented.
[0194] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0195] The medical video caching method and apparatus, electronic device, and storage medium provided in embodiments of the present application obtain raw medical video data and perform video frame extraction on the raw medical video data to obtain raw video frames, laying the foundation for subsequent video caching. Next, the raw video frames are gridded, subdividing the video frames into multiple spatial grids, facilitating more refined analysis and storage of video content. The raw grid video frames are divided into regions, splitting them into background grid video frames and lesion grid video frames, and assigning association tags to clearly distinguish between critical and non-critical information, facilitating targeted processing of video frames in different regions. Furthermore, video compression is performed on the background grid video frames, effectively reducing the amount of data stored and conserving storage space. A lesion storage space is created based on the lesion grid video frames, and the lesion grid video frames are cached in the lesion storage space without compression, ensuring the clarity of the lesion video frames. The target background video frames are cached in a preset background storage space, conserving storage space. Finally, the lesion storage space is associated with the background storage space via association tags. By reducing the storage volume of medical video data, the storage time of data can be shortened, efficient storage of medical video data is achieved, and the efficiency of video storage is improved.
[0196] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0197] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0198] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0199] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.
[0200] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0201] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0202] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0203] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0204] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0205] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: various media that can store programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0206] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.
Claims
1. A medical video caching method, characterized in that: The method comprises: Acquiring original medical video data, and performing video frame extraction on the original medical video data to obtain original video frames; Performing gridding processing on the original video frame to obtain an original grid video frame; Performing region division on the original grid video frame to obtain a background grid video frame and a lesion grid video frame; wherein the background grid video frame and the lesion grid video frame have associated labels; Performing video compression on the background grid video frame to obtain a target background video frame; Creating a lesion storage space based on the lesion grid video frame, and caching the lesion grid video frame into the lesion storage space; The target background video frame is cached in a preset background storage space, and the lesion storage space is associated with the background storage space based on the association tag.
2. The method according to claim 1, characterized in that The creating a lesion storage space based on the lesion grid video frame includes: Performing type recognition on the lesion grid video frame to obtain a lesion type; determining a lesion heat parameter based on the lesion type; Performing heat calculation based on the lesion heat parameter to obtain a target access heat factor; A storage space is created based on the target access heat factor to obtain the lesion storage space.
3. The method according to claim 1, characterized in that The background grid video frame includes a plurality of background spatial grid units, each of which includes a background sub-video frame; and the video compression of the background grid video frame to obtain a target background video frame includes: For each of the background spatial grid units, grouping the background sub-video frames to obtain multiple groups of continuous video frames; For each group of the continuous video frames, performing similarity calculation on the continuous video frames to obtain inter-frame similarity data; The background grid video frame is compressed according to the inter-frame similarity data to obtain the target background video frame.
4. The method according to claim 3, characterized in that The performing video compression on the background grid video frame according to the inter-frame similarity data to obtain the target background video frame includes: Performing compression region division on the background grid video frame according to the inter-frame similarity data to obtain compressed grid video frames and uncompressed grid video frames; Performing video frame compression on the compressed grid video frame to obtain a compressed video frame; The compressed video frame and the uncompressed grid video frame are fused to obtain the target background video frame.
5. The method according to claim 4, characterized in that The compressed grid video frame includes a grid position parameter and a grid time parameter; and the video frame compression is performed on the compressed grid video frame to obtain a compressed video frame, including: Obtaining original feature values of each of the compressed grid video frames; Performing pooling processing on the original eigenvalues to obtain pooled eigenvalues; Performing singular value decomposition on the pooled eigenvalues to obtain the compressed video frame.
6. The method according to any one of claims 1 to 5, characterized in that The step of dividing the original grid video frame into regions to obtain a background grid video frame and a lesion grid video frame includes: Extracting motion features from the original grid video frame to obtain grid motion features; The original grid video frame is segmented into regions based on the grid motion features to obtain the background grid video frame and the lesion grid video frame.
7. The method according to any one of claims 1 to 5, characterized in that The gridding process of the original video frame to obtain the original grid video frame includes: Obtaining a resolution parameter of the original video frame; Performing grid calculation based on preset grid size parameters and the resolution parameters to obtain grid division ratio data; The original video frame is segmented based on the grid division ratio data to obtain the original grid video frame.
8. A medical video caching device, characterized in that: The device comprises: A video data acquisition module is used to acquire original medical video data and perform video frame extraction on the original medical video data to obtain original video frames; A video gridding processing module, configured to perform gridding processing on the original video frame to obtain an original grid video frame; A video region division module is used to divide the original grid video frame into regions to obtain a background grid video frame and a lesion grid video frame; wherein the background grid video frame and the lesion grid video frame have associated labels; A video compression module, configured to compress the background grid video frame to obtain a target background video frame; a lesion video caching module, configured to create a lesion storage space based on the lesion grid video frame, and cache the lesion grid video frame into the lesion storage space; The compressed video cache module is used to cache the target background video frame to a preset background storage space, and associate the lesion storage space with the background storage space based on the association tag.
9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Method and device for detecting lesions by using ultrasound images and computer vision
CN111227864A
Video data transmission method and device based on key frame and storage medium
CN113938666A
Data storage optimization method based on medical field and electronic equipment
CN118092788A
Case image resource library construction method of smart medical system
CN119181472A
Breathing endoscope video collecting and processing method, electronic equipment and storage medium
CN119893052A