Low-light medical video enhancement method, system, device and medium with style guidance and temporal compensation
By employing a bidirectional transformation branch network framework, an inter-frame compensation module, and a text-guided style representation module, the robustness and imaging style adaptability issues in low-light medical video enhancement were addressed, achieving high-quality medical video enhancement and improving the accuracy and safety of clinical diagnosis and surgery.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN UNIV
- Filing Date
- 2025-07-16
- Publication Date
- 2026-05-01
AI Technical Summary
Existing low-light medical video enhancement methods lack robustness and generalization ability in medical scenarios, are difficult to adapt to complex lighting environments, and ignore the diversity of imaging styles and the temporal continuity of video, affecting the accuracy of clinical diagnosis and the safety of surgery.
A bidirectional transformation branch network framework is adopted, which combines an inter-frame compensation module and a text-guided style representation module. Inter-frame compensation features are generated through deformable alignment and dual attention mechanism, and style-guided features are generated using cross attention mechanism to achieve temporal compensation and style adaptation enhancement of images.
It improves the temporal consistency and imaging style adaptability of medical videos, achieving more natural, smooth, and high-quality medical video enhancement effects to meet the needs of clinical diagnosis and surgery.
Smart Images

Figure CN120878151B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of image processing technology and computer-aided medical diagnosis technology, and in particular to a style-guided and time-compensated method, system, device and medium for enhancing low-light medical videos. Background Technology
[0002] Medical examinations provide direct visualization of internal organ structures and lesion morphology, serving as a common tool for assisting clinical diagnosis and surgical treatment. However, in practical clinical applications, due to complex examination environments, convoluted and narrow organ passages, limited operating angles, and unstable lighting conditions, the acquired medical videos often suffer from various distortions such as insufficient brightness, low contrast, severe noise, and temporal flicker, resulting in poor video quality. This not only affects the doctor's identification of lesions but may also interfere with the surgical procedure. Existing research indicates that low-quality medical videos may lead to a high rate of missed diagnoses, seriously impacting the accuracy of clinical diagnosis and the safety of treatment. Therefore, developing an efficient and practical low-light medical video enhancement method is of great significance for ensuring the smooth progress of clinical diagnosis and surgery.
[0003] Traditional low-light video enhancement methods often rely on manually designed image priors and manually adjusted parameters to improve image brightness and contrast. These methods often lack robustness in complex lighting environments, their enhancement effects are easily limited, and they lack good generalization ability, making it difficult to meet the high-quality video requirements of real-world clinical scenarios. With the development of deep learning technology, neural network-based image enhancement methods have gradually become the mainstream solution due to their powerful feature modeling capabilities. However, most current mainstream low-light video enhancement methods focus on image processing in natural scenes, and training typically relies on strictly paired low-light and normal-light videos as supervision signals. But in medical scenarios, obtaining paired video samples is extremely challenging. Therefore, some research has begun to shift towards unpaired supervised learning mechanisms, proposing style transfer enhancement under conditions without paired data. Nevertheless, these methods often overlook the impact of the diversity of imaging styles and the temporal continuity of video on clinical usability in medical scenarios. Summary of the Invention
[0004] To address the aforementioned issues, this application provides a style-guided and time-compensated method, system, device, and medium for enhancing low-light medical videos, thereby resolving the difficulty in obtaining medical matching data.
[0005] According to the first aspect of this application, a style-guided and temporal-compensated method for enhancing low-light medical videos is provided, the method comprising:
[0006] A bidirectional conversion branch network framework is constructed, comprising two generators and two discriminators. The two generators are a first generator and a second generator, and the two discriminators are a first discriminator and a second discriminator. The first generator is used to convert continuous video frames in the low-light domain into continuous video frames in the normal-light domain, and the second generator is used to convert continuous video frames in the normal-light domain into continuous video frames in the low-light domain. The first discriminator and the second discriminator are used to determine the authenticity of the output images of the first generator and the second generator, respectively.
[0007] The two generators respond to the input image data and process it through three parallel branches. These three parallel branches include branch one, branch two, and branch three. Branch one inputs the supporting frame and its adjacent reference frames into the inter-frame compensation module, generating inter-frame compensation features through deformable alignment and a dual attention mechanism. Branch two extracts the spatial features of the reference frames and fuses them with the inter-frame compensation features to obtain fused features. Branch three inputs a preset text style description and fused features into the text-guided style representation module, generating style-guided features through a cross-attention mechanism. The supporting frame and its adjacent reference frames originate from the input image data.
[0008] The inter-frame compensation features, fusion features, and style guidance features are fused together and then decoded to generate continuous video frames in either the normal illumination domain or the low illumination domain.
[0009] Furthermore, the image data input to the first generator includes directly input and / or low-light domain continuous video frames generated by the second generator, and the image data input to the second generator includes directly input and / or normal-light domain continuous video frames generated by the first generator.
[0010] Furthermore, branch one will support the input of the frame and its adjacent reference frames into the inter-frame compensation module, and the method for generating inter-frame compensation features through deformable alignment and dual attention mechanisms includes:
[0011] Obtain the first reference frame Second reference frame and support frame F t L The corresponding spatial feature representations are generated through a feature extraction network, which are respectively the features of the first reference frame. Second reference frame features and support frame features
[0012] Employing a deformable modeling strategy for supporting frame features Perform spatial alignment operations to generate features consistent with the first reference frame. Second reference frame features Spatially consistent first alignment feature Second alignment feature
[0013] Calculate the first alignment feature respectively Second alignment feature With the first reference feature Second reference feature The difference information is used to obtain the first difference feature and the second difference feature;
[0014] The first and second difference features are spatially downsampled and compressed, and their channel dimensions are expanded, respectively, and then fused into a unified temporal representation.
[0015] Joint attention weights are generated through a dual-guided approach combining spatial and channel attention mechanisms, and then weighted for processing. Obtain the fused features
[0016] Will Input residual learning units, and output inter-frame compensation features after pooling and structuring.
[0017] Furthermore, the third branch inputs the preset text style description and fusion features into the text guidance style representation module, and generates style guidance features through a cross-attention mechanism in the following ways:
[0018] Obtain preset text style descriptions and fusion features;
[0019] The preset text style description is encoded using a text encoder to extract text features represented by style vectors.
[0020] Linear mapping is performed on the fused features and the text features respectively to generate a query, key, and value matrix for the cross-attention mechanism, and the fused cross-attention features are obtained through the attention mechanism.
[0021] The cross-attention features are transformed using layer normalization and a feedforward fully connected network to generate an enhanced attention feature representation.
[0022] Represent the attention features The feature weight map is mapped by the Sigmoid function and then multiplied channel by channel with the fused features to obtain the weighted result.
[0023] The weighted results are processed by convolutional layers, and residual connections are combined to output image feature representations with style preservation capabilities.
[0024]
[0025] Furthermore, the method also includes: optimizing and training the bidirectional transformation branch network framework based on a set loss function.
[0026] Furthermore, the loss function includes transmission loss, brightness loss, and structural loss.
[0027] Furthermore, the transmission loss consists of adversarial loss, cycle consistency loss, and identity mapping loss.
[0028] According to the second technical solution of this application, a style-guided and time-compensated low-light medical video enhancement system is provided, the system comprising:
[0029] The framework building module is configured to build a bidirectional conversion branch network framework, including two generators and two discriminators. The two generators are a first generator and a second generator, and the two discriminators are a first discriminator and a second discriminator. The first generator is used to convert continuous video frames in the low-light domain into continuous video frames in the normal-light domain, and the second generator is used to convert continuous video frames in the normal-light domain into continuous video frames in the low-light domain. The first discriminator and the second discriminator are used to determine the authenticity of the output images of the first generator and the second generator, respectively.
[0030] A multi-branch parallel processing module is configured to input image data into a generator and process it through three parallel branches. These three parallel branches include branch one, branch two, and branch three. Branch one inputs a supporting frame and its adjacent reference frames into an inter-frame compensation module, generating inter-frame compensation features through deformable alignment and a dual attention mechanism. Branch two extracts spatial features from the reference frames and fuses them with the inter-frame compensation features to obtain fused features. Branch three inputs a preset text style description into a text-guided style representation module, generating style-guided features through a cross-attention mechanism. The supporting frame and its adjacent reference frames originate from the input image data.
[0031] The feature fusion module is configured to fuse the inter-frame compensation features, fusion features, and style guidance features, and generate continuous video frames in the normal illumination domain or the low illumination domain via a decoder.
[0032] According to the third technical solution of this application, an electronic device is provided, the electronic device comprising: a memory for storing a computer program; and a processor for executing the computer program to implement the method described above.
[0033] According to the fourth technical solution of this application, a non-transitory computer-readable storage medium storing instructions is provided, which, when executed by a processor, performs the method described above.
[0034] The style-guided and timing-compensated low-light medical video enhancement methods, systems, devices, and media according to the various schemes of this application have at least the following technical effects:
[0035] To address the issues of insufficient temporal consistency and monotonous imaging style in video enhancement, this application designs an inter-frame compensation module and a text-guided style representation module, which are used to improve the temporal consistency of enhanced videos and adapt to various imaging styles, thereby achieving a more natural, smooth, and style-adaptive medical video enhancement effect.
[0036] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0037] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:
[0038] Figure 1 A flowchart illustrating a style-guided and timing-compensated low-light medical video enhancement method provided in this application embodiment;
[0039] Figure 2 This is a structural diagram of the bidirectional conversion branch network framework provided in the embodiments of this application;
[0040] Figure 3 A flowchart illustrating the inter-frame compensation feature generation process provided in this application embodiment;
[0041] Figure 4 A schematic diagram of the inter-frame compensation module provided in an embodiment of this application;
[0042] Figure 5 A flowchart for generating style-guided features provided in this application embodiment;
[0043] Figure 6 A schematic diagram of the text guidance style representation module provided in an embodiment of this application;
[0044] Figure 7 This is a structural diagram of a style-guided and timing-compensated low-light medical video enhancement system provided in an embodiment of this application. Detailed Implementation
[0045] To enable those skilled in the art to better understand the technical solution of this application, the application will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0046] This application provides a style-guided and timing-compensated method for enhancing low-light medical videos. Figure 1This is a flowchart illustrating a style-guided and timing-compensated low-light medical video enhancement method provided in an embodiment of this application. Figure 1 As shown, the method may include the following steps S10 to S30.
[0047] S10, construct a bidirectional conversion branch network framework, including two generators and two discriminators; wherein, the two generators are a first generator and a second generator, and the two discriminators are a first discriminator and a second discriminator. The first generator is used to convert continuous video frames in the low-light domain into continuous video frames in the normal-light domain, and the second generator is used to convert continuous video frames in the normal-light domain into continuous video frames in the low-light domain. The first discriminator and the second discriminator are used to determine the authenticity of the output images of the first generator and the second generator, respectively.
[0048] S20, the two generators respond to the input image data and process it through three parallel branches; the three parallel branches include branch one, branch two, and branch three. Branch one inputs the supporting frame and its adjacent reference frames into the inter-frame compensation module, and generates inter-frame compensation features through deformable alignment and dual attention mechanism; branch two extracts the spatial features of the reference frames and fuses them with the inter-frame compensation features to obtain fused features; branch three inputs the preset text style description and fused features into the text-guided style representation module, and generates style-guided features through cross-attention mechanism; wherein, the supporting frame and its adjacent reference frames come from the input image data;
[0049] S30 fuses inter-frame compensation features, fusion features, and style guidance features, and generates continuous video frames in the normal illumination domain or the low illumination domain through the decoder.
[0050] In some embodiments, such as Figure 2 The diagram shown illustrates the structure of a bidirectional transformation branch network framework provided in this embodiment. This framework includes two generators and two discriminators. The two generators are a first generator and a second generator, and the two discriminators are a first discriminator and a second discriminator. The two generators are used to implement image style transfer between the low-light domain and the normal-light domain, respectively; the two discriminators are used to determine the realism of the generated images, thereby constructing an end-to-end adversarial enhancement framework. In this bidirectional structure, the first generation module receives consecutive video frames from the low-light video domain. As input, the corresponding normal lighting image is generated; the second generation module then uses continuous video frames in the normal lighting domain. The system generates corresponding low-light style video frames as input. Through the bidirectional mapping process described above, the system not only achieves enhancement but also effectively constrains the style consistency of the generated results. The generated video frames are then input into the corresponding discrimination modules to determine whether the image originates from the real data distribution. The entire generation-discrimination process employs multiple loss functions for joint constraints to ensure comprehensive optimization of the output video frames in terms of sharpness, stability, and style consistency. In the specific design of the generator, a three-way parallel input structure is used to fully integrate information from different modalities. The first branch receives consecutive frames of low-light video and extracts temporal compensation features between the support frames and reference frames through an inter-frame compensation mechanism. The second branch processes the reference frame image, extracts its spatial features, and fuses them with the aforementioned compensation features. The third branch inputs text information containing style cues, extracts semantic features through a text encoding module, and uses a style guidance module to adaptively model the relationship between text and image features to extract high-level representations related to the image imaging style. The fused features are then processed by the decoder to generate an enhanced image, which is then sent to the discrimination module for authenticity determination. The decoder section achieves image restoration through a multi-layer upsampling structure, while the discrimination module consists of a multi-layer convolutional structure, combined with normalization and non-linear activation mechanisms to enhance discrimination capabilities, and finally outputs the discrimination result.
[0051] In some embodiments, such as Figure 3 As shown, branch one inputs the support frame and its adjacent reference frames into the inter-frame compensation module, and generates inter-frame compensation features through deformable alignment and dual attention mechanisms, which is achieved through the following steps:
[0052] S301: Obtain the first reference frame Second reference frame and support frame F t L The corresponding spatial feature representations are generated through a feature extraction network, which are respectively the features of the first reference frame. Second reference frame features and support frame features
[0053] S302: Employing a deformable modeling strategy for supporting frame features Perform spatial alignment operations to generate features consistent with the first reference frame. Second reference frame features Spatially consistent first alignment feature Second alignment feature
[0054] S303: Calculate the first alignment feature respectively Second alignment feature With the first reference feature Second reference feature The difference information is used to obtain the first difference feature and the second difference feature;
[0055] S304: Spatial downsampling compression and channel dimension expansion are performed on the first and second difference features respectively, and then they are fused into a unified time-series representation.
[0056] S305: Joint attention weights are generated through dual guidance of spatial attention and channel attention mechanisms, and then weighted processing is performed. Obtain the fused features
[0057] S306: Merged features Input residual learning units, and output inter-frame compensation features after pooling and structuring.
[0058] In some embodiments, such as Figure 4 As shown, to improve the temporal consistency and natural smoothness of the video during the enhancement process, this embodiment proposes an inter-frame compensation module (TC). This module dynamically models the motion information between frames based on the temporal relationship between the reference frame and its adjacent frames, thereby achieving effective alignment and compensation for supporting frames. Specifically, the TC module uses the reference frame (… and ) and support frame F t L As input, spatial feature representations of each frame are first obtained through a feature extraction network. and Considering the potential for significant displacement or non-rigid motion between video frames, this module introduces a deformable modeling strategy to better achieve inter-frame alignment. This strategy aligns the features of supporting frames, obtaining alignment features consistent with the reference frame space. and Subsequently, the TC module extracts differential information representing motion changes from the alignment and reference features by constructing a temporal difference operation. This difference feature is further spatially compressed through downsampling and expanded in the channel dimension to enhance its expressive power. The difference features at different time points are then uniformly fused and reshaped into a unified temporal fusion representation. To further highlight keyframe regions and suppress redundant information, the module introduces a dual guidance strategy based on spatial attention and channel attention mechanisms. The spatial attention mechanism enhances spatially salient regions, while the channel attention mechanism strengthens the expressive power of semantic channels. The attention weights output by both mechanisms are jointly fused to improve the discriminativeness of the final compensated features. Finally, the fused features... The residual learning units, after pooling and structuring, are further processed to generate the final inter-frame compensated feature representation.
[0059] In some embodiments, such as Figure 5 As shown, the third branch inputs the preset text style description and fusion features into the text-guided style representation module, and generates style-guided features through a cross-attention mechanism, which is achieved through the following steps;
[0060] S501: Obtain preset text style descriptions and fusion features;
[0061] S502: Encode the preset text style description using a text encoder to extract text features represented by style vectors.
[0062] S503: Perform linear mapping on the fused features and the text features respectively to generate a query, key, and value matrix for the cross-attention mechanism, and obtain the fused cross-attention features through the attention mechanism;
[0063] S504: The cross-attention features are transformed through layer normalization and a feedforward fully connected network to generate an enhanced attention feature representation.
[0064] S505: Represent the attention features The feature weight map is mapped by the Sigmoid function and then multiplied channel by channel with the fused features to obtain the weighted result.
[0065] S506: The weighted results are processed through convolutional layers and combined with residual connections to output image feature representations with style preservation capabilities.
[0066] In some embodiments, given the existence of colonoscopy devices from various manufacturers (such as Olympus and Fuji) in current clinical settings, with significant differences in their imaging styles, and the fact that physicians often have preferences for specific imaging styles, when enhancing endoscopic videos under low-light conditions, it is necessary to maintain the stylistic characteristics of the original image in addition to improving image quality. This embodiment proposes a text-guided style representation module (SG) to achieve controllable preservation of the imaging style in the enhancement result. For example... Figure 6 As shown, the style representation module achieves explicit modeling and guidance of imaging style by fusing medical text prompts with image fusion features. Specifically, a text encoder is first used to encode the preset style description statements, extracting the text features represented by style vectors. To preserve spatial structure information, the SG module performs linear mapping on the fused features and the reshaped text features, generating query (Q), key (K), and value (V) matrices for the cross-attention mechanism. This is then transformed into fused cross-attention features through the attention mechanism. To improve feature stability, these cross-features are further transformed using layer normalization and a feedforward fully connected network to generate enhanced attention feature representations. Next, the attention feature is mapped to a feature weight map using a sigmoid function and then multiplied channel-wise with the original fused image features. This enhances image region features highly relevant to textual semantics while suppressing style-irrelevant interference. Subsequently, this module further processes the above results through convolutional units and employs a residual connection mechanism to improve its stability and feature transfer capability, ultimately outputting an image feature representation with style preservation capabilities.
[0067] In some embodiments, the bidirectional conversion branch network framework is optimized and trained based on a set loss function. For example, multiple loss functions are used for joint optimization, including a transmission loss consisting of adversarial loss, cycle consistency loss, and identity mapping loss, as well as brightness loss and structure loss.
[0068] The following section will introduce the dataset sources and partitioning ratios during the training phase, as well as the training environment, optimizer settings, batch size, and epoch size.
[0069] This embodiment uses the collected medical video dataset LLCVD, which contains a total of 300 colonoscopy video samples acquired under white light conditions, including 200 samples under low light conditions and 100 samples under normal light conditions. For experimental convenience, in the network training experiment, all video images were uniformly adjusted to 256×256 resolution. 160 low-light videos and 80 normal-light videos were used for model training, and the remaining 40 low-light videos and 20 normal-light videos were used for testing.
[0070] The network model in this embodiment is implemented using the PyTorch deep learning framework and runs on the Ubuntu 18.04 operating system. The network model uses the Adam optimizer with a learning rate of 0.0001, a batch size of 4 training data, and 200 epochs.
[0071] This embodiment uses BRISQUE, a traditional no-reference image quality assessment model based on natural scene statistical theory, which is widely accepted in the field of medical image and video enhancement, to measure network performance. Generally, a smaller BRISQUE value indicates better performance.
[0072] Experiments were conducted using the methods described above. The results show that the proposed method achieves a BRISQUE coefficient of 27.0311 on the LLCCVD dataset, demonstrating the feasibility and advancement of the proposed method.
[0073] Another aspect of this application provides a style-guided and timing-compensated low-light medical video enhancement system, such as... Figure 7 The diagram shown is a structural diagram of a style-guided and timing-compensated low-light medical video enhancement system provided in an embodiment of this application. The style-guided and timing-compensated low-light medical video enhancement system includes:
[0074] The framework construction module 701 is configured to construct a bidirectional conversion branch network framework, including two generators and two discriminators. The two generators are a first generator and a second generator, and the two discriminators are a first discriminator and a second discriminator. The first generator is used to convert continuous video frames in the low-light domain into continuous video frames in the normal-light domain, and the second generator is used to convert continuous video frames in the normal-light domain into continuous video frames in the low-light domain. The first discriminator and the second discriminator are used to determine the authenticity of the output images of the first generator and the second generator, respectively.
[0075] A multi-branch parallel processing module 702 is configured to input image data into a generator and process it through three parallel branches. The three parallel branches include branch one, branch two, and branch three. Branch one inputs a support frame and its adjacent reference frames into an inter-frame compensation module, generating inter-frame compensation features through deformable alignment and a dual attention mechanism. Branch two extracts spatial features from the reference frames and fuses them with the inter-frame compensation features to obtain fused features. Branch three inputs a preset text style description into a text-guided style representation module, generating style-guided features through a cross-attention mechanism. The support frame and its adjacent reference frames originate from the input image data.
[0076] The feature fusion module 703 is configured to fuse the inter-frame compensation features, fusion features and style guidance features, and generate continuous video frames in the normal illumination domain or the low illumination domain via a decoder.
[0077] In some embodiments, the image data input to the first generator includes directly input and / or low-light domain continuous video frames generated by the second generator, and the image data input to the second generator includes directly input and / or normal-light domain continuous video frames generated by the first generator.
[0078] In some embodiments, the multi-branch parallel processing module is further configured to:
[0079] Obtain the first reference frame Second reference frame and support frame F t L The corresponding spatial feature representations are generated through a feature extraction network, which are respectively the features of the first reference frame. Second reference frame features and support frame features
[0080] Employing a deformable modeling strategy for supporting frame features Perform spatial alignment operations to generate features consistent with the first reference frame. Second reference frame features Spatially consistent first alignment feature Second alignment feature
[0081] Calculate the first alignment feature respectively Second alignment feature With the first reference feature Second reference feature The difference information is used to obtain the first difference feature and the second difference feature;
[0082] The first and second difference features are spatially downsampled and compressed, and their channel dimensions are expanded, respectively, and then fused into a unified temporal representation.
[0083] Joint attention weights are generated through a dual-guided approach combining spatial and channel attention mechanisms, and then weighted for processing. Obtain the fused features
[0084] Will Input residual learning units, and output inter-frame compensation features after pooling and structuring.
[0085] In some embodiments, the multi-branch parallel processing module is further configured to;
[0086] Obtain preset text style descriptions and fusion features;
[0087] The preset text style description is encoded using a text encoder to extract text features represented by style vectors.
[0088] Linear mapping is performed on the fused features and the text features respectively to generate a query, key, and value matrix for the cross-attention mechanism, and the fused cross-attention features are obtained through the attention mechanism.
[0089] The cross-attention features are transformed using layer normalization and a feedforward fully connected network to generate an enhanced attention feature representation.
[0090] Represent the attention features The feature weight map is mapped by the Sigmoid function and then multiplied channel by channel with the fused features to obtain the weighted result.
[0091] The weighted results are processed by convolutional layers, and residual connections are combined to output image feature representations with style preservation capabilities.
[0092]
[0093] In some embodiments, the system further includes a training module configured to perform optimized training on the bidirectional transformation branch network framework based on a set loss function.
[0094] In some embodiments, the loss function includes transmission loss, brightness loss, and structural loss.
[0095] In some embodiments, the transmission loss consists of adversarial loss, cycle consistency loss, and identity mapping loss.
[0096] It should be noted that the style-guided and timing-compensated low-light medical video enhancement device provided in the above embodiments and the style-guided and timing-compensated low-light medical video enhancement method provided in the aforementioned embodiments belong to the same concept. The specific way in which each module and unit performs operations has been described in detail in the method embodiments, and will not be repeated here.
[0097] Another aspect of this application provides an electronic device, including: a controller; and a memory for storing one or more programs, which, when executed by the controller, perform the methods described in the various embodiments above.
[0098] Another aspect of this application provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the method as described above. This computer-readable storage medium may be included in the electronic device described in the above embodiments, or it may exist independently and not assembled into the electronic device.
[0099] Another aspect of this application provides a computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various embodiments described above.
[0100] According to one aspect of the embodiments of this application, a computer system is also provided, including a Central Processing Unit (CPU), which can perform various appropriate actions and processes based on a program stored in read-only memory (ROM) or a program loaded from storage into random access memory (RAM), such as performing the methods described above. Various programs and data required for system operation are also stored in the RAM. The CPU, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.
[0101] For example, a computer system includes a Central Processing Unit (CPU), which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) or loaded from storage into random access memory (RAM), such as executing the methods described in the above embodiments. The RAM also stores various programs and data required for system operation. The CPU, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.
[0102] The following components are connected to the I / O interface: input sections including keyboards, mice, etc.; output sections including cathode ray tubes (CRTs), liquid crystal displays (LCDs), and speakers; storage sections including hard drives; and communication sections including network interface cards such as LAN (Local Area Network) cards and modems. The communication sections perform communication processing via networks such as the Internet. Drives are also connected to the I / O interface as needed. Removable media, such as disks, optical discs, magneto-optical discs, semiconductor memories, etc., are installed on the drive as needed so that computer programs read from them can be installed into the storage section as required.
[0103] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program including a computer program for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication component, and / or installed from a removable medium. When the computer program is executed by a central processing unit (CPU), it performs various functions defined in the system of this application.
[0104] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying a computer-readable computer program. The transmitted data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.
[0105] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0106] The module units described in the embodiments of this application can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.
[0107] The above embodiments are only used to illustrate this application and are not intended to limit this application. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of this application. Therefore, all equivalent technical solutions also fall within the scope of this application, and the patent protection scope of this application should be defined by the claims.
Claims
1. A style-guided and temporal-compensated method for enhancing low-light medical videos, characterized in that, The method includes: A bidirectional conversion branch network framework is constructed, comprising two generators and two discriminators. The two generators are a first generator and a second generator, and the two discriminators are a first discriminator and a second discriminator. The first generator is used to convert continuous video frames in the low-light domain into continuous video frames in the normal-light domain, and the second generator is used to convert continuous video frames in the normal-light domain into continuous video frames in the low-light domain. The first discriminator and the second discriminator are used to determine the authenticity of the output images of the first generator and the second generator, respectively. The two generators respond to the input image data and process it through three parallel branches. These three parallel branches include branch one, branch two, and branch three. Branch one inputs the supporting frame and its adjacent reference frames into the inter-frame compensation module, generating inter-frame compensation features through deformable alignment and a dual attention mechanism. Branch two extracts the spatial features of the reference frames and fuses them with the inter-frame compensation features to obtain fused features. Branch three inputs a preset text style description and fused features into the text-guided style representation module, generating style-guided features through a cross-attention mechanism. The supporting frame and its adjacent reference frames originate from the input image data. The inter-frame compensation features, fusion features, and style guidance features are fused together and then decoded to generate continuous video frames in either the normal illumination domain or the low illumination domain.
2. The method according to claim 1, characterized in that, The image data input to the first generator includes directly input and / or low-light domain continuous video frames generated by the second generator, and the image data input to the second generator includes directly input and / or normal-light domain continuous video frames generated by the first generator.
3. The method according to claim 1, characterized in that, The first branch will support the input of the frame and its adjacent reference frames into the inter-frame compensation module, and the methods for generating inter-frame compensation features through deformable alignment and dual attention mechanisms include: Obtain the first reference frame Second reference frame and support frames The corresponding spatial feature representations are generated through a feature extraction network, which are respectively the features of the first reference frame. Second reference frame features and support frame features Employing a deformable modeling strategy for supporting frame features Perform spatial alignment operations to generate features consistent with the first reference frame. Second reference frame features Spatially consistent first alignment feature Second alignment feature Calculate the first alignment feature respectively Second alignment feature With the first reference feature Second reference feature The difference information is used to obtain the first difference feature and the second difference feature; The first and second difference features are spatially downsampled and compressed, and their channel dimensions are expanded, respectively, and then fused into a unified temporal representation. Joint attention weights are generated through a dual-guided approach combining spatial and channel attention mechanisms, and then weighted for processing. Obtain the fused features Will Input residual learning units, and output inter-frame compensation features after pooling and structuring.
4. The method according to claim 1, characterized in that, The third branch inputs the preset text style description and fusion features into the text guidance style representation module, and generates style guidance features through a cross-attention mechanism in the following ways: Obtain preset text style descriptions and fusion features; The preset text style description is encoded using a text encoder to extract text features represented by style vectors. Linear mapping is performed on the fused features and the text features respectively to generate a query, key, and value matrix for the cross-attention mechanism, and the fused cross-attention features are obtained through the attention mechanism. The cross-attention features are transformed using layer normalization and a feedforward fully connected network to generate an enhanced attention feature representation. Represent the attention features The feature weight map is mapped by the Sigmoid function and then multiplied channel by channel with the fused features to obtain the weighted result. The weighted results are processed by convolutional layers, and residual connections are combined to output image feature representations with style preservation capabilities.
5. The method according to any one of claims 1 to 4, characterized in that, The method further includes: optimizing and training the bidirectional transformation branch network framework based on a set loss function.
6. The method according to claim 5, characterized in that, The loss function includes transmission loss, brightness loss, and structural loss.
7. The method according to claim 6, characterized in that, The transmission loss consists of adversarial loss, cycle consistency loss, and identity mapping loss.
8. A style-guided and time-compensated low-light medical video enhancement system, characterized in that, The system includes: The framework building module is configured to build a bidirectional conversion branch network framework, including two generators and two discriminators. The two generators are a first generator and a second generator, and the two discriminators are a first discriminator and a second discriminator. The first generator is used to convert continuous video frames in the low-light domain into continuous video frames in the normal-light domain, and the second generator is used to convert continuous video frames in the normal-light domain into continuous video frames in the low-light domain. The first discriminator and the second discriminator are used to determine the authenticity of the output images of the first generator and the second generator, respectively. A multi-branch parallel processing module is configured to input image data into a generator and process it through three parallel branches. These three parallel branches include branch one, branch two, and branch three. Branch one inputs a supporting frame and its adjacent reference frames into an inter-frame compensation module, generating inter-frame compensation features through deformable alignment and a dual attention mechanism. Branch two extracts spatial features from the reference frames and fuses them with the inter-frame compensation features to obtain fused features. Branch three inputs a preset text style description into a text-guided style representation module, generating style-guided features through a cross-attention mechanism. The supporting frame and its adjacent reference frames originate from the input image data. The feature fusion module is configured to fuse the inter-frame compensation features, fusion features, and style guidance features, and generate continuous video frames in the normal illumination domain or the low illumination domain via a decoder.
9. An electronic device, characterized in that, The electronic device includes: Memory, used to store computer programs; A processor for executing the computer program to implement the method as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium storing instructions, characterized in that, When the instructions are executed by the processor, the method according to any one of claims 1 to 7 is performed.