Ship navigation monitoring video super-resolution method and system based on mixed attention
By combining deep learning and hybrid attention mechanisms, using dual cameras to acquire videos and perform adaptive alternating superposition, the problems of low resolution and blurred details in ship navigation videos are solved, and efficient video super-resolution and real-time requirements are achieved.
Patent Information
- Application Number
- CN202510628956.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-16
AI Technical Summary
The prior art has problems of low resolution, high noise and blurred details in ship navigation videos, making it difficult to adapt to complex marine environments and difficult to meet real-time requirements.
Using a deep learning and mixed attention method, by using two cameras with different accuracy to acquire videos, a hybrid attention module is built to perform keyframe extraction of high-precision and low-precision videos, adaptive alternating superimposition, and a super-resolution reconstruction network is constructed for adversarial training to generate super-resolution video sequences.
It significantly improves the super-resolution quality of ship video, enhances the detailed characteristics of ship targets, adapts to complex marine environments, and achieves real-time requirements.
Smart Images

Figure CN120147138A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video processing, and more specifically, to a ship navigation monitoring video super-resolution method and system based on hybrid attention. Background Art
[0002] Ship navigation monitoring systems play a crucial role in fields such as ocean transportation, port management, and maritime search and rescue. With the rapid development of computer vision and video surveillance technologies, video-based ship navigation monitoring has become an important means of modern maritime supervision. However, in practical applications, due to factors such as camera hardware performance, transmission bandwidth, and environmental interference (such as fog, waves, and low light), the collected ship navigation videos often have problems such as low resolution, high noise, and blurred details, seriously affecting the accuracy of subsequent tasks such as target detection, tracking, and behavior analysis.
[0003] Traditional video super-resolution methods are mainly based on interpolation or reconstruction-based methods, but these methods usually rely on manually designed prior knowledge and are difficult to adapt to complex marine environments. In recent years, deep learning-based super-resolution technologies have made significant progress, especially the introduction of convolutional neural networks (CNNs) and generative adversarial networks (GANs), which have significantly improved the super-resolution performance of images and videos. Existing methods have limitations in ship navigation videos, such as being affected by motion blur and complex backgrounds and having difficulty meeting real-time requirements. In recent years, the attention mechanism has been introduced into the super-resolution task to improve the reconstruction performance by focusing on key regions. Hybrid attention combines mechanisms such as channel attention, spatial attention, and temporal attention, and can better model the spatio-temporal dependencies and multi-scale features in videos. Therefore, how to combine deep learning and the hybrid attention mechanism to improve video resolution is an urgent problem to be solved currently. Summary of the Invention
[0004] To solve the above technical problems, the present invention proposes a ship navigation monitoring video super-resolution method and system based on hybrid attention, which combines deep learning and the hybrid attention mechanism to improve video resolution while enhancing the detailed features of ship targets and adapting to complex marine environments.
[0005] The first aspect of the present invention provides a ship navigation monitoring video super-resolution method based on hybrid attention, including the following steps: Collect ship navigation video sequences using two cameras with different precisions and perform preprocessing to obtain two preprocessed ship navigation video sequences; Construct hybrid attention modules for the two ship navigation video sequences respectively, and use the hybrid attention modules to extract key frames from high-precision videos and low-precision videos to obtain key frames of high-precision videos and low-precision videos; Adaptive alternately stack the key frames of the high-precision video and the low-precision video to generate a stacked video sequence; Construct a super-resolution reconstruction network, use the stacked video sequence as the model input, enhance the stacked video sequence, and output a super-resolution video sequence through adversarial training.
[0006] In this solution, two cameras with different precisions are used to collect the ship navigation video sequence and perform preprocessing. Specifically: Configure a high-precision camera and a low-precision camera, perform hardware synchronization on the high-precision camera and the low-precision camera, and monitor and align based on timestamps to obtain two-way ship navigation video sequences respectively; Use the SIFT feature detection algorithm in the two-way ship navigation video sequence to perform scale-space extreme value detection, locate the key points of the video frame, calculate the main direction of the key points, generate SIFT feature descriptors, and use the ORB feature detection algorithm to quickly locate the key points and generate BRIEF feature descriptors; Based on the SIFT feature descriptors and BRIEF feature descriptors, a feature point set of the high-precision video frame and the low-precision video frame is generated, and the feature point set is respectively SIFT-matched through the Euclidean distance and ORB-matched through the Hamming distance; Obtain the key point matching result and eliminate the wrong matches, optimize the key point matching pairs to achieve spatial alignment, and eliminate the perspective difference.
[0007] In this solution, a hybrid attention module is constructed for the two-way ship navigation video sequence respectively, and the hybrid attention module of the high-precision video sequence is used to extract frames from the high-precision video. Specifically: For the high-precision video sequence, a hybrid attention module is constructed using spatial attention and semantic attention. The ship navigation video sequence collected by the preprocessed high-precision camera is used as the module input. In the spatial attention branch of the hybrid attention module, the Laplacian gradient function is used for clarity evaluation, and the spatial attention is configured according to the local clarity evaluation result of the video frame; In the semantic attention branch of the hybrid attention module, an encoder based on MobileNetV2 is used to encode the video frame into the embedding space to obtain single-frame visual features. A pre-trained classification network is constructed using the single-frame visual features, and class activation mapping is used to extract the visual semantics corresponding to the video frame; For the task objective of obtaining ship navigation videos, class activation mapping uses global average pooling to obtain the relationship between single-frame visual features and task objective categories, projects the single-frame visual features using the weights of the output layer of the classification network to obtain a class activation map with task objective category labels, activates the class activation map using the Softmax function, and configures the semantic attention of video frames according to the obtained probabilities; Combine the spatial attention and semantic attention of video frames, and filter out video frames that meet the preset weight threshold according to the combined attention weights as key frames of the ship navigation video sequence collected by the high-precision camera for output.
[0008] In this solution, a hybrid attention module of a low-precision video sequence is used for low-precision video frame extraction, specifically: For a low-precision video sequence, a hybrid attention module is constructed using temporal attention and motion attention. The preprocessed ship navigation video sequence collected by the low-precision camera is used as the module input, and the optical flow vector of the task objective in the video frame is calculated in the hybrid attention module; In the temporal attention branch, the optical flow vectors of each video frame are encoded and represented, and a GRU unit is used for temporal modeling to capture the long-term dependence of the optical flow vector sequence, and temporal attention weights are generated according to the hidden state of the GRU unit; In the motion attention branch, calculate the amplitude of each optical flow vector, obtain the average motion intensity of the video frame according to the amplitude of the optical flow vector, and obtain the main direction of the optical flow vector. Calculate the angle between the motion directions of all optical flow vectors in the video frame and the main direction for motion consistency evaluation to obtain the motion direction consistency evaluation result; Generate the motion attention weights of the video frame through the average motion intensity of the video frame and the motion direction consistency evaluation result, combine the temporal attention weights and motion attention weights of the video frame, and filter out video frames that meet the preset weight threshold as key frames of the ship navigation video sequence collected by the low-precision camera for output.
[0009] In this solution, the key frames of the high-precision video and the low-precision video are adaptively and alternately stacked to generate a stacked video sequence, specifically: Obtain the key frames of the high-precision video sequence and the low-precision video sequence, use the motion offset between the key frames of the two video sequences for motion compensation alignment, and obtain the absolute value of the motion feature difference between the high-precision video key frame and the low-precision video key frame; When the absolute value of the motion feature difference is greater than the preset threshold, select the key frame with a larger motion feature value. Otherwise, select the high-precision video key frame and the low-precision video key frame according to the odd-even alternating strategy; Obtain the superimposed video sequence, establish an inter-frame structural similarity metric for monitoring, and perform visual consistency checks. When the preset visual consistency standard is met, the superimposed video sequence is output.
[0010] In this solution, a super-resolution reconstruction network is constructed. The superimposed video sequence is used as the model input, and the superimposed video sequence is enhanced to output a super-resolution video sequence. Specifically: Group the features of the superimposed video sequence according to different time steps. In each group, select the target frame according to the motion feature value. Use a 3D dense block with residual connections in each group to obtain spatio-temporal features, and integrate the spatio-temporal features within the group to obtain group-level features; Convolve the group-level features of each group to obtain a single-channel feature map. Connect the single-channel feature maps along the time axis, introduce a self-attention mechanism, calculate the self-attention weights for each position across channels using softmax, and use the self-attention weights to perform element-wise multiplication of the group-level features of different groups at the same position; Import the result of multiplying the cascaded component features by the self-attention weights into a 3D dense block to synthesize the local features of different groups. Add a 2D dense block on top of the 3D dense block to further fuse the features and generate an aggregated feature map; Use channel attention, spatial attention, and temporal attention to perform hybrid attention enhancement on the aggregated feature map. Use a feature pyramid structure to fuse the features of the enhanced aggregated feature map. Upsample the fused features to obtain a high-resolution image consistent with the video frame scale, and connect them along the time axis to obtain a super-resolution video sequence.
[0011] In this solution, adversarial training is introduced into the super-resolution reconstruction network. Specifically: Introduce a discriminator module into the super-resolution reconstruction network. The discriminator module is divided into a spatial discriminator branch and a temporal discriminator branch. Import the obtained high-resolution image into the spatial discriminator for feature discrimination to determine whether the input high-resolution image is a real high-resolution frame or a generated super-resolution frame; Import the obtained super-resolution video sequence into the temporal discriminator for feature discrimination to determine whether the input super-resolution video sequence has real temporal dynamic characteristics; Obtain the discriminant matrices output by the spatial discriminator branch and the temporal discriminator branch. According to the adversarial mechanism, use the discriminant matrices to guide the training of the super-resolution reconstruction network. After iterative training, stop training until the discriminator module can no longer distinguish between the generated super-resolution video sequence and the real super-resolution video sequence, and obtain the trained super-resolution reconstruction network.
[0012] In the second aspect of the present invention, a ship navigation monitoring video super-resolution system based on hybrid attention is provided, including a data acquisition and preprocessing unit, a hybrid attention frame extraction unit, a key frame superposition unit, a super-resolution reconstruction unit, and a post-processing optimization unit; The data acquisition and preprocessing unit uses two cameras with different precisions to collect ship navigation video sequences and perform preprocessing to obtain two preprocessed ship navigation video sequences; The hybrid attention frame extraction unit constructs a hybrid attention module for each of the two ship navigation video sequences, and uses the hybrid attention module to perform high-precision video frame extraction and low-precision video frame extraction to obtain the most informative key frames; The key frame superposition unit aligns the key frames of the high-precision video and the low-precision video, and performs adaptive alternating superposition to generate a superposed video sequence; The super-resolution reconstruction unit constructs a super-resolution reconstruction network, uses the hybrid attention mechanism to enhance the superposed video sequence, and reconstructs the super-resolution video sequence through adversarial training; The post-processing optimization unit performs post-processing optimization on the super-resolution results of the ship navigation video to enhance the edges and reduce inter-frame jitter, and lightweight the super-resolution reconstruction network to achieve real-time super-resolution.
[0013] Compared with the prior art, the beneficial effects of the present invention are as follows: By combining dual-camera data fusion with the hybrid attention mechanism, the present invention significantly improves the quality of ship video super-resolution, and significantly improves the analysis efficiency of ship videos while maintaining real-time performance. By combining deep learning and the hybrid attention mechanism, this method is expected to enhance the detailed features of ship targets while improving the video resolution, and adapt to complex marine environments, providing a high-quality data basis for subsequent ship detection, recognition, and behavior analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the technical solutions in the embodiments or the exemplary of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the exemplary description. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the drawings shown without creative efforts.
[0015] Figure 1 Shows a flowchart of a ship navigation monitoring video super-resolution method based on hybrid attention; Figure 2 Shows a flowchart of alternately superposing the key frames of the high-precision video and the low-precision video; Figure 3Shows the flowchart of constructing a super-resolution reconstruction network for video super-resolution; Figure 4 Shows the block diagram of a ship navigation monitoring video super-resolution system based on hybrid attention. Detailed implementation manners
[0016] In order to more clearly understand the above objects, features and advantages of the present invention, the present invention will be further described in detail below with reference to the drawings and specific implementation manners. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other.
[0017] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited by the specific embodiments disclosed below.
[0018] Figure 1 Shows the flowchart of a ship navigation monitoring video super-resolution method based on hybrid attention.
[0019] As Figure 1 shown, in the first embodiment of the present invention, a ship navigation monitoring video super-resolution method based on hybrid attention is provided, including: S102, using two cameras with different precisions to collect ship navigation video sequences and perform preprocessing to obtain two preprocessed ship navigation video sequences; S104, respectively constructing a hybrid attention module for the two ship navigation video sequences, and using the hybrid attention module to perform high-precision video frame extraction and low-precision video frame extraction to obtain key frames of the high-precision video and the low-precision video; S106, adaptively and alternately superimposing the key frames of the high-precision video and the low-precision video to generate a superimposed video sequence; S108, constructing a super-resolution reconstruction network, using the superimposed video sequence as the model input, enhancing the superimposed video sequence, and outputting a super-resolution video sequence through adversarial training.
[0020] It should be noted that a high-precision camera and a low-precision camera are configured. The high-precision camera provides clear but limited detail information, and the low-precision camera provides smoother but blurred motion information. The high-precision camera and the low-precision camera are hardware-synchronized and monitored and aligned based on timestamps to respectively obtain two-channel ship navigation video sequences. In the two-channel ship navigation video sequences, the SIFT feature detection algorithm is used for feature extraction. Scale-space extreme value detection is performed through a Gaussian difference pyramid to eliminate low-contrast and edge response points, locate the key points of the video frame, calculate the main direction of the key points, and generate SIFT feature descriptors. Additionally, the ORB feature detection algorithm is used for feature extraction. The FAST corner detection is used to quickly locate the key points, and BRIEF feature descriptors are generated. Based on the SIFT feature descriptors and the BRIEF feature descriptors, a feature point set of the high-precision video frame and the low-precision video frame is generated. The feature point set is respectively SIFT-matched through the Euclidean distance and ORB-matched through the Hamming distance. In the SIFT matching, the KNN is used to screen the nearest neighbor matching pairs, and in the ORB matching, cross-validation is adopted to improve the matching accuracy. The key point matching result is obtained and false matches are eliminated. The average error of the matching point pairs is calculated, and the key point matching point pairs are optimized to achieve spatial alignment and eliminate the perspective difference. When the error is less than the preset threshold, it is determined that the alignment fails and re-matching is triggered. If the initial alignment fails, the image is downsampled and re-matched. By combining the accuracy of SIFT and the speed of ORB, the performance and efficiency are balanced, and the quality of video preprocessing is improved. Additionally, in the preprocessing process, guided filtering is used to reduce sea wave interference.
[0021] It should be noted that a hybrid attention module is constructed for the two-channel ship navigation video sequences respectively. For the high-precision video sequence, a hybrid attention module is constructed using spatial attention and semantic attention. The preprocessed ship navigation video sequence collected by the high-precision camera is used as the module input. In the spatial attention branch of the hybrid attention module, the Laplacian gradient function is used for clarity evaluation. The spatial attention is configured according to the local clarity evaluation result of the video frame, and high-sharpness video frames are selected according to the spatial attention.
[0022] In the semantic attention branch of the hybrid attention module, an encoder based on MobileNetV2 is used to encode video frames and map them to an embedding space to obtain single-frame visual features. A pre-trained classification network is constructed using the single-frame visual features. Preferably, the pre-trained classification network can be constructed by a lightweight CNN. Class activation mapping is used to extract the visual semantics corresponding to the video frames. Class activation mapping uses a global average pooling (GAP) layer to generate a visual activation mapping by capturing the relationship between the feature map and different classes. The task objective of the ship navigation video (such as a ship target) is obtained. The class activation mapping uses global average pooling to obtain the relationship between the single-frame visual features and the task objective class, projects the single-frame visual features using the weights of the output layer of the classification network, obtains a class activation map with the task objective class label, activates the class activation map using the Softmax function, and configures the semantic attention of the video frames according to the obtained probability, preferentially selecting frames with a high ship occupancy ratio. The spatial attention and semantic attention of the video frames are combined, and the video frames that meet the preset weight threshold are screened according to the combined attention weights and output as key frames of the ship navigation video sequence collected by the high-precision camera to provide high-quality details.
[0023] For the low-precision video sequence, a hybrid attention module is constructed using temporal attention and motion attention. The preprocessed ship navigation video sequence collected by the low-precision camera is used as the module input. In the hybrid attention module, the Pyramidal Lucas-Kanade algorithm is used for optical flow estimation to construct an optical flow field, and the optical flow vectors of the task objective in the video frames are calculated. In the temporal attention branch, the optical flow vectors of each video frame are encoded and represented, and the GRU unit is used for temporal modeling to capture the long-term dependencies of the optical flow vector sequence. The temporal attention weights are generated according to the hidden state of the GRU unit, and the generated attention weights can accurately reflect the key motion regions in the video.
[0024] In the motion attention branch, for each optical flow vector calculate its magnitude , and obtain the average motion intensity of the video frame according to the magnitude of the optical flow vector , is the number of valid optical flow vectors, and obtain the main direction of the optical flow vector . The main direction represents the overall motion direction of the ship. Calculate the angle between the motion direction of all optical flow vectors in the video frame and the main direction, obtain the number of vectors with an angle greater than the preset angle, calculate the ratio of the obtained number of vectors to the valid optical flow vector data for motion consistency evaluation, and obtain the motion direction consistency evaluation result ; The motion attention weights of the video frames are generated through the average motion intensity of the video frames and the motion direction consistency evaluation result , , As an adjustment parameter, when the motion is more intense and the direction is more consistent, the corresponding motion weight is higher. By combining the temporal attention weight and the motion attention weight of the video frames, the video frames that meet the preset weight threshold are selected as the key frames of the ship navigation video sequence collected by the low-precision camera for output. Frames with significant motion (such as ship turning and acceleration) are selected to provide motion compensation information.
[0025] Figure 2 The flowchart of alternately superimposing the key frames of the high-precision video and the low-precision video is shown.
[0026] According to the embodiments of the present invention, the key frames of the high-precision video and the low-precision video are adaptively alternately superimposed to generate a superimposed video sequence. Specifically: S202, obtain the key frames of the high-precision video sequence and the low-precision video sequence, use the motion offset between the key frames of the two video sequences for motion compensation alignment, and obtain the absolute value of the motion feature difference between the high-precision video key frame and the low-precision video key frame; S204, when the absolute value of the motion feature difference is greater than the preset threshold, select the key frame with a larger motion feature value; otherwise, select the high-precision video key frame and the low-precision video key frame according to the odd-even alternating strategy; S206, obtain the superimposed video sequence, establish an inter-frame structural similarity index monitoring for visual consistency check, and when the preset visual consistency standard is met, output the superimposed video sequence.
[0027] It should be noted that the motion offset between the key frames of the high-precision video sequence and the low-precision video sequence is estimated by the optical flow method for warping alignment. The motion scene is determined according to the absolute value threshold of the motion feature difference, and the threshold can be dynamically adjusted, and the threshold is appropriately reduced under low light conditions. When intense motion is detected, the key frame with a larger motion feature value is selected. For example, when the ship suddenly turns, the low-precision camera can better capture the fast motion and automatically select the low-precision video key frame for anti-blur processing. When the ship is sailing smoothly, the high-precision camera provides a clearer picture and the high-precision video key frame is preferentially selected. In the case of no significant motion difference, the high-precision video sequence and the low-precision video sequence are alternately selected in a fixed order, and through the odd-even alternating strategy, it is ensured that the key frames of the high-precision video sequence and the low-precision video sequence strictly alternate.
[0028] Figure 3 The flowchart of constructing a super-resolution reconstruction network for video super-resolution is shown.
[0029] According to an embodiment of the present invention, a super-resolution reconstruction network is constructed. The superimposed video sequence is used as the model input, and the superimposed video sequence is enhanced to output a super-resolution video sequence. Specifically: S302, perform feature grouping on the superimposed video sequence according to different time steps. In each group, select the target frame according to the motion feature value. Obtain spatio-temporal features using 3D dense blocks with residual connections in each group, and integrate the spatio-temporal features within the group to obtain group-level features; S304, perform convolution on the group-level features of each group to obtain a single-channel feature map. Connect the single-channel feature maps along the time axis, introduce a self-attention mechanism to calculate the self-attention weights for each position across channels using softmax, and use the self-attention weights to perform element-wise multiplication on the group-level features of different groups at the same position; S306, import the result of multiplying the cascaded component features by the self-attention weights into a 3D dense block to synthesize the local features of different groups, and add a 2D dense block at the top of the 3D dense block to further fuse the features to generate an aggregated feature map; S308, perform hybrid attention enhancement on the aggregated feature map using channel attention, spatial attention, and temporal attention. Use a feature pyramid structure to perform feature fusion on the enhanced aggregated feature map, and obtain a high-resolution image consistent with the video frame scale through upsampling of the fused features, and connect along the time axis to obtain a super-resolution video sequence.
[0030] It should be noted that when constructing the super-resolution reconstruction network, grouping is performed and group attention is introduced in the reconstruction of the superimposed video sequence to align, fuse, and integrate the video frame information with different degrees of motion. Feature grouping of the superimposed video sequence provides different complementary information for the restoration of the target frame. The attention mechanism is used to adaptively utilize the complementary information to restore the missing details of the target frame and enhance the performance of video super-resolution. After grouping, the temporal features of the group-level features are extracted using the self-attention mechanism to obtain the context information of the group. When a certain area of a group is occluded, complementary information is extracted from the remaining groups to restore the unclear local details, achieving mutual complementation. Each layer of the 3D dense block inputs the feature maps spliced from all previous layers, captures both spatial details and temporal motion at the same time, and realizes multi-scale feature reuse through dense connections. A 2D dense block is added at the top of the 3D dense block to further fuse the features, and the information of each group is deeply integrated to generate an aggregated feature map.
[0031] The aggregated feature map is enhanced with hybrid attention using channel attention, spatial attention, and temporal attention. Channel attention dynamically adjusts the importance weights of feature channels, enhances ship feature channels (such as specific color channels like hull numbers and draft lines), and suppresses interference channels such as wave reflections. Spatial attention focuses on key spatial regions of the image, automatically locks in the areas of interest of the ship, enhances the hull edge regions, and suppresses irrelevant sea surface regions. Temporal attention models the inter-frame motion dependency to enhance motion-significant frames (such as turning and acceleration phases), suppresses the contribution of blurred frames, and improves the continuity of the ship's track. While maintaining real-time processing performance, the hybrid attention mechanism is used to specifically enhance the key parts of the ship (hull numbers, hull markings), and the stable output brought about by spatio-temporal consistency constraints.
[0032] It should be noted that a discriminator module is introduced into the super-resolution reconstruction network. The discriminator module is divided into a spatial discriminator branch and a temporal discriminator branch. The obtained high-resolution image is imported into the spatial discriminator to perform unpooled downsampling through stride convolution for feature discrimination, to determine whether the input high-resolution image is a real high-resolution frame or a generated super-resolution frame; through adversarial training, it forces the superimposed video reconstruction generation network to produce more realistic texture details, especially improving high-frequency content such as ship hull numbers and flags. The obtained super-resolution video sequence is imported into the temporal discriminator, and a 3D convolutional neural network is used for feature discrimination to determine whether the input super-resolution video sequence has real temporal dynamic characteristics; through adversarial training, problems such as inter-frame jitter and inconsistent motion blur in the generated video are identified. The discriminant matrices output by the spatial discriminator branch and the temporal discriminator branch are obtained, and according to the adversarial mechanism, the discriminant matrices are used to guide the training of the super-resolution reconstruction network. After iterative training, the training is terminated until the discriminator module can no longer distinguish between the generated super-resolution video sequence and the real super-resolution video sequence, and the trained super-resolution reconstruction network is obtained.
[0033] It should be noted that the super-resolution video sequence output by the super-resolution reconstruction network is optimized to improve real-time performance. For example, Unsharp Masking is used to enhance the edges, and Kalman filtering is adopted to reduce inter-frame jitter. In addition, knowledge distillation or channel pruning is used for the super-resolution reconstruction network to perform model lightweight processing to reduce the computational load and achieve real-time super-resolution.
[0034] The super-resolution video sequence output by the super-resolution reconstruction network is used to obtain the ship detection area by the YOLOv5 network for ship identification. Based on the ship identification result, the navigation trajectory of the target ship is generated, and the space-time coding is performed on the navigation trajectory. The prediction head performs trajectory prediction based on the space-time coding. The trajectory prediction sequence within a preset time is compared with the historical berthing trajectory sequence of the ship in the port database. The dynamic time warping algorithm is used to obtain the time warping distance between the sequences. According to the time warping distance representing the trajectory similarity, when the trajectory similarity is greater than the preset similarity threshold, the target ship is marked for berthing, and the real-time marking is updated to achieve real-time berthing analysis.
[0035] Figure 4 Fig. shows a block diagram of a ship navigation monitoring video super-resolution system based on hybrid attention.
[0036] The second embodiment of the present invention provides a ship navigation monitoring video super-resolution system 4 based on hybrid attention. The system includes a data acquisition and preprocessing unit 401, a hybrid attention frame extraction unit 402, a key frame superposition unit 403, a super-resolution reconstruction unit 404, and a post-processing optimization unit 405; The data acquisition and preprocessing unit uses two cameras with different precisions to collect ship navigation video sequences and perform preprocessing to obtain two preprocessed ship navigation video sequences; The hybrid attention frame extraction unit constructs a hybrid attention module for each of the two ship navigation video sequences, and uses the hybrid attention module to perform high-precision video frame extraction and low-precision video frame extraction to obtain the most informative key frames; The key frame superposition unit aligns the key frames of the high-precision video and the low-precision video, and performs adaptive alternating superposition to generate a superposed video sequence; The super-resolution reconstruction unit constructs a super-resolution reconstruction network, uses the hybrid attention mechanism to enhance the superposed video sequence, and reconstructs the super-resolution video sequence through adversarial training; The post-processing optimization unit performs post-processing optimization on the super-resolution result of the ship navigation video to enhance the edges and reduce the inter-frame jitter, and lightweight the super-resolution reconstruction network to achieve real-time super-resolution.
[0037] The third embodiment of the present invention provides a computer-readable storage medium, which includes a program for the method of ship navigation monitoring video super-resolution based on hybrid attention. When the program for the method of ship navigation monitoring video super-resolution based on hybrid attention is executed by a processor, the steps of the method of ship navigation monitoring video super-resolution based on hybrid attention are implemented.
[0038] In several embodiments provided in the present application, it should be understood that the disclosed method can be implemented in other ways. The embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed with each other can be through some interfaces, indirect coupling or communication connection of devices or units, and can be electrical, mechanical, or other forms.
[0039] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: removable storage devices, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), magnetic disks, or optical discs and other various media that can store program codes.
[0040] Alternatively, if the above-mentioned integrated unit of the present invention is implemented in the form of a software functional unit and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present invention. And the foregoing storage medium includes: removable storage devices, ROM, RAM, magnetic disks, or optical discs and other various media that can store program codes.
[0041] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention.
Claims
1. A ship navigation monitoring video super-resolution method based on mixed attention, characterized in that: The following steps are involved: Use two cameras with different precisions to collect and preprocess the ship navigation video sequences, and obtain two preprocessed ship navigation video sequences; Constructing a hybrid attention module for two ship navigation video sequences respectively, using the hybrid attention module to extract high-precision video frames and low-precision video frames to obtain key frames of high-precision video and low-precision video; Adaptively and alternately superimpose key frames of high-precision video and low-precision video to generate a superimposed video sequence; A super-resolution reconstruction network is constructed, the superimposed video sequence is used as a model input, the superimposed video sequence is enhanced, and a super-resolution video sequence is output through adversarial training.
2. The ship navigation monitoring video super-resolution method based on mixed attention according to claim 1 is characterized in that: Two cameras with different precision are used to collect the ship navigation video sequence and preprocess it, specifically: A high-precision camera and a low-precision camera are configured, hardware synchronization is performed on the high-precision camera and the low-precision camera, and monitoring alignment is performed based on timestamps to obtain two channels of ship navigation video sequences respectively; The SIFT feature detection algorithm is used to detect scale space extreme values in the two-channel ship navigation video sequences, locate the key points of the video frames, calculate the main direction of the key points, and generate SIFT feature descriptors. In addition, the ORB feature detection algorithm is used to quickly locate the key points and generate BRIEF feature descriptors. Based on the SIFT feature descriptor and the BRIEF feature descriptor, a feature point set of a high-precision video frame and a low-precision video frame is generated, and the feature point set is matched by SIFT through Euclidean distance and by ORB through Hamming distance respectively; Obtain key point matching results and eliminate false matches, optimize key point matching pairs to achieve spatial alignment, and eliminate perspective differences.
3. The ship navigation monitoring video super-resolution method based on mixed attention according to claim 1 is characterized in that: A hybrid attention module is constructed for the two ship navigation video sequences respectively, and a high-precision video frame extraction is performed using the hybrid attention module of the high-precision video sequence, specifically: For high-precision video sequences, a hybrid attention module is constructed using spatial attention and semantic attention. The pre-processed ship navigation video sequence collected by the high-precision camera is used as the module input. The Laplacian gradient function is used in the spatial attention branch of the hybrid attention module to evaluate the clarity, and the spatial attention is configured according to the local clarity evaluation result of the video frame. In the semantic attention branch of the hybrid attention module, a MobileNetV2-based encoder is used to encode and map the video frame to the embedding space, a single-frame visual feature is obtained, a pre-trained classification network is constructed using the single-frame visual feature, and a class activation map is used to extract the visual semantics corresponding to the video frame; The task target of the ship navigation video is obtained, and the class activation map uses global average pooling to obtain the relationship between the single-frame visual features and the task target category. The single-frame visual features are projected using the weights of the classification network output layer to obtain a class activation map with the task target category label. The class activation map is activated using the Softmax function, and the semantic attention of the video frame is configured according to the obtained probability; The spatial attention and semantic attention of the video frames are combined, and the video frames that meet the preset weight threshold are selected according to the combined attention weights and output as the key frames of the ship navigation video sequence captured by the high-precision camera.
4. The ship navigation monitoring video super-resolution method based on mixed attention according to claim 3 is characterized in that: Use the hybrid attention module of low-precision video sequences to extract low-precision video frames, specifically: For low-precision video sequences, temporal attention and motion attention are used to construct a hybrid attention module, and the pre-processed ship navigation video sequence collected by the low-precision camera is used as the module input, and the optical flow vector of the task target in the video frame is calculated in the hybrid attention module; In the temporal attention branch, the optical flow vector of each video frame is encoded and represented, and the GRU unit is used for temporal modeling to capture the long-term dependency of the optical flow vector sequence, and the temporal attention weight is generated according to the hidden state of the GRU unit; In the motion attention branch, the amplitude of each optical flow vector is calculated, and the average motion intensity of the video frame is obtained according to the amplitude of the optical flow vector, and the main direction of the optical flow vector is obtained, and the angle between the motion direction of all optical flow vectors in the video frame and the main direction is calculated, and motion consistency evaluation is performed to obtain the motion direction consistency evaluation result; The motion attention weight of the video frame is generated by the average motion intensity and motion direction consistency evaluation results of the video frame. The temporal attention weight and motion attention weight of the video frame are combined, and the video frames that meet the preset weight threshold are screened and output as the key frames of the ship navigation video sequence captured by a low-precision camera.
5. The ship navigation monitoring video super-resolution method based on mixed attention according to claim 1 is characterized in that: The key frames of the high-precision video and the low-precision video are adaptively and alternately superimposed to generate a superimposed video sequence, specifically: Obtain key frames of high-precision video sequences and low-precision video sequences, use the motion offset between the key frames of the two video sequences to perform motion compensation alignment, and obtain the absolute value of the motion feature difference between the high-precision video key frames and the low-precision video key frames; When the absolute value of the motion feature difference is greater than a preset threshold, a key frame with a larger motion feature value is selected; otherwise, a high-precision video key frame and a low-precision video key frame are selected according to an odd-even alternating strategy; The superimposed video sequence is obtained to establish the inter-frame structural similarity index monitoring for visual consistency check. When the preset visual consistency standard is met, the superimposed video sequence is output.
6. The ship navigation monitoring video super-resolution method based on mixed attention according to claim 1 is characterized in that: A super-resolution reconstruction network is constructed, the superimposed video sequence is used as a model input, the superimposed video sequence is enhanced, and a super-resolution video sequence is output, specifically: The superimposed video sequences are grouped according to different time steps. The target frame is selected in each group according to the motion feature value. The 3D dense blocks connected by residual connection are used in each group to obtain the spatiotemporal features. The spatiotemporal features in the group are integrated to obtain the group-level features. Convolve the group-level features of each group to obtain a single-channel feature map, connect the single-channel feature map along the time axis, introduce a self-attention mechanism to use softmax to calculate the self-attention weight for each position across channels, and use the self-attention weight to multiply the group-level features of different groups at the same position element by element; The result of multiplying the cascaded component features and the self-attention weights is imported into the 3D dense block to integrate the local features of different groups. A 2D dense block is added on top of the 3D dense block to further fuse the features and generate an aggregated feature map. The aggregated feature map is enhanced by hybrid attention using channel attention, spatial attention and temporal attention. The enhanced aggregated feature map is fused using a feature pyramid structure. The fused features are upsampled to obtain a high-resolution image consistent with the video frame scale, and are connected along the time axis to obtain a super-resolution video sequence.
7. The ship navigation monitoring video super-resolution method based on mixed attention according to claim 6 is characterized in that: Adversarial training is introduced in the super-resolution reconstruction network, specifically: A discriminator module is introduced into the super-resolution reconstruction network. The discriminator module is divided into a spatial discriminator branch and a temporal discriminator branch. The acquired high-resolution image is imported into the spatial discriminator for feature recognition to determine whether the input high-resolution image is a real high-resolution frame or a generated super-resolution frame. The acquired super-resolution video sequence is imported into the temporal discriminator for feature identification to determine whether the input super-resolution video sequence has real temporal dynamic characteristics; The discriminant matrices output by the spatial discriminator branch and the temporal discriminator branch are obtained, and the training of the super-resolution reconstruction network is guided according to the discriminant matrix according to the adversarial mechanism. After iterative training, the training is terminated until the discriminator module cannot distinguish whether the generated super-resolution video sequence is a real super-resolution video sequence, and the trained super-resolution reconstruction network is obtained.
8. A ship navigation monitoring video super-resolution system based on mixed attention, characterized in that: Implementing the ship navigation monitoring video super-resolution method based on hybrid attention as described in any one of claims 1 to 7, comprising a data acquisition and preprocessing unit, a hybrid attention frame extraction unit, a key frame superposition unit, a super-resolution reconstruction unit and a post-processing optimization unit; The data acquisition and preprocessing unit uses two cameras with different precisions to acquire and preprocess the ship navigation video sequence, and obtains two channels of preprocessed ship navigation video sequences; The hybrid attention frame extraction unit constructs a hybrid attention module for the two ship navigation video sequences respectively, and uses the hybrid attention module to perform high-precision video frame extraction and low-precision video frame extraction to obtain the key frame with the most information; The key frame superposition unit aligns the key frames of the high-precision video and the low-precision video, and performs adaptive alternating superposition to generate a superimposed video sequence; The super-resolution reconstruction unit constructs a super-resolution reconstruction network, uses a hybrid attention mechanism to enhance the superimposed video sequence, and reconstructs the super-resolution video sequence through adversarial training; The post-processing optimization unit performs post-processing optimization on the super-resolution result of the ship navigation video to enhance the edge and reduce the inter-frame jitter, and lightweight super-resolution reconstruction network to achieve real-time super-resolution.
Citation Information
Patent Citations
A key frame extraction method for ship surveillance video based on bidirectional GRU and attention mechanism
CN109508642A
Video super-resolution method based on multi-frame attention mechanism progressive fusion
CN112991183A
Video generation method and device and related equipment
CN113099146A
Multi-camera face super-resolution method and system based on domain migration fusion network
CN115131205A
High-frame-rate super-resolution improvement method based on multi-modal acquisition
CN115393194A
Cited By
Electric power scene illegal border crossing detection method and device and storage medium
CN121259742A
A power scene illegal boundary crossing detection method and device and a storage medium
CN121259742B