Ship Target Recognition and Tracking Method Based on Multi-Source Video Data Fusion

By using a combination of three-dimensional surface wavelet transformation and motion energy in video fusion, multi-source videos are fused, which solves the problems of information loss and time-space consistency caused by video feature differences under different environmental conditions, and achieves efficient video fusion and target recognition.

CN117173606BActive Publication Date: 2025-06-20XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310935734.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-27
Publication Date
2025-06-20
Estimated Expiration
2043-07-27

AI Technical Summary

Technical Problem

The feature differences of different types of videos acquired under different environmental conditions in the prior art have not been fully considered, resulting in information loss and low temporal and spatial consistency, making it difficult to achieve efficient video fusion and target recognition.

Method used

A video fusion method based on three-dimensional surface wavelet transformation is adopted to fuse multi-source videos. By detecting and tracking target feature points, a feature trajectory is constructed, and a candidate parameter vector is calculated using an error function to finally achieve efficient fusion of videos.

Benefits of technology

It improves the temporal stability and consistency of video fusion, can better extract spatio-temporal information, suppress noise, and improve video quality, so that the ship can efficiently and accurately identify surrounding targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117173606B_ABST
    Figure CN117173606B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for ship target recognition and tracking based on multi-source video data fusion, including: obtaining multi-source videos of ship targets; using the video in the front-facing direction as the main-direction video, aligning different-source videos from other angles with the main-direction video respectively to obtain the aligned multi-source videos; adopting a video fusion method combining motion energy based on three-dimensional surface wavelet transform to fuse multiple frames of images of the aligned multi-source videos to obtain a fused video; and using the fused video for ship target recognition and tracking. This method combines the motion-based fusion rule and the energy-based fusion rule, taking into account both temporal information and spatial information. It not only obtains good spatial information extraction ability, but also achieves good effects in terms of temporal stability and consistency, achieving a better fusion effect, improving the video quality, and enabling ships to efficiently and accurately recognize surrounding targets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of target recognition, and particularly relates to a ship target recognition and tracking method based on multi-source video data fusion. Background Art

[0002] With the continuous development of satellite and remote sensing technologies, in the 1970s, satellite remote sensing technology was increasingly applied to aerial photography and became the most widely used aerial surveying technology. Currently, satellite remote sensing technology has become an essential tool in fields such as military reconnaissance, ocean monitoring, and meteorological observation. There are significant risks in maritime operations. Using satellite imaging technology and artificial intelligence technology to strengthen the observation of the ocean and achieve rapid and effective identification of targets of maritime ships is an important technology for the maritime strategic layout, which helps to improve the efficiency of maritime operations and the ability of prediction and early warning. However, the maritime weather is complex and changeable, easily forming obstructions such as clouds, fog, and waves, which affect subsequent identification and tracking after target capture.

[0003] Videos contain rich target information. Common videos include thermal infrared videos and visible light videos. Visible light videos usually have rich textures; while the target may be unclear due to camouflage, including being hidden in thick fog at sea, in dark conditions at night, etc.; in such cases, the target cannot be clearly seen. Infrared videos can represent hidden targets in the low-frequency components; while infrared videos usually lack textures. Since visible light and thermal infrared videos have different complementary advantages, visible light and thermal infrared videos can be fused to improve the overall perception quality. However, visible light videos taken under different weather conditions (e.g., on foggy days or at night) may exhibit different characteristics, and thermal infrared videos are easily affected by changes in environmental temperature; this sensitivity to environmental conditions makes the fusion of visible light and thermal infrared videos a challenge. In addition, there are different types of video data such as remote sensing and multispectral, so multi-source video fusion is a very important issue.

[0004] A video usually contains many regions with different types of features, such as time motion targets and spatial geometric features. If these static image fusion rules are simply extended and applied to video fusion, all regions with different types of features in the source videos will be treated equally, which will reduce the spatio-temporal consistency to a certain extent. Currently, the proposed video fusion methods are mainly divided into two categories, namely video-based fusion methods and frame-based fusion methods.

[0005] Due to the special relationship between video and image, videos are generally fused frame by frame using static image fusion algorithms. The focus of existing frame-based fusion methods is to preserve the details of the images and maximize the texture of the sources. In video-based methods, existing methods have proposed three-dimensional non-separable multi-scale transform (MST) tools such as three-dimensional surfacelet transform (ST), dual-tree complex wavelet transform, and undecimated discrete curvelet transform (UDCT). Using these 3D-MST tools, new methods for video fusion have been developed. However, these 3D transforms have a high computational complexity and require the complete information of the entire video, which is only applicable to offline applications.

[0006] In frame-based methods, two videos are fused frame by frame, and the fusion can be performed in the spatial domain or the transform domain. Some people use genetic algorithms for adaptive fusion, but genetic algorithms require multiple iterations and thus are inefficient. There are also image fusion methods that consider the time complexity and use guided filters. However, these algorithms ignore the differences between different types of video frames, which easily causes information loss and affects the global performance of the fusion result.

[0007] Therefore, existing algorithms do not fully consider the different characteristics of different types of videos obtained under different environmental conditions and adopt the same processing method for them. This single-frame-based fusion algorithm only considers the spatial information of the input video, wastes the rich temporal information in the input video, and has low spatio-temporal consistency. Summary of the Invention

[0008] To solve the above problems existing in the prior art, the present invention provides a ship target recognition and tracking method based on multi-source video data fusion. The technical problems to be solved by the present invention are realized through the following technical solutions:

[0009] An embodiment of the present invention provides a ship target recognition and tracking method based on multi-source video data fusion, including the steps of:

[0010] Obtain multi-source videos of ship targets;

[0011] Use the video in the front-facing direction as the main direction video, and align the different source videos from other angles with the main direction video respectively to obtain the aligned multi-source videos;

[0012] Adopt a video fusion method combining motion energy based on three-dimensional surfacelet transform to fuse multiple frames of images of the aligned multi-source videos to obtain a fused video;

[0013] Use the fused video for ship target recognition and tracking.

[0014] In an embodiment of the present invention, taking the video in the front view direction as the main direction video, aligning different source videos from other angles with the main direction video respectively to obtain the aligned multi-source video, including:

[0015] Detect and track the feature points of the target in the multi-source video, and construct a number of feature trajectories;

[0016] Estimate the basic attributes of each of the feature trajectories;

[0017] Construct an initial table for preliminary matching between the trajectories based on the basic attributes;

[0018] Select a pair of target trajectories from the initial table, calculate the error function between the target trajectories, obtain the candidate parameter vector when the error function is minimized, and repeat this step several times to obtain several candidate parameter vectors;

[0019] Select the candidate parameter vector with the highest score from the several candidate parameter vectors as the target parameter;

[0020] Taking the video in the front view direction as the main direction video, use the target parameter to align different source videos from other angles with the main direction video respectively to obtain the aligned multi-source video.

[0021] In an embodiment of the present invention, the basic attributes include one or more of dynamic attributes, spatial attributes, and shape attributes.

[0022] In an embodiment of the present invention, the calculation formula of the error function is:

[0023]

[0024] where, [t0,..., t n is the time support of the spatio-temporal trajectory γ, is the spatial position of the spatio-temporal point at time t, is the spatial position of the spatio-temporal point in another sequence at time t′ = s·t + Δt, s is the ratio between the frame rates of two cameras, and Δt is an unknown parameter;

[0025] The calculation formula of the candidate parameter vector is:

[0026]

[0027] where, Γ is the set of all trajectories in the main direction video, d(·) is dist H (·) or dist F (·), for the three-dimensional case, is the error metric, distF (l, q) is the distance in pixels between point q and the epipolar line l, and F is the fundamental matrix in the 3D case.

[0028] In one embodiment of the present invention, a video fusion method combining motion energy based on three-dimensional surface wavelet transform is used to fuse multiple frames of images of the aligned multi-source videos, obtaining a fused video, including:

[0029] Using the three-dimensional surface wavelet transform method to decompose each of the aligned videos into subbands of different scales and directions, where the subbands of different scales and directions include bandpass direction subbands and low-pass direction subbands;

[0030] Using a multi-source video fusion method based on ST motion to fuse the bandpass direction subband coefficients, obtaining bandpass direction subband fusion coefficients, and using the bandpass direction subband fusion coefficients to fuse the aligned multi-source videos, obtaining a bandpass direction fused video;

[0031] Frame by frame, merging the low-pass direction subband coefficients in the ST domain, obtaining low-pass direction subband fusion coefficients, and using the low-pass direction subband fusion coefficients to fuse the aligned multi-source videos, obtaining a subband direction fused video;

[0032] Performing an inverse 3D-ST reconstruction on the bandpass direction fused video and the subband direction fused video to obtain the fused video.

[0033] In one embodiment of the present invention, using a multi-source video fusion method based on ST motion to fuse the bandpass direction subband coefficients, obtaining bandpass direction subband fusion coefficients, includes the steps of:

[0034] Using an improved ST domain z-score method, combining the bandpass direction subband coefficients to distinguish the temporal motion information and spatial geometric information of each frame of video in the aligned multi-source videos, obtaining a motion detection map;

[0035] According to the motion detection map, dividing each bandpass direction subband of the aligned multi-source videos into a first region, a second region, and a third region, where the first region represents that motion information is detected in one video, the second region represents that motion information is detected in at least 2 videos, and the third region represents that no motion information is detected in all videos;

[0036] Fusing the bandpass direction subband coefficients according to the number of videos in which motion information is detected in the first region, the second region, and the third region, obtaining bandpass direction subband fusion coefficients for different regions.

[0037] In one embodiment of the present invention, by using an improved ST-domain z-score method and combining the band-pass direction sub-band coefficients, the temporal motion information and spatial geometric information of each frame of video in the aligned multi-source video are separated to obtain a motion detection map, including:

[0038] At a fixed scale and sub-direction, judge whether there is a moving object in the current frame by using the temporal variation of the band-pass direction sub-band coefficients of the video; wherein, the temporal variation of the band-pass direction sub-band coefficients is:

[0039] Δ (s,j,k) (x,y,t) = Y (s,j,k) (x,y,t + 1) - Y (s,j,k) (x,y,t)

[0040] wherein, s is the scale, (j,k) is the sub-direction, Y (s,j,k) (x,y,t) is the ST band-pass direction sub-band coefficient, t is the time dimension, and (x,y) is the spatial position;

[0041] When there is a moving object in the current frame, use the modified z-score method to perform two motion detections on the moving object in each sub-band to obtain a motion detection map:

[0042]

[0043]

[0044]

[0045] wherein, is the median of Δ (s,j,k) (x,y,t) along the x and y dimensions in the current frame t, MAD (s,j,k) (t) is the median of absolute deviations, β is a predefined experimental threshold, and Z (s,j,k) (x,y,t) is the modified z-score.

[0046] In one embodiment of the present invention, the definitions of the first region, the second region, and the third region are:

[0047]

[0048] wherein, is the motion mapping, and the region indicates that motion information is detected in a video Vm, indicates that motion information is detected in multiple or all videos, It is indicated that no motion information is detected in all videos. Vm represents the input video, where m ∈ 1, 2, ..., n, and n represents the total number of input videos.

[0049] In an embodiment of the present invention, the band-pass direction sub-band coefficients are fused according to the number of videos in which motion information is detected in the first region, the second region, and the third region, to obtain the band-pass direction sub-band fusion coefficients for different regions, including:

[0050] For the first region, the band-pass direction sub-band fusion coefficient takes the band-pass direction sub-band coefficient of the video in which motion information is detected:

[0051]

[0052] For the second region, the coefficients with significant motion information are selected as the band-pass direction sub-band fusion coefficients:

[0053]

[0054] Wherein, is the motion significance, defined as:

[0055]

[0056] Wherein, K s,j represents the total number of directions of the s-th scale and the j-th hourglass branch, T1 is the local frame number, and τ represents the time variation.

[0057] For the third region, the band-pass direction sub-band coefficients are fused by using a fusion rule based on spatio-temporal energy matching to obtain the band-pass direction sub-band fusion coefficients for the third region.

[0058] In an embodiment of the present invention, for the third region, the band-pass direction sub-band coefficients are fused by using a fusion rule based on spatio-temporal energy matching to obtain the band-pass direction sub-band fusion coefficients for the third region, including:

[0059] For each scale, direction, and spatial position, the average value of the ST band-pass direction sub-band coefficients of each video along the time dimension is calculated to obtain the average value:

[0060]

[0061] Wherein, T(s, j, k) represents the number of frames included in the current sub-band, represents the ST band-pass direction sub-band coefficient of each video, t is the time dimension, s is the scale, (j, k) is the sub-direction, and (x, y) is the spatial position;

[0062] According to the average value, the energy of the local spatial region centered on the spatial position in each video is calculated:

[0063]

[0064] Among them, A1×B1 is the size of the local spatial region;

[0065] Calculate the similarity between the energies of multiple videos:

[0066]

[0067] Among them, VC represents multiple videos;

[0068] For each frame in the current sub-band, determine the band-pass direction sub-band fusion coefficient of the third region according to the comparison result between the similarity and the adaptive threshold:

[0069]

[0070] Among them, is the local weight, m ∈ (1, n), γ (s,j,k) is the adaptive threshold, represents the band-pass direction sub-band coefficient of the video when the video energy is the largest, represents the largest video energy.

[0071] Compared with the prior art, the beneficial effects of the present invention:

[0072] The present invention adopts a video fusion method combining motion energy based on three-dimensional surface wavelet transform to fuse the aligned multi-source videos; independently detects and fuses the time motion target regions based on the motion-based fusion rule, has high time stability and consistency, can be used for real-time fusion, and has high calculation efficiency; based on the energy-based fusion rule, uses the motion selectivity of three-dimensional surface wavelet transform, and adopts the fusion rule based on spatio-temporal region energy to extract spatio-temporal information from the input videos, has good spatio-temporal information extraction ability, and can suppress noise; combines the two algorithms, taking into account both time information and space information, not only obtains good spatio-temporal information extraction ability, but also achieves good results in terms of time stability and consistency, thus achieving a better fusion effect, improving the video quality, enabling the ship to efficiently and accurately identify surrounding targets, and being conducive to making judgments on subsequent navigation behaviors such as timely tracking of the ship. Brief Description of the Drawings

[0073] Figure 1 is a schematic flowchart of a ship target recognition and tracking method based on multi-source video data fusion provided by an embodiment of the present invention;

[0074] Figure 2Schematic diagram of another ship target recognition and tracking method based on multi-source video data fusion provided by an embodiment of the present invention;

[0075] Figure 3 Block diagram schematic of video alignment provided by an embodiment of the present invention;

[0076] Figure 4 Block diagram schematic of video fusion provided by an embodiment of the present invention. Detailed implementation manners

[0077] The present invention will be further described in detail below in conjunction with specific embodiments, but the implementation manners of the present invention are not limited thereto.

[0078] Embodiment 1

[0079] Please refer to Figure 1 and Figure 2 , Figure 1 which is a schematic flow diagram of a ship target recognition and tracking method based on multi-source video data fusion provided by an embodiment of the present invention, Figure 2 and which is a schematic diagram of another ship target recognition and tracking method based on multi-source video data fusion provided by an embodiment of the present invention.

[0080] This method is a video-based fusion method that directly performs video-to-video fusion on multi-source video data of ship targets. The entire framework of this method is divided into two parts. The first part is video alignment, and the second part is the video fusion part. In the video alignment part, videos are taken with the front view direction as the main direction, and videos of different types from other angles are respectively aligned with the video in the front view direction. Here, a video sequence matching method based on spatio-temporal trajectories is adopted. In the video fusion part, an improved video fusion method combining motion and energy based on three-dimensional surface wavelet transform (3D-ST) is adopted. 3D-ST is used to fuse multiple frames of images of the aligned input video as a whole. The video signal is regarded as existing in a special three-dimensional space (two spatial dimensions and one time dimension), and the signals are fused simultaneously.

[0081] This method includes the steps of:

[0082] S1. Obtain multi-source videos of ship targets.

[0083] Specifically, the multi-source videos may include infrared videos, visible light videos, remote sensing videos, multi-spectral videos, etc.

[0084] S2. Use the video in the front view direction as the main direction video, and align videos from different sources at other angles with the main direction video respectively to obtain the aligned multi-source videos.

[0085] Let S and S' be two input image sequences, where S is the "reference" sequence and S' is the second sequence. Let is a spatio-temporal point in the reference sequence S, i.e., a pixel (x, y) at frame (time) t. Let be the matching spatio-temporal point in the sequence S′. The recorded scene can change dynamically, that is, it can contain moving objects. The camera can be either stationary or moving jointly with fixed (but unknown) internal and relative external parameters. In this setting, the spatio-temporal correspondence between video sequences can be described / modeled by a set of parameters . The goal of video alignment is to recover these parameters.

[0086] When there is a time shift (offset) between two input sequences (e.g., if the cameras are not activated simultaneously), and / or when they have different frame rates (e.g., PAL vs. NTSC sequences), a temporal misalignment occurs. This temporal misalignment can be modeled by a one-dimensional affine transformation of time t′ = s·t + Δt, and is usually in sub-frame time units. Note that in most cases s is known, which is the ratio between the frame rates of the two cameras (e.g., for PAL and NTSC sequences, s = 25 / 30 = 5 / 6). Therefore, in this case, it contains only one unknown parameter Δt.

[0087] To model the spatial parameters, the spatial part of the spatio-temporal point is discussed. Let represent the homogeneous coordinates of the spatial component of the spatio-temporal point in S. The spatial misalignment between the two sequences is due to the two cameras having different external and internal calibration parameters. Two possible cases are considered here: the two-dimensional case and the three-dimensional case.

[0088] In the two-dimensional case, the distance between the camera projection centers is negligible compared to the distance from the camera to the scene, or if the scene is approximately planar. In this two-dimensional case, the spatio-temporal relationship between the two sequences is represented by an unknown 3×3 homography H and an unknown Δt:

[0089]

[0090] In one example, i.e., are 9 spatial parameters:

[0091]

[0092] defined up to a scale factor (h ij are the 9 entries of H) and

[0093] In the three-dimensional case, where the cameras do not intersect and the scene contains observable 3D changes, the spatio-temporal relationship between two sequences is represented by an unknown fundamental matrix F and an unknown Δt:

[0094]

[0095] where [·]T represents the transpose of a vector.

[0096] In this case, the spatial relationship parameters are: where f ij are the 9 entries of the 3×3 fundamental matrix F (up to a scale factor),

[0097] Note that in both cases, since the cameras are relatively fixed with respect to each other (both the internal parameters and the external parameters between the cameras are fixed), F or H is shared by all pairs of temporally corresponding frames.

[0098] Furthermore, let be a spatio-temporal trajectory, represent spatio-temporal points, and let Γ and Γ′ denote the sets of all trajectories in sequences S and S′ respectively. Then, the spatio-temporal matching between the two sequences is recovered by establishing correspondences between the trajectories in the sets Γ and Γ′.

[0099] See Figure 3 Figure 3 which is a block diagram schematic of the video alignment provided by an embodiment of the present invention. The video sequence matching method based on spatio-temporal trajectories includes the steps of:

[0100] S21, Detect and track the feature points of the targets in multi-source videos, and construct several feature trajectories.

[0101] Specifically, calculate the trajectory of a moving object by tracking the unique point on the blob of the moving object as a feature point. For example, the trajectory of a moving object can be calculated by tracking the centroid of the moving object or the vertex on the contour of the moving object, thereby constructing several feature trajectories.

[0102] S22, Estimate the basic attributes of each feature trajectory. Among them, the basic attributes include one or more of dynamic attributes, spatial attributes, and shape attributes.

[0103] ​Specifically, in the case of a large number of trajectories, trajectory attributes can be used to reduce the matching complexity. The dynamic trajectories (moving objects) in one sequence are only matched with the dynamic trajectories in another sequence. When it is expected that the cameras have similar photometric attributes, the spatial attributes of the features (e.g., the size, gradient, or color distribution of the moving object) can also be used. When significant changes in appearance are expected, the shape attributes of the trajectories can still be used (e.g., normalized length, average speed, curvature). Although some of them are not photometric invariants, they are useful in the initial search for rough tentative matches, i.e., step S23.

[0104] S23. Construct an initial table of preliminary matches between trajectories based on basic attributes.

[0105] S24. Select a pair of target trajectories from the initial table, calculate the error function between the target trajectories, and obtain the candidate parameter vector when the error function is minimized. Repeat this step several times to obtain several candidate parameter vectors.

[0106] First, randomly select a pair of potentially corresponding trajectories and estimate the candidate parameter vector In each trial, calculate the parameter set that minimizes the error function The parameter set that minimizes

[0107] A pair of corresponding trajectories γ and γ′ can uniquely define: (i) the spatial relationship. (ii) the temporal relationship, (iii) the residual measurement (i.e., the error function):

[0108]

[0109] where, [t0,..., t n is the time support of the spatio-temporal trajectory γ, is the spatio-temporal point The spatial position (i.e., pixel coordinates) at time t (homogeneous coordinates), is the spatial position of the spatio-temporal point in another sequence at time t′ = s · t + Δt, s is the ratio between the frame rates of the two cameras, and Δt is the unknown parameter;

[0110] For the fundamental matrix in the three-dimensional case, the error metric is: is the error metric, dist F (l, q) is the distance in pixels between the point q and the epipolar line l, and F is the fundamental matrix in the 3D case.

[0111] The matching of a pair of trajectories on two sequences results in multiple point correspondences on the camera views. These point pairs are used to calculate the spatio-temporal relationship between the two sequences. Find a candidate parameter vector or Among them, h 11 , ..., h 33 or f 11 , ..., f 33 are the components of the homography H or the fundamental matrix F respectively, minimizing the following candidate parameter vectors:

[0112]

[0113] Among them, Γ is the set of all trajectories in the dominant direction video, and d(·) is dist H (·) or dist F (·), and d(·) depends on whether the scene is a two-dimensional or three-dimensional case (the summation is only for the selected trajectories).

[0114] The minimization of the above formula is completed by iterating the following two steps:

[0115] (i) Fix Δt and approximate H (or F) using standard methods

[0116] (ii) Fix H (or F) and refine Δt. Since t′ = s·t + Δt is not necessarily an integer value (allowing sub-frame time shift), it is interpolated from the positions of adjacent (integer-time) points: and Search for α = t′ - t1 (1 ≥ α ≥ 0) that minimizes the following term:

[0117]

[0118] Use a finite number of refinement iterations (10 to 20), or stop early if the remaining error does not change. The initial (integer) approximation of Δt is obtained by exhaustive search within a small fixed time interval (20 - 25 frames).

[0119] An error function corresponds to a Therefore, after calculating multiple corresponding and select when it is the minimum value of

[0120] Then, repeat the above steps several times until the difference between the two calculated is less than the preset threshold, and then stop repeating to obtain multiple candidate parameter vectors.

[0121] S25. Select the candidate parameter vector with the highest score from several candidate parameter vectors as the target parameter.

[0122] Specifically, select the candidate parameter vector with the largest value from multiple candidate parameter vectors as the target parameter.

[0123] S26. Use the video in the frontal view direction as the main direction video, and align the different source videos at other angles with the main direction video respectively by using the target parameters to obtain the multi-source video after alignment.

[0124] S3. Adopt a video fusion method combining motion energy based on three-dimensional surface wavelet transform to fuse multiple frames of images of the aligned multi-source video to obtain a fused video.

[0125] Please refer to Figure 4 , Figure 4 which is a block diagram schematic of the video fusion provided by the embodiment of the present invention. Figure 4 Adopt a video fusion method combining motion energy based on three-dimensional surface wavelet transform (3D-ST) for video fusion. The idea is as follows: First, 3D-ST decomposes the input video into subbands of different scales and directions; second, adopt a fusion rule to fuse the subband coefficients of the input video; finally, perform inverse 3D-ST reconstruction on the fused video. Among them,

[0126] Specifically, the video fusion method includes the following steps:

[0127] S31. Adopt the three-dimensional surface wavelet transform method to decompose each aligned video into subbands of different scales and directions. The subbands of different scales and directions include band-pass direction subbands and low-pass direction subbands.

[0128] S32. Adopt a multi-source video fusion method based on ST motion to fuse the band-pass direction subband coefficients to obtain band-pass direction subband fusion coefficients, and use the band-pass direction subband fusion coefficients to fuse the aligned multi-source video to obtain a band-pass direction fused video.

[0129] S33. Frame by frame, merge the low-pass direction subband coefficients in the ST domain to obtain low-pass direction subband fusion coefficients, and use the low-pass direction subband fusion coefficients to fuse the aligned multi-source video to obtain a subband direction fused video.

[0130] S34. Perform inverse 3D-ST reconstruction on the band-pass direction fused video and the subband direction fused video to obtain a fused video.

[0131] In the above video fusion method, steps S31 and S34 are implemented by using existing methods; step S33 can adopt existing static image fusion rules to frame by frame merge the low-pass direction subband coefficients in the ST domain, so as to obtain low-pass direction subband fusion coefficients.

[0132] For step S32, the band-pass direction sub-band coefficients contain motion information, and motion is the most important feature for distinguishing videos from static images. In the ST domain, coefficients with larger absolute values may correspond to significant spatial structures or moving objects. Therefore, in this embodiment, an improved z-score is first used to detect temporal motion information in the ST domain to distinguish the temporal motion information and static background information (i.e., spatial geometric information) of the aligned multi-source video; then, in order to improve the fusion performance and computational efficiency, a motion-based fusion rule is adopted to process the temporal motion information and spatial geometric information separately to obtain the band-pass direction sub-band fusion coefficients. Step S32 specifically includes the steps:

[0133] S321. Use the improved ST-domain z-score method to distinguish the temporal motion information and spatial geometric information of each frame of the aligned multi-source video by combining the band-pass direction sub-band coefficients, and obtain a motion detection map. Specifically, it includes:

[0134] 1). At a fixed scale and sub-direction, judge whether there are moving objects in the current frame by using the temporal variation of the band-pass direction sub-band coefficients of the video.

[0135] Specifically, assume that the ST band-pass direction sub-band coefficients of the video are represented as {Y (s,j,k) (x, y, t)}. At a fixed scale s and sub-direction (j, k), the temporal variation of the ST band-pass direction sub-band coefficients is:

[0136] A (s,j,k) (x,y,t) = Y (s,j,k) (x,y,t + 1) - Y (s,j,k) (x,y,t)

[0137] where s is the scale, (j, k) is the sub-direction, Y (s,j,k) (x, y, t) is the ST band-pass direction sub-band coefficient, t is the time dimension, and (x, y) is the spatial position.

[0138] If there are no moving objects in the current frame, the differences of the obtained ST band-pass direction sub-band coefficients at each scale and direction follow a normal distribution with a mean of zero and a small standard deviation. If there are moving objects in the current frame, the differences of some ST band-pass direction sub-band coefficients will change greatly. Under the assumption that the moving objects only occupy a small part of the frame, the ST coefficient differences corresponding to the moving objects will become outliers of the normal distribution. Therefore, the problem of the coefficients corresponding to the motion is similar to the problem of detecting outliers in the set of ST coefficient differences.

[0139] 2). When there are moving objects in the current frame, use the modified z-score method to perform two motion detections on the moving objects in each sub-band to obtain a motion detection map.

[0140] Specifically, the corrected z-score is calculated as follows:

[0141]

[0142] Where is the Δ along the x and y dimensions of the current frame t (s,j,k) (x, y, t) is the median, MAD (s,j,k) (t) is the median of the absolute deviation,

[0143] In 3D-ST, the resampling matrix adopted varies with the scale and the hourglass filtering branch. Therefore, motion detection is performed scale by scale in each hourglass filtering branch, and then all results in different directions are combined through a logical OR operation to obtain the motion map for each scale and hourglass filtering branch, that is, by comparing the modified Z-score Z (s,j,k) (x, y, t) with a predefined time-delay threshold β to obtain the coefficient corresponding to the motion:

[0144]

[0145] Where β is the experimental threshold, which can be set to 10.

[0146] To eliminate ghosts in the outlier detection mask, two corrected z-score tests are performed to obtain the final motion detection map:

[0147]

[0148] It can be seen from the above formula that when motion information is detected both times, the conclusion that there is motion information is drawn.

[0149] S322. According to the motion detection map, each band-pass direction sub-band of the aligned multi-source video is segmented into a first region, a second region, and a third region. Among them, the first region indicates that motion information is detected in one video, the second region indicates that motion information is detected in at least two videos, and the third region indicates that no motion information is detected in all videos.

[0150] Specifically, through the motion detection algorithm in step S321, each band-pass direction sub-band of the input videos V1, V2,..., Vn+1 can be segmented into multiple regions, respectively represented as m ∈ 1, 2,..., n, where n represents the total number of input videos. In this embodiment, the multiple regions include a first region, a second region, and a third region. Among them, the first region indicates that motion information is detected in one video, the second region indicates that motion information is detected in at least two videos, and the third region indicates that no motion information is detected in all videos. The definitions of the first region, the second region, and the third region are as follows:

[0151]

[0152] Among them, is the motion mapping defined in step S321, that is, the motion detection map.

[0153] For m, there is only one value that satisfies the first region indicating that motion information is detected in only one video Vm. For at least m, there are two values that satisfy the second region indicating that motion information is detected in more than one video, that is, in multiple or all videos. For the third region indicating that no motion is detected in all videos, that is, the region only contains static background (spatial geometry) information.

[0154] S323. According to the number of videos in which motion information is detected in the first region, the second region, and the third region, fuse the band-pass direction sub-band coefficients to obtain the band-pass direction sub-band fusion coefficients for different regions.

[0155] Specifically, in order to better process the temporal motion information and the static background information, different fusion rules need to be adopted for different regions.

[0156] 1). For the first region, the band-pass direction sub-band fusion coefficient takes the band-pass direction sub-band coefficient of the video in which motion information is detected.

[0157] Specifically, for the region m has only one value that satisfies Only the sub-band coefficients corresponding to the input video Vm contain motion information. Therefore, the band-pass direction sub-band coefficient of the fused video in this region directly takes the band-pass direction sub-band coefficient of the input video Vm in which motion information is detected

[0158]

[0159] 2). For the second region, select the coefficients with significant motion information as the band-pass direction sub-band fusion coefficients.

[0160] For the second region at least m has two values that satisfy Introduce a motion selectivity index to determine how to determine the sub-band coefficients corresponding to the input video, and select the coefficients with significant motion information as the fusion coefficients. Then the proposed fusion rule is:

[0161]

[0162] Among them, is the motion saliency, defined as:

[0163]

[0164] where K s,j represents the total number of directions of the sth scale and the jth hourglass branch, T1 is the local frame number, which can be set to 3 here, and τ represents the time variation.

[0165] 3) For the third region, fuse the band-pass direction sub-band coefficients using the fusion rule based on spatio-temporal energy matching to obtain the fused band-pass direction sub-band coefficients of the third region.

[0166] The third region contains static background information. For each spatial position (x, y), the coefficient is almost the same along the time dimension t. To improve the computational efficiency, the fusion rule based on spatio-temporal energy matching is used for merging the band-pass direction sub-band coefficients of the region. Specifically, it includes the steps:[[]]

[0167] First, calculate the average value of the ST band-pass direction sub-band coefficients of each video along the time dimension t for each scale s, sub-direction (j, k), and spatial position (x, y) respectively, to obtain the average value:

[0168]

[0169] where T(s, j, k) represents the number of frames contained in the current sub-band.

[0170] Second, calculate the energy of the local spatial region centered on the spatial position in each video according to the average value.

[0171] Specifically, for the average value calculate the energy of a local spatial region of size A1×B1 centered on the spatial position (x, y):

[0172]

[0173] where A1×B1 is the size of the local spatial region, and A1×B1 is set to 3×3.

[0174] Third, calculate the similarity between the energies of multiple videos.

[0175] Specifically, calculate the similarity between m = 1, 2,..., n

[0176]

[0177] ​Among them, VC represents multiple videos.

[0178] Finally, for each frame in the current subband, according to the similarity with the adaptive threshold γ (s,j,k) to determine the bandpass direction subband fusion coefficient of the third region:

[0179]

[0180] Among them, is the local weight, m ∈ (1, n), γ (s,j,k) is the adaptive threshold, represents the bandpass direction subband coefficient of the video when the video energy is the largest, represents the largest video energy.

[0181] In the above formula, for the case of η (s,j,k) (x, y) ≥ γ, the same fusion graph merging coefficient is adopted

[0182] After obtaining the bandpass direction subband fusion coefficients of each region, an existing method is used to fuse the aligned multi-source videos using the bandpass direction subband fusion coefficients to obtain a bandpass direction fused video.

[0183] S4. Use the fused video for ship target recognition and tracking.

[0184] Adopt an existing method to use the fused video for ship target recognition and tracking.

[0185] Compared with the traditional single-frame-based fusion algorithm, this embodiment has higher temporal stability and consistency. In addition, 3D-ST has motion selectivity, and even without a motion detection pre-step, the motion information of the input video can be correctly captured by the 3D-ST coefficients. In particular, compared with other 3D multiscale transforms such as 3D discrete wavelet transform, 3D-ST has higher directional resolution and can extract motion information from the input video more accurately.

[0186] In this embodiment, multiple frames of images of the input video are regarded as a special three-dimensional space, and the two-dimensional space (x, y) and the one-dimensional time t are simultaneously subjected to signal and fusion, rather than frame-by-frame fusion. Due to the motion selectivity of 3D-ST, these fusion algorithms can well extract important spatio-temporal information from the input video. An improved z-score-based motion detection is also adopted to distinguish temporal motion information and spatial geometric information, making it significantly superior to traditional frame-based and motion-based methods in terms of spatio-temporal information extraction and temporal stability and consistency.

[0187] In this embodiment, a video fusion method combining motion energy based on three-dimensional surface wavelet transform is adopted to fuse the aligned multi-source videos; independent detection and fusion of the time motion target area are performed based on the motion-based fusion rule, which has high time stability and consistency, can be used for real-time fusion, and has high computational efficiency; based on the energy-based fusion rule, with the motion selectivity of the three-dimensional surface wavelet transform, the spatio-temporal information is extracted from the input video by adopting the fusion rule based on spatio-temporal region energy, which has good spatio-temporal information extraction ability and can suppress noise. By combining the two algorithms and taking into account both time information and space information, not only good spatio-temporal information extraction ability is obtained, but also good results are achieved in terms of time stability and consistency, thus achieving a better fusion effect, improving the video quality, enabling the ship to efficiently and accurately identify surrounding targets, and facilitating the judgment of subsequent navigation behaviors such as timely tracking of the ship.

[0188] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A method for ship target recognition and tracking based on multi-source video data fusion, characterized in that, Including the steps: Obtain multi-source videos of the ship target; Take the video in the frontal direction as the main direction video, and align the different source videos from other angles with the main direction video respectively to obtain the aligned multi-source video; Adopt a video fusion method combining motion energy based on 3D surface wavelet transform to fuse multiple frames of images of the aligned multi-source video to obtain a fused video; including: decomposing each aligned video into sub-bands of different scales and directions by using the 3D surface wavelet transform method, where the sub-bands of different scales and directions include band-pass direction sub-bands and low-pass direction sub-bands; adopting a multi-source video fusion method based on ST motion to fuse the coefficients of the band-pass direction sub-bands to obtain band-pass direction sub-band fusion coefficients, and using the band-pass direction sub-band fusion coefficients to fuse the aligned multi-source video to obtain a band-pass direction fused video; merging the coefficients of the low-pass direction sub-bands frame by frame in the ST domain to obtain low-pass direction sub-band fusion coefficients, and using the low-pass direction sub-band fusion coefficients to fuse the aligned multi-source video to obtain a sub-band direction fused video; performing inverse 3D-ST reconstruction on the band-pass direction fused video and the sub-band direction fused video to obtain the fused video; Use the fused video for target recognition and tracking of the ship.

2. The method for ship target recognition and tracking based on multi-source video data fusion according to claim 1, characterized in that, Taking the video in the frontal direction as the main direction video, and aligning the different source videos from other angles with the main direction video respectively to obtain the aligned multi-source video, including: Detect and track the feature points of the target in the multi-source video, and construct several feature trajectories; Estimate the basic attributes of each of the feature trajectories; Construct an initial table for preliminary matching between the trajectories based on the basic attributes; Select a pair of target trajectories from the initial table, calculate the error function between the target trajectories, obtain the candidate parameter vector when the error function is minimized, and repeat this step several times to obtain several candidate parameter vectors; Select the candidate parameter vector with the highest score from the several candidate parameter vectors as the target parameter; Taking the video in the frontal direction as the main direction video, use the target parameter to align the different source videos from other angles with the main direction video respectively to obtain the aligned multi-source video.

3. The method for ship target recognition and tracking based on multi-source video data fusion according to claim 2, characterized in that, The basic attributes include one or more of dynamic attributes, spatial attributes, and shape attributes.

4. The method for ship target recognition and tracking based on multi-source video data fusion according to claim 2, characterized in that, The calculation formula of the error function is: Among them, is the time support of the spatio-temporal trajectory , is the spatio-temporal point at time under the spatial position of is the time when the spatio-temporal point in another sequence of spatial position is the ratio between the frame rates of two cameras is an unknown parameter; The calculation formula of the candidate parameter vector is: Among them, is the set of all trajectories in the main direction video, is or For the three-dimensional case, is the error metric, is the point and the epipolar line the distance in pixels between them, is the fundamental matrix in the 3D case.

5. The method for ship target recognition and tracking based on multi-source video data fusion according to claim 1, characterized in that, Adopt a multi-source video fusion method based on ST motion to fuse the coefficients of the band-pass direction sub-bands to obtain band-pass direction sub-band fusion coefficients, including the steps: Use the improved ST domain z-score method to separate the temporal motion information and spatial geometric information of each frame of video in the aligned multi-source video in combination with the coefficients of the band-pass direction sub-bands to obtain a motion detection map; According to the motion detection map, divide each band-pass direction sub-band of the aligned multi-source video into a first region, a second region, and a third region, where the first region represents that motion information is detected in one video, the second region represents that motion information is detected in at least 2 videos, and the third region represents that no motion information is detected in all videos; Fuse the band-pass direction sub-band coefficients according to the number of videos with motion information detected in the first region, the second region, and the third region, to obtain the band-pass direction sub-band fusion coefficients for different regions.

6. The method for ship target recognition and tracking based on multi-source video data fusion according to claim 5, characterized in that, Using the improved ST-domain z-score method, separate the temporal motion information and the spatial geometric information of each frame of video in the aligned multi-source videos in combination with the band-pass direction sub-band coefficients, to obtain a motion detection map, including: At a fixed scale and sub-direction, judge whether there are moving objects in the current frame by using the temporal variation of the band-pass direction sub-band coefficients of the video; wherein, the temporal variation of the band-pass direction sub-band coefficients is: Among them, is the scale, ( ) is the sub-direction, is the ST bandpass direction sub-band coefficient, is the time dimension, is the spatial position; When there are moving objects in the current frame, use the modified z-score method to perform two motion detections on the moving objects in each sub-band, to obtain a motion detection map: Among them, is the median of the current frame t along the x and y dimensions , , is the median of the absolute deviation, , is a predefined experimental threshold, is the corrected z-score.

7. The ship target recognition and tracking method based on multi-source video data fusion according to claim 6, characterized in that, The definitions of the first region, the second region, and the third region are: Among them, is the motion mapping, and the area indicates that motion information is detected in a video while indicates that motion information is detected in multiple or all videos, indicates that no motion information is detected in all videos, indicates the input video, , indicates the total number of input videos.

8. The ship target recognition and tracking method based on multi-source video data fusion according to claim 7, characterized in that, Fuse the band-pass direction sub-band coefficients according to the number of videos with motion information detected in the first region, the second region, and the third region, to obtain the band-pass direction sub-band fusion coefficients for different regions, including: For the first region, the band-pass direction sub-band fusion coefficient takes the band-pass direction sub-band coefficients of the videos with detected motion information: ; For the second region, select the coefficients with significant motion information as the band-pass direction sub-band fusion coefficients: Among them, is the motion saliency, defined as: Among them, represents the total number of directions of the s-th scale and the j-th hourglass branch, is the local frame number, represents the time variation; For the third region, adopt a fusion rule based on spatio-temporal energy matching to fuse the band-pass direction sub-band coefficients, to obtain the band-pass direction sub-band fusion coefficients of the third region.

9. The ship target recognition and tracking method based on multi-source video data fusion according to claim 8, characterized in that, For the third region, adopt a fusion rule based on spatio-temporal energy matching to fuse the band-pass direction sub-band coefficients, to obtain the band-pass direction sub-band fusion coefficients of the third region, including: Calculate the average value of the ST band-pass direction sub-band coefficients of each video along the time dimension for each scale, direction, and spatial position, to obtain the average value: Among them, represents the number of frames included in the current sub-band, represents the ST bandpass direction sub-band coefficients of each video, is the time dimension, is the scale, is the sub-direction, is the spatial position; According to the average value, calculate the energy of the local spatial region centered on the spatial position in each video: Among them, is the size of the local spatial region; Calculate the similarity between the energies of multiple videos: Among them, represents multiple videos; For each frame in the current sub-band, determine the band-pass direction sub-band fusion coefficients of the third region according to the comparison result between the similarity and the adaptive threshold: Among them, is the local weight, , is the adaptive threshold, , represents the bandpass direction subband coefficient of the video when the video energy is the largest, represents the maximum video energy.

Citation Information

Patent Citations

  • Ship target identification system and method based on artificial intelligence image processing

    CN109766830A

  • Indoor moving target positioning method based on Doppler sensor network

    CN111208507A