Water surface 3D target detection method based on multi-view space-time fusion
By adopting multi-view space-time fusion method in water surface object detection, the feature compression backbone network and decoder are used to perform semantic feature extraction and target embedding vector generation, and combining timing memory and alignment modules for motion alignment processing, the problem of difficult to balance perception and calculation amount and perception range in the 3D space in the prior art is solved, and efficient and accurate 3D object detection on water surface is achieved.
Patent Information
- Application Number
- CN202510218093.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-27
AI Technical Summary
The existing water surface object detection technology cannot be perceived in 3D space, and the lack of target distance information leads to safety risks. The multi-view image fusion algorithm is difficult to weigh between the calculation amount and the perception range.
The 3D object detection method of water surface based on multi-view space-time fusion is adopted, and multi-view images are obtained through devices equipped with cameras of different viewing angles, and semantic feature extraction and target embedding vector generation is used to combine timing memory and alignment modules to perform motion alignment processing to achieve efficient 3D object detection.
It improves the accuracy and efficiency of 3D object detection on the water surface, enhances the detection range, reduces the computational complexity, avoids spatial misalignment caused by camera shake, and significantly improves the detection accuracy.
Smart Images

Figure CN120220128A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a water surface 3D target detection method based on multi-view spatiotemporal fusion. Background Art
[0002] With the continuous global attention to the development and utilization of water resources, the inland and ocean shipping industries have developed rapidly. In recent years, the number of ships, transportation capacity and navigation speed have increased significantly, which has promoted the prosperity of the shipping industry. However, the problem that comes with it is that the incidence of ship collision accidents is also increasing. Therefore, how to perceive and warn surface targets is an important issue that needs to be solved urgently. Timely and accurate detection of surface targets plays an extremely important role in preventing ship collisions and is an important means to ensure the safe navigation of ships.
[0003] In the past, surface target detection usually used a single image as input to a trained deep learning model, and output a 2D target bounding box and category based on the image coordinates. Due to the camera's field of view, the perception range is often limited, and the surrounding environment cannot be fully perceived. It is also impossible to perceive in 3D space, and the lack of target distance information leads to safety hazards. Therefore, there is great potential in using multi-camera image fusion for 3D target perception.
[0004] In recent years, thanks to the rapid development of the field of autonomous driving, the visual 3D object detection algorithm based on surround multi-view has become a key algorithm for autonomous driving perception and has been widely studied by academia and industry. Specifically, the algorithm inputs multiple surround camera images and requires the output of the category, 3D center coordinates, size, yaw angle, speed and other information of all surrounding objects of interest to achieve 3D environment perception centered on the vehicle itself.
[0005] Some recent methods transform multi-view image features into bird's-eye view features, which can alleviate the problem of target occlusion. However, the complex perspective transformation operations and dense feature representations reduce the algorithm's reasoning speed, and it is difficult to balance the amount of computation and the perception range. Increasing the perception range is accompanied by a square increase in the amount of computation, and cannot adapt well to long-distance perception in open water environments.
[0006] Other methods extend the Transformer to target detection, cross-attention between sparse query and multi-view image features, and each target query will only predict one target, without the need to build dense bird's-eye view features, which improves the running speed. However, it cannot be well applied to surface target detection. This patent proposes a surface 3D target detection method based on multi-view spatiotemporal fusion, which makes full use of spatiotemporal context information, further improves the running efficiency, and avoids the negative impact of hull shaking on feature interaction. Summary of the invention
[0007] Aiming at the problems existing in the prior art, the purpose of the present invention is to provide a water surface 3D target detection method based on multi-view spatio-temporal fusion.
[0008] To achieve the above object of the invention, the technical solution adopted by the present invention is as follows:
[0009] The present invention provides a water surface 3D target detection method based on multi-view spatio-temporal fusion, comprising the following steps:
[0010] S1: Obtain multi-view images of the water surface at the same moment by using a device equipped with cameras with different perspectives;
[0011] S2: Input the target information features of the previous frame of multi-view images and the current frame of multi-view images into the feature compression backbone network. The feature compression backbone network extracts semantic features from the current frame of multi-view images based on the target information features of the previous frame of multi-view images, and obtains and outputs the feature map of the current frame of multi-view images;
[0012] S3: Input the target information features of the previous frame of multi-view images, the feature map of the current frame of multi-view images, and the current frame initialization target query into the decoder for decoding processing, and obtain and output the target embedding vector of the current frame of multi-view images;
[0013] S4: Send the target embedding vector of the current frame of multi-view images obtained in step S3 into the regression and classification head for processing, and obtain and output the target detection result; the target detection result includes target bounding box attributes and target categories;
[0014] S5: Input the target embedding vector of the current frame of multi-view images, the target bounding box attributes obtained in step S4, the pose matrix of the camera-equipped device, and the time interval between adjacent frames into the temporal memory and alignment module for motion alignment processing, and obtain and store the target information features of the multi-view images at the current moment for use in the target detection of the next frame of multi-view images.
[0015] According to the above water surface 3D target detection method, preferably, the feature compression backbone network includes an image preprocessing module, a first image encoding layer, a first fully connected layer, a second fully connected layer, and a second image encoding layer.
[0016] According to the above water surface 3D target detection method, preferably, the specific operation of step S2 is:
[0017] S21: Input the current-frame multi-view image into the image preprocessing module of the feature compression backbone network. The image preprocessing module divides each input image into image patches of a fixed size, and uses a linear projection layer to map each image patch into a high-dimensional vector to obtain the initial image embedding of each image patch. Then, add the initial image embedding of each image patch to the corresponding image patch position information to obtain and output the initial image embedding containing position information for each image patch.
[0018] S22: Input the initial image embedding containing position information for each image patch output in step S1 into the first image encoding layer. The first image encoding layer extracts features from the initial image embedding containing position information to obtain a first image embedding residual, and add the first image embedding residual to the initial image embedding containing position information to obtain and output the first updated image embedding for each image patch.
[0019] S23: Input the first updated image embedding for each image patch output in step S2 into the first fully-connected layer. The first fully-connected layer aligns the dimensions of the first updated image embedding with the target information features of the previous-frame multi-view image to obtain and output the dimension-aligned updated image embedding for each image patch.
[0020] S24: Perform matrix multiplication on the dimension-aligned updated image embedding for each image patch output in step S23 and the target information features of the previous-frame multi-view image to obtain an attention map. Input the attention map into the second fully-connected layer for feature extraction and regression processing to obtain and output the importance score S of the first updated image embedding for each image patch.
[0021] S25: Sort the first updated image embeddings for each image patch in descending order according to the importance score S. Denote the first updated image embeddings in the top P% after sorting as the first updated important image embeddings, and denote the remaining (100 - P)% of the first updated image embeddings as the first updated non-important image embeddings.
[0022] S26: Input the first updated important image embeddings into the second image encoding layer. The second image encoding layer extracts features from the first updated important image embeddings to obtain a second image embedding residual, and add the second image embedding residual to the first updated important image embeddings to obtain and output the second updated important image embedding for each image patch.
[0023] S27: Take the second updated important image embedding output from step S26 as the input, and repeat steps S23 - S26 until after repeating L1 times, output the updated important image embedding obtained after the L1 - th repetition and all the non - important image embeddings updated from the 1st to the L1 - th time; then, according to the image patch positions corresponding to the updated important image embedding obtained after the L1 - th repetition and all the non - important image embeddings updated from the 1st to the L1 - th time, rearrange the updated important image embedding obtained after the L1 - th repetition and all the non - important image embeddings updated from the 1st to the L1 - th time to obtain the current - frame multi - view image feature map and output it.
[0024] According to the above - mentioned water surface 3D target detection method, preferably, in step S24, the calculation formula for the matrix multiplication is as follows:
[0025]
[0026] In the formula, A is the attention map, is the dimension - aligned updated image embedding, is the feature of the previous - frame multi - view image target information, N is the number of pixels in the current - frame multi - view image feature map, N q is the number of targets of the target features in the previous - frame multi - view image, C q is the feature dimension number of the updated image embedding after dimension alignment.
[0027] According to the above - mentioned water surface 3D target detection method, preferably, the first image encoding layer is the same as the second image encoding layer, and both are composed of a multi - head self - attention network and a feed - forward network.
[0028] According to the above - mentioned water surface 3D target detection method, preferably, in step S3, the decoder is composed of a cross - attention module, a projection sampling interaction module, and a feed - forward network connected in sequence.
[0029] According to the above - mentioned water surface 3D target detection method, preferably, the specific operation of step S3 is as follows:
[0030] S31: Input the current - frame initialized target query and the previous - frame multi - view image target information feature into the cross - attention module for processing to obtain the first target query residual, and add the first target query residual to the current - frame initialized target query to obtain the first updated target query and output it;
[0031] S32: Input the current - frame multi - view feature map and the first updated target query output in step S31 into the projection interaction sampling module for projection interaction sampling processing to obtain the second target query residual, and add the second target query residual to the first target query residual to obtain the second updated target query and output it;
[0032] S33: Optimize and update the second updated target query output in step S32 by inputting it into the feedforward to obtain a third updated target query;
[0033] S34: Use the third updated target query output in step S33 as the input to the cross-attention module, and repeat steps S31 - S33 until repeated L2 times, then output the updated target query obtained after the L2-th repetition, denoted as the current frame target embedding vector.
[0034] According to the above-mentioned water surface 3D target detection method, preferably, the projection interaction sampling module includes three fully connected layers.
[0035] According to the above-mentioned water surface 3D target detection method, preferably, the specific operation of step S32 is as follows:
[0036] S321: Project the 3D reference coordinates corresponding to each target query onto the current frame multi-view image feature map to obtain 2D reference coordinates located on the image feature image planes of different views of the current frame and output them;
[0037] S322: Input the first updated target query into the first fully connected layer for processing to obtain multi-head multi-coordinate offsets; add the multi-head multi-coordinate offsets to the 2D reference coordinates of the target query to obtain multi-head multi-sampling points, and use the multi-head multi-sampling points to perform bilinear interpolation sampling on the current frame multi-view feature map to obtain multi-head multi-sampling features and output them;
[0038] S323: Input the first updated target query into the second fully connected layer for processing to obtain multiple attention score values, perform softmax processing on the attention score values to obtain multiple attention weights, and perform weighted summation of the attention weights and the multi-head multi-sampling features output in step S322 to obtain multi-head sampling features and output them;
[0039] S324: Input the multi-head sampling features output in step S323 into the third fully connected layer for projection processing to obtain sampling features; add each first updated target query to its corresponding sampling feature to obtain a second updated target query and output it.
[0040] According to the above-mentioned water surface 3D target detection method, preferably, in step S321, project the 3D reference coordinates corresponding to each target query onto the current frame multi-view image feature map according to Equation 2,
[0041] Equation 2 is as follows:
[0042] P 2D = K[R|t]P q Equation 2
[0043] wherein, R and t are respectively the rotation matrix and the translation vector for transforming the world coordinates to the camera coordinate system, and K is the camera internal parameter, are the 2D reference coordinates in the image plane.
[0044] According to the above-mentioned 3D water surface target detection method, preferably, in step S5, the calculation formula for motion alignment is shown in Equation 3:
[0045]
[0046] where Q e represents the target embedding vector of the current frame multi-view image, B p represents the target bounding box attributes, represents the pose transformation matrix from the coordinate system of the camera device in the previous frame to the coordinate system of the camera device in the current frame, Δt represents the time interval between adjacent frames, and MLP represents a multi-layer perceptron.
[0047] According to the above-mentioned 3D water surface target detection method, preferably, in steps S4 and S5, the target bounding box attributes include the position coordinates, size, direction, and speed of the target.
[0048] According to the above-mentioned 3D water surface target detection method, preferably, in step S3, the current frame initializes the target query as a set of learnable embedding vector parameters, and the current frame initializes the target query to be iteratively updated by the decoder.
[0049] According to the above-mentioned 3D water surface target detection method, preferably, in step S1, the device equipped with cameras with different perspectives is a ship.
[0050] According to the above-mentioned 3D water surface target detection method, preferably, in step S27, the number of L1 times is a preset value.
[0051] According to the above-mentioned 3D water surface target detection method, preferably, in step S34, the number of L2 times is a preset value.
[0052] Compared with the prior art, the positive and beneficial effects obtained by the present invention are as follows:
[0053] (1) The 3D water surface target detection method based on multi-view spatio-temporal fusion proposed by the present invention makes full use of spatio-temporal context information, fuses multi-view spatial features and historical target features, increases the detection range and improves the detection accuracy, and realizes the efficient and accurate detection of 3D water surface targets.
[0054] (2) The present invention proposes a feature compression backbone network. When using the feature compression backbone network to extract features from the current frame multi-view image, the present invention introduces the target information features of the previous frame multi-view image stored in the temporal memory and alignment module, and uses the target information features of the previous frame multi-view image to better predict the importance score, reduce the number of image embeddings participating in the calculation, and greatly improve the calculation efficiency.
[0055] (3) The decoder of the present invention contains a projection sampling interaction module. This projection sampling interaction module not only fuses multi-view features, but also the interaction method based on projection sampling avoids the spatial misalignment caused by the shaking of the camera lens, greatly improving the accuracy and precision of target detection. Brief Description of the Drawings
[0056] Figure 1 is the flow framework diagram of the water surface 3D target detection method based on multi-view spatio-temporal fusion of the present invention;
[0057] Figure 2 is the schematic diagram of the architecture of the feature compression backbone network in Embodiment 1 of the present invention;
[0058] Figure 3 is the schematic diagram of the network architecture of the decoder in Embodiment 1 of the present invention;
[0059] Figure 4 is the schematic diagram of the data processing process of the projection sampling interaction module in Embodiment 1 of the present invention;
[0060] Figure 5 is the schematic diagram of the data processing process of the temporal memory and alignment module. Detailed Embodiment
[0061] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.
[0062] The following further elaborates on the embodiments of the present invention with reference to the accompanying drawings.
[0063] Embodiment 1:
[0064] A water surface 3D target detection method based on multi-view spatio-temporal fusion, as Figure 1 shown, specifically includes the following steps:
[0065] S1: Use a device equipped with cameras with different perspectives to obtain multi-view images of the water surface at the same moment. Among them, the device can be a ship, and there are panoramic multi-cameras on the ship.
[0066] S2: Input the target information features of the previous frame of multi-view image and the current frame of multi-view image into the feature compression backbone network. The feature compression backbone network extracts semantic features from the current frame of multi-view image based on the target information features of the previous frame of multi-view image, and obtains and outputs the feature map of the current frame of multi-view image. Among them, the feature compression backbone network (as Figure 2 shown) includes an image preprocessing module, a first image encoding layer, a first fully connected layer, a second fully connected layer, and a second image encoding layer. The first image encoding layer is the same as the second image encoding layer, and both are composed of a multi-head self-attention network and a feed-forward network; the first connection layer is the same as the second connection layer.
[0067] S3: Input the target information features of the previous frame of multi-view image, the feature map of the current frame of multi-view image, and the current frame initialization target query into the decoder for decoding processing, and obtain and output the target embedding vector of the current frame of multi-view image. Among them, the current frame initialization target query is a set of learnable embedding vector parameters, and the current frame initialization target query is iteratively updated after being input into the decoder. The decoder is composed of a cross-attention module, a projection sampling interaction module, and a feed-forward network connected in sequence (as Figure 3 shown).
[0068] S4: Input the target embedding vector of the current frame of multi-view image obtained in step S3 into the regression and classification head for processing, and obtain and output the target detection result; the target detection result includes target bounding box attributes and target categories; the target bounding box attributes include the position coordinates, size, direction, and speed of the target.
[0069] S5: Input the target embedding vector of the current frame of multi-view image, the target bounding box attributes obtained in step S4, the pose matrix of the camera-equipped device, and the time interval between adjacent frames into the temporal memory and alignment module for motion alignment processing, and obtain and store the target information features of the multi-view image at the current moment for input to the feature compression backbone network and decoder during the target detection of the next frame of multi-view image.
[0070] Among them, as a preferred implementation manner, the specific operation of the above step S2 is (as Figure 2 shown):
[0071] S21: Input the current-frame multi-view image into the image preprocessing module of the feature compression backbone network. The image preprocessing module divides each input image into image patches of a fixed size, and uses a linear projection layer to map each image patch into a high-dimensional vector to obtain the initial image embedding of each image patch. Then, add the initial image embedding of each image patch to the corresponding image patch position information to obtain and output the initial image embedding containing position information of each image patch.
[0072] S22: Input the initial image embedding containing position information of each image patch output in step S1 into the first image encoding layer. The first image encoding layer extracts features from the initial image embedding containing position information to obtain a first image embedding residual, and add the first image embedding residual to the initial image embedding containing position information to obtain and output the first updated image embedding of each image patch.
[0073] S23: Input the first updated image embedding of each image patch output in step S2 into the first fully-connected layer. The first fully-connected layer aligns the dimensions of the first updated image embedding with the target information features of the previous-frame multi-view image to obtain and output the dimension-aligned updated image embedding of each image patch.
[0074] S24: Perform matrix multiplication calculation on the dimension-aligned updated image embedding of each image patch output in step S23 and the target information features of the previous-frame multi-view image according to Equation 1 to obtain an attention map. Input the attention map into the second fully-connected layer for feature extraction and regression processing to obtain and output the importance score S of the first updated image embedding of each image patch.
[0075] Equation 1 is specifically as follows:
[0076]
[0077] In the formula, A is the attention map, is the dimension-aligned updated image embedding, is the feature of the target information of the previous-frame multi-view image, N is the number of pixels in the feature map of the current-frame multi-view image, N q is the number of targets of the target features in the previous-frame multi-view image, C q is the feature dimension number of the updated image embedding after dimension alignment.
[0078] S25: Sort the first updated image embeddings of each image patch in descending order according to the importance score S. Denote the first updated image embeddings in the top P% after sorting as the first updated important image embeddings, and denote the remaining (100 - P)% of the first updated image embeddings as the first updated unimportant image embeddings.
[0079] S26: Embed the first-updated important image into the input second image encoding layer. The second image encoding layer extracts features from the embedding of the first-updated important image to obtain a second image embedding residual, and adds the second image embedding residual to the embedding of the first-updated important image to obtain the second-updated important image embedding for each image block and output it.
[0080] S27: Use the second-updated important image embedding output in step S26 as the input, and repeat steps S23 - S26 until after repeating L1 times (L1 is a preset number of times), output the updated important image embedding obtained after the L1th repetition and all non-important image embeddings updated from the 1st to the L1th time; then, according to the image block positions corresponding to the updated important image embedding obtained after the L1th repetition and all non-important image embeddings updated from the 1st to the L1th time, rearrange the updated important image embedding obtained after the L1th repetition and all non-important image embeddings updated from the 1st to the L1th time to obtain the current frame multi-view image feature map and output it.
[0081] As a preferred implementation manner, the specific operation of step S3 above is (as Figure 3 shown):
[0082] S31: Input the current frame initialization target query and the previous frame multi-view image target information features into the cross-attention module for processing to obtain a first target query residual, and add the first target query residual to the current frame initialization target query to obtain the first-updated target query and output it; the current frame initialization target query is a set of learnable embedding vector parameters.
[0083] S32: Input the current frame multi-view feature map and the first-updated target query output in step S31 into the projection interaction sampling module for projection interaction sampling processing to obtain a second target query residual, and add the second target query residual to the first target query residual to obtain the second-updated target query and output it.
[0084] S33: Input the second-updated target query output in step S32 into the feed-forward for optimization and update to obtain the third-updated target query.
[0085] S34: Use the third-updated target query output in step S33 as the input of the cross-attention module, and repeat steps S31 - S33 until after repeating L2 times, output the updated target query obtained after the L2th repetition, denoted as the current frame target embedding vector.
[0086] As a preferred implementation manner, the projection interaction sampling module includes three fully connected layers. The specific operation of step S32 above is:
[0087] S321: Project the 3D reference coordinates corresponding to each target query onto the current-frame multi-view image feature map according to Equation 2, obtain the 2D reference coordinates on the image feature image planes of different views in the current frame, and output them;
[0088] Equation 2 is as follows:
[0089] P 2D = K[R|t}P q Equation 2
[0090] In the formula, R and t are respectively the rotation matrix and translation vector for transforming the world coordinates to the camera coordinate system, and K is the camera internal parameter. is the 2D reference coordinate on the image plane;
[0091] S322: Input the first-updated target query into the first fully-connected layer for processing to obtain multi-head multi-coordinate offsets; add the multi-head multi-coordinate offsets to the 2D reference coordinates of the target query to obtain multi-head multi-sampling points, and perform bilinear interpolation sampling on the current-frame multi-view feature map using the multi-head multi-sampling points to obtain multi-head multi-sampling features and output them;
[0092] S323: Input the first-updated target query into the second fully-connected layer for processing to obtain multiple attention score values, perform softmax processing on the attention score values to obtain multiple attention weights, and perform weighted summation on the multi-head multi-sampling features output in step S322 using the attention weights to obtain multi-head sampling features and output them;
[0093] S324: Input the multi-head sampling features output in step S323 into the third fully-connected layer for projection processing to obtain sampling features; add each first-updated target query to its corresponding sampling feature to obtain the second-updated target query and output it.
[0094] As a preferred implementation manner, in the above step S5, input the target embedding vector of the current-frame multi-view image, the target bounding box attributes obtained in step S4, the pose matrix of the camera device, and the time interval between adjacent frames into the temporal memory and alignment module for motion alignment processing according to Equation 3 (as Figure 5 shown), and after the motion alignment processing, the target embedding vector of the current-frame multi-view image is obtained as the target information feature of the current-frame multi-view image and stored for use as the input to the feature compression backbone network and decoder during the target detection of the next-frame multi-view image. Among them, Equation 3 is as follows:
[0095]
[0096] Among them, Q e represents the target embedding vector of the current-frame multi-view image, and B pRepresents the target bounding box attributes, represents the pose transformation matrix from the coordinate system of the camera device carried in the previous frame to the coordinate system of the camera device in the current frame, Δt represents the time interval between adjacent frames, and MLP represents a multi-layer perceptron.
[0097] Verification of the target detection effect of the 3D water surface target detection method based on multi-view spatio-temporal fusion described in Embodiment 1 of the present invention:
[0098] The verification experiment steps are as follows:
[0099] 1) Make a 3D water surface target detection dataset (the 3D water surface target detection dataset includes multiple consecutive multi-view images, a timestamp sequence, a pose transformation matrix sequence of the camera platform carried, and 3D annotation information related to the target object in the multi-view images). Randomly select 10% of the samples from the 3D water surface target detection dataset as the test set, and the remaining 90% of the samples as the training set.
[0100] 2) Select the PETR (Position Embedding Transformation for Multi-View 3D Object Detection) model as the baseline model, and apply the 3D water surface target detection method based on multi-view spatio-temporal fusion described in Embodiment 1 of the present invention to the baseline model to obtain an improved model. Then, use the training set to train the baseline model and the improved model respectively.
[0101] 3) Use the test set to test the trained baseline model and the trained improved model, and use mAP (mean Average Precision) as the evaluation index to quantitatively analyze the detection performance of the trained baseline model and the trained improved model.
[0102] Verification experiment results: It is found through verification by the test set that the improved model equipped with the 3D water surface target detection method based on multi-view spatio-temporal fusion described in Embodiment 1 of the present invention shows high detection accuracy in multiple scenarios; moreover, the improved model equipped with the 3D water surface target detection method based on multi-view spatio-temporal fusion described in Embodiment 1 of the present invention achieves a mAP of 68.5% on the test set, while the mAP of the baseline model PETR is only 41.2%; therefore, compared with the existing PETR model, the 3D water surface target detection method based on multi-view spatio-temporal fusion of the present invention has a significant improvement in the mAP index, indicating the reliability and superiority of the 3D water surface target detection method of the present invention in practical applications.
[0103] Finally, it should be noted that the above embodiments are only preferred embodiments of the present invention, and are not intended to limit the present invention in other forms. Any person skilled in the relevant art may make changes or modifications using the above technical content as inspiration. These equivalent embodiments with equivalent changes. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the technical concept of the present invention still fall within the protection scope of the claims of the present invention.
Claims
1. A method for detecting 3D objects on a water surface based on multi-view spatiotemporal fusion, characterized in that: The following steps are involved: S1: Use a device equipped with cameras with different viewing angles to obtain multiple view images of the water surface at the same time; S2: Input the target information features of the previous multi-view image and the current multi-view image into the feature compression backbone network. The feature compression backbone network extracts semantic features of the current multi-view image based on the target information features of the previous multi-view image, obtains the feature map of the current multi-view image and outputs it; S3: Input the target information features of the previous multi-view image, the feature map of the current multi-view image, and the initialization target query of the current frame into the decoder for decoding processing, obtain the target embedding vector of the current multi-view image and output it; S4: sending the target embedding vector of the multi-view image of the current frame obtained in step S3 to the regression and classification head for processing, obtaining and outputting the target detection result; the target detection result includes the target bounding box attributes and the target category; S5: Input the target embedding vector of the multi-view image of the current frame, the target bounding box attributes obtained in step S4, the pose matrix of the camera device and the time interval between adjacent frames into the timing memory and alignment module for motion alignment processing to obtain and store the target information features of the multi-view image at the current moment.
2. The method for detecting 3D objects on a water surface according to claim 1, characterized in that: The feature compression backbone network includes an image preprocessing module, a first image coding layer, a first fully connected layer, a second fully connected layer, and a second image coding layer; The specific operations of step S2 are: S21: Input the multi-view image of the current frame into the image preprocessing module of the feature compression backbone network, wherein the image preprocessing module divides each input image into image blocks of a fixed size, and uses a linear projection layer to map each image block to a high-dimensional vector, thereby obtaining an initial image embedding of each image block; then, the initial image embedding of each image block is added to the corresponding image block position information, thereby obtaining and outputting the initial image embedding of each image block including the position information; S22: inputting the initial image embedding containing position information of each image block outputted in step S1 into the first image coding layer, the first image coding layer extracting features from the initial image embedding containing position information to obtain a first image embedding residual, and adding the first image embedding residual to the initial image embedding containing position information to obtain a first updated image embedding of each image block and outputting the first updated image embedding; S23: inputting the first updated image embedding of each image block outputted in step S2 into the first fully connected layer, the first fully connected layer dimensionally aligning the first updated image embedding with the target information feature of the previous frame of multi-view image, obtaining the dimensionally aligned updated image embedding of each image block and outputting it; S24: Perform matrix multiplication calculation on the dimensionally aligned updated image embedding of each image block output in step S23 and the target information feature of the previous frame of multi-view image to obtain an attention map; input the attention map into the second fully connected layer for feature extraction and regression processing, obtain the importance score S of the first updated image embedding of each image block and output it; S25: Arrange the first updated image embeddings of each image block in descending order according to the importance score S, record the first updated image embeddings in the top P% after arrangement as first updated important image embeddings, and record the remaining (100-P)% of the first updated image embeddings as first updated unimportant image embeddings; S26: inputting the first updated important image embedding into the second image coding layer, the second image coding layer extracting features from the first updated important image embedding to obtain a second image embedding residual, and adding the second image embedding residual to the first updated important image embedding to obtain and output a second updated important image embedding for each image block; S27: Take the second updated important image embedding output by step S26 as input, repeat steps S23 to S26 until it is repeated L1 times, and then output the updated important image embedding obtained after the L1th repetition and all the non-important image embeddings updated from the 1st to the L1th; then, according to the image block positions corresponding to the updated important image embedding obtained after the L1th repetition and all the non-important image embeddings updated from the 1st to the L1th, rearrange the updated important image embedding obtained after the L1th repetition and all the non-important image embeddings updated from the 1st to the L1th to obtain the multi-view image feature map of the current frame and output it.
3. The method for detecting 3D targets on a water surface according to claim 2, characterized in that: In step S24, the calculation formula of the matrix multiplication calculation is as follows: Where A is the attention map, is the dimensionally aligned updated image embedding, is the feature of the target information of the previous multi-view image, N is the number of pixels of the feature map of the current multi-view image, N q is the number of target features in the previous multi-view image, C q is the number of feature dimensions of the updated image embedding after dimension alignment.
4. The method for detecting 3D objects on the water surface according to claim 3, characterized in that: The first image coding layer and the second image coding layer are both composed of a multi-head self-attention network and a feedforward network.
5. The method for detecting 3D objects on a water surface according to claim 1, characterized in that: In step S3, the decoder is composed of a cross-attention module, a projection sampling interaction module and a feedforward network connected in sequence.
6. The method for detecting 3D targets on a water surface according to claim 5, characterized in that: The specific operations of step S3 are: S31: inputting the current frame initialization target query and the target information features of the previous frame multi-view image into the cross attention module for processing to obtain a first target query residual, and adding the first target query residual to the current frame initialization target query to obtain a first updated target query and output it; S32: Input the multi-view feature map of the current frame and the first updated target query outputted in step S31 into the projection interaction sampling module for projection interaction sampling processing to obtain a second target query residual, and add the second target query residual to the first target query residual to obtain a second updated target query and output it; S33: inputting the second updated target query outputted from step S32 into the feedforward for optimization and updating to obtain a third updated target query; S34: Use the third updated target query output from step S33 as the input of the cross-attention module, repeat steps S31 to S33 until L2 times, and output the updated target query obtained after the L2th repetition, which is recorded as the current frame target embedding vector.
7. The method for detecting 3D objects on a water surface according to claim 6, characterized in that: The projection interaction sampling module includes three fully connected layers. The specific operation of step S32 is: S321: Project the 3D reference coordinates corresponding to each target query onto the multi-view image feature map of the current frame, obtain the 2D reference coordinates of the image feature image planes located in different views of the current frame and output them; S322: Input the first updated target query into the first fully connected layer for processing to obtain multi-head multi-coordinate offset; Adding the multi-head multi-coordinate offset to the 2D reference coordinate of the target query to obtain multi-head multi-sampling points, using the multi-head multi-sampling points to perform bilinear interpolation sampling on the multi-view feature map of the current frame to obtain multi-head multi-sampling features and output them; S323: Input the first updated target query into the second fully connected layer for processing to obtain multiple attention score values, and obtain multiple attention weights by softmax processing of the attention score values. Perform weighted summation of the attention weights and the multi-head multi-sampling features output in step S322 to obtain the multi-head sampling features and output them; S324: Input the multi-head sampling features outputted in step S323 into the third fully connected layer for projection processing to obtain sampling features; Each first updated target query is added to its corresponding sampled feature to obtain the second updated target query and output it.
8. The method for detecting 3D objects on a water surface according to claim 1, characterized in that: In step S5, the calculation formula of the motion alignment is as follows: Among them, Q e represents the target embedding vector of the multi-view image of the current frame, B p represents the target bounding box attributes, represents the pose transformation matrix from the coordinate system of the camera device in the previous frame to the coordinate system of the camera device in the current frame, Δt represents the time interval between adjacent frames, and MLP represents a multi-layer perceptron.