Image-based Unmanned Boat Control Method, System, Device, Product and Medium
By installing cameras and navigation radar on unmanned boats, the fusion processing of visual images and radar images is solved, and the problem of unmanned boats being difficult to track targets and autonomous navigation when they cannot receive external signals is achieved, and autonomous tracking and heading control of unmanned boats is realized.
Patent Information
- Application Number
- CN202510007442.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-01-03
AI Technical Summary
Unmanned boats have difficulty in achieving continuous tracking and autonomous navigation of targets when GPS systems and other external signals are not received, especially in complex and changing marine environments.
By installing cameras and navigation radar on the unmanned boat, visual images and radar images are used for fusion processing, spatial characteristics and target characteristics are extracted, and navigation images are generated, thereby calculating the azimuth difference of markers and generating navigation control signals.
It realizes that unmanned boats can independently track targets and perform heading control when they cannot receive external navigation signals, ensuring mission execution and safety of unmanned boats.
Smart Images

Figure CN119396164B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of unmanned boat control, and particularly to an image-based unmanned boat control method, system, device, product and medium. Background Art
[0002] In recent years, the topic of using visual feedback to control robotic systems such as mobile robots and robotic arms has become an important topic in the field of control engineering, and image-based control is one of the hottest topics among them. The visual servo control of unmanned boats has broad application prospects and can be applied to many fields such as military, marine resource exploration, environmental sampling and monitoring. Unmanned boats can be used for rescue operations and detecting irregularly moving boats near ports. In order to perform these tasks in a complex and changing marine environment, an advanced sensing system is needed to accurately capture the details of the surrounding environment to ensure the unmanned boat's tracking of the target and its own navigation control. When the GPS and other sensors fail or are restricted in operation, such as in shallow water areas, sewers or areas with poor satellite coverage, the vision-based motion control system can ensure the execution of tasks and the safety of the unmanned boat. In contrast, although the inertial navigation system does not need to rely on the external environment, it will have problems such as drift during operation, which will cause deviation of the unmanned boat's heading, and the inertial navigation system cannot achieve tracking of a specified target. Summary of the Invention
[0003] The present invention aims to solve at least one of the technical problems existing in the related art. For this reason, the present invention provides an image-based unmanned boat control method, which can still control the unmanned boat to continuously track a certain target or autonomously control the navigation of the unmanned boat when the GPS system, remote control equipment, etc. cannot receive external signals.
[0004] The present invention provides an image-based unmanned boat control method, including:
[0005] S1: Arrange a camera and a navigation radar on the unmanned boat, determine the shooting area and shoot the shooting area through the camera to obtain an image to be processed, perform Fourier transform and adaptive feature extraction on the image to be processed to obtain a target amplitude spectrum, and obtain spatial features through the target amplitude spectrum;
[0006] S2: Perform global grid filling on the image to be processed to obtain a filled image, perform defogging processing on the filled image to obtain an initial defogged image, and perform reverse filling on the initial defogged image to obtain a target defogged image;
[0007] S3: Fuse the spatial features with the target defogged image to obtain an image to be extracted, perform feature extraction on the image to be extracted to obtain a target extracted image;
[0008] S4: Detect through the navigation radar to obtain an initial radar image, perform spatial folding on the initial radar image and calculate the channel weights to obtain a channel-weighted radar image, calculate the spatial weighting matrix of the channel-weighted radar image, and obtain a target radar image through the spatial weighting matrix;
[0009] S5: Fuse the target radar image with the target extraction image to obtain a navigation image, calculate the azimuth difference of the marker through the navigation image, and obtain a navigation control signal according to the azimuth difference of the marker. The unmanned boat completes the heading control according to the navigation control signal.
[0010] According to the image-based unmanned boat control method provided by the present invention, step S1 specifically includes:
[0011] S11: Determine the installation conditions of the camera and the navigation radar and the equipment environment of the unmanned boat, install the camera and the navigation radar on the unmanned boat according to the installation conditions and the equipment environment, determine the shooting area and shoot the shooting area through the camera to obtain an image to be processed;
[0012] S12: Obtain the channel dimension of the image to be processed, perform block processing on the image to be processed according to the channel dimension, and perform average pooling processing and maximum pooling processing on the image to be processed to obtain a pooled image;
[0013] S13: Perform Fourier transform on the pooled image to obtain an amplitude spectrum, perform sub-band feature extraction on the amplitude spectrum to obtain spatial adaptive features, perform rearrangement and stacking on the spatial adaptive features to obtain a target amplitude spectrum, and use the minimax algorithm to obtain the spatial features through the target amplitude spectrum.
[0014] According to the image-based unmanned boat control method provided by the present invention, step S2 specifically includes:
[0015] S21: Perform edge filling and global feature grid filling on the image to be processed to obtain a filled image;
[0016] S22: Select an initial defogging algorithm, change the internal parameters of the defogging algorithm of the initial defogging algorithm and calculate the defogging loss of the initial defogging algorithm, retain the internal parameters of the defogging algorithm when the defogging loss is the smallest to obtain the defogging algorithm, and perform defogging processing on the filled image through the defogging algorithm to obtain the initial defogged image;
[0017] S23: Perform reverse edge filling and reverse global filling on the initial defogged image to obtain the target defogged image.
[0018] According to the image-based unmanned boat control method provided by the present invention, step S3 specifically includes:
[0019] S31: Fuse the spatial features with the target dehazed image to obtain an image to be extracted, select an initial feature extraction algorithm, change the internal parameters of the feature extraction algorithm of the initial feature extraction algorithm, and calculate the classification error and bounding box regression error of the initial feature extraction algorithm;
[0020] S32: Calculate the loss error of the initial feature extraction algorithm through the classification error and the bounding box regression error of the initial feature extraction algorithm, retain the internal parameters of the feature extraction algorithm when the loss error is the smallest to obtain a feature extraction algorithm, and perform feature extraction on the image to be extracted through the feature extraction algorithm to obtain the target extraction image.
[0021] According to the image-based unmanned boat control method provided by the present invention, step S4 specifically includes:
[0022] S41: Detect the shooting area of the camera through a navigation radar to obtain an initial radar image;
[0023] S42: Perform linear projection on the initial radar image to obtain a linearly projected radar image, fold the spatial dimension of the linearly projected radar image to obtain a spatially folded radar image, calculate the channel weight of the spatially folded radar image to obtain a channel weighted matrix feature map, and obtain the channel weighted radar image through the channel information of the spatially folded radar image and the channel weighted matrix feature map;
[0024] S43: Activate and convolve the channel weighted radar image to obtain the spatial weighted matrix, multiply the spatial weighted matrix element by element with the linearly projected radar image to obtain the target radar image.
[0025] According to the image-based unmanned boat control method provided by the present invention, the camera is an RGBD camera, and the image depth information is included in the image to be processed captured by the RGBD camera.
[0026] The present invention also provides an image-based unmanned boat control system, including:
[0027] Spatial feature acquisition module: used to arrange a camera and a navigation radar on an unmanned boat, determine a shooting area, shoot the shooting area through the camera, obtain an image to be processed, perform Fourier transform and adaptive feature extraction on the image to be processed to obtain a target amplitude spectrum, and obtain spatial features through the target amplitude spectrum;
[0028] Image dehazing module: used to perform global grid filling on the to-be-processed image to obtain a filled image, perform dehazing processing on the filled image to obtain an initial dehazed image, and perform reverse filling on the initial dehazed image to obtain a target dehazed image;
[0029] Feature extraction module: used to fuse the spatial feature with the target dehazed image to obtain an image to be extracted, and perform feature extraction on the image to be extracted to obtain a target extracted image;
[0030] Radar image acquisition module: used to detect through a navigation radar to obtain an initial radar image, perform spatial folding on the initial radar image and calculate channel weights to obtain a channel-weighted radar image, calculate the spatial weighting matrix of the channel-weighted radar image, and obtain a target radar image through the spatial weighting matrix;
[0031] Unmanned boat control module: used to fuse the target radar image with the target extracted image to obtain a navigation image, calculate the azimuth difference of the landmark through the navigation image, and obtain a navigation control signal according to the azimuth difference of the landmark. The unmanned boat completes the heading control according to the navigation control signal.
[0032] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the image-based unmanned boat control method as described in any one of the above are implemented.
[0033] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the image-based unmanned boat control method as described in any one of the above are implemented.
[0034] The present invention also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the steps of the image-based unmanned boat control method as described in any one of the above.
[0035] One or more of the above technical solutions in the embodiments of the present invention have at least one of the following technical effects:
[0036] The image-based unmanned boat control method, system, device, product and medium provided by the present invention control the unmanned boat through visual images and radar images, and can perform target tracking and unmanned boat heading control in extreme situations where the coastal visibility is reduced or the unmanned boat cannot receive navigation signals from the outside world. At the same time, an artificial intelligence algorithm is introduced. During the target detection process, more prominent image features are obtained by performing Fourier transform on the image features and adaptive feature extraction. During the dehazing process, the dehazing effect is improved by screening the dehazing model. At the same time, the radar image is weighted and the visual image and the radar image are fused to obtain a navigation control signal and control the heading of the unmanned boat.
[0037] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. Description of the Drawings
[0038] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0039] Figure 1 It is a schematic flowchart of the image-based unmanned boat control method provided by the present invention.
[0040] Figure 2 It is a schematic structural diagram of the image-based unmanned boat control system provided by the present invention.
[0041] Figure 3 It is a schematic structural diagram of the image-based unmanned boat control device provided by the present invention.
[0042] Reference Signs:
[0043] 100, Spatial Feature Acquisition Module; 200, Image Dehazing Module; 300, Feature Extraction Module; 400, Radar Image Acquisition Module; 500, Unmanned Boat Control Module; 810, Processor; 820, Communication Interface; 830, Memory; 840, Communication Bus. Detailed Embodiments
[0044] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention. The following embodiments are used to illustrate the present invention, but shall not be used to limit the scope of the present invention.
[0045] In the description of the embodiments of the present invention, it should be noted that the terms "first", "second", and "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.
[0046] In the description of the embodiments of the present invention, it should be noted that unless otherwise clearly specified and defined, the terms "connected" and "coupled" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms in the embodiments of the present invention can be understood according to specific circumstances.
[0047] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the embodiments of the present invention. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0048] The following is combined with Figures 1 to 3 to describe the implementation scheme of the present invention:
[0049] Figure 1 It is a schematic flow chart of an image-based unmanned boat control method. Among them, S1 is to arrange a camera and a navigation radar, obtain the image to be processed, and obtain spatial features through the image to be processed; S2 is to perform global grid filling and defogging processing on the image to be processed to obtain a target defogged image; S3 is to obtain the image to be extracted, perform feature extraction on the image to be extracted, so as to obtain a target extracted image; S4 is to detect through the navigation radar and perform channel weighting and spatial weighting to obtain a target extracted image; S5 is to obtain a navigation image, calculate the azimuth difference of the landmark through the navigation image, and finally complete the control of the unmanned boat.
[0050] The image-based unmanned boat control method provided by the present invention includes:
[0051] S1: Arrange a camera and a navigation radar on the unmanned boat, determine the shooting area, shoot the shooting area through the camera to obtain an image to be processed, perform Fourier transform and adaptive feature extraction on the image to be processed to obtain a target amplitude spectrum, and obtain spatial features through the target amplitude spectrum;
[0052] Further, the purpose of this stage is to deploy the navigation radar and the camera, obtain the image to be processed through the camera, and finally obtain the spatial features of the image to be processed. Step S1 specifically includes:
[0053] S11: Determine the installation conditions of the camera and the navigation radar and the equipment environment of the unmanned boat, install the camera and the navigation radar on the unmanned boat according to the installation conditions and the equipment environment, determine the shooting area, and shoot the shooting area through the camera to obtain an image to be processed;
[0054] S12: Obtain the channel dimension of the image to be processed, perform block processing on the image to be processed according to the channel dimension, and perform average pooling processing and max pooling processing on the image to be processed to obtain a pooled image;
[0055] S13: Perform Fourier transform on the pooled image to obtain an amplitude spectrum, perform sub-band feature extraction on the amplitude spectrum to obtain spatial adaptive features, rearrange and stack the spatial adaptive features to obtain a target amplitude spectrum, and use the minimax algorithm to obtain the spatial features through the target amplitude spectrum.
[0056] For the above steps, the specific implementation manners in this embodiment are as follows:
[0057] During the working process of the unmanned boat, it may encounter some extreme situations, such as thunderstorm weather, strong electromagnetic interference, etc., which will greatly deteriorate the electromagnetic environment where the unmanned boat is located, so that it cannot normally receive external GPS satellite signals, remote control signals, etc. for its own navigation and control. At this time, although it can rely on the inertial navigation system, on the one hand, the inertial navigation system is prone to drift, and on the other hand, it cannot achieve target tracking. Therefore, it is necessary to perform its navigation control through visual images and radar images.
[0058] First, determine the installation conditions for the camera and the navigation radar. Under normal circumstances, the installation conditions for both the camera and the navigation radar require that the field of view in their detection directions should be as open as possible, the position should be as high as possible and there should be no obstacles, and for the navigation radar, there should be no other electronic devices nearby that can interfere with its signal. For the camera, when multiple cameras are installed on the unmanned boat, the cameras should be at the same height as much as possible to avoid introducing additional errors into the calculation of spatial features due to the height difference between different cameras when they work together. The equipment environment is the working environment provided for various equipment by the positions reserved on the unmanned boat for installing equipment, and it needs to be jointly determined according to the power supply situation, wiring situation, and orientation of the reserved position of the unmanned boat. Provide the best possible equipment environment for the camera and the navigation radar on the premise of meeting the installation conditions. Then, install the camera and the navigation radar on the unmanned boat according to the installation conditions and the equipment environment. Here, the camera can be an RGBD (RGB Depth map) camera. The images captured by the RGBD camera contain relatively rich depth information of the images. Therefore, the images to be processed obtained by shooting include relatively rich depth information of the images. In this way, the spatial features extracted from the images to be processed also include this depth information, that is, the distance between the objects in the images and the camera. Determine the shooting area according to the perspective of the camera in the surrounding environment where the unmanned boat is located and take pictures, and judge whether there are tracking targets or fixed reference objects that can provide a basis for navigation, such as lighthouses, shore walls, etc. in the captured images, and use the images including the tracking targets or fixed reference objects as the images to be processed.
[0059] Next, obtain the channel dimension of the image to be processed, and divide the image to be processed according to the channel dimension. In this embodiment, the image content with the same channel dimension is divided into the same first image block. Integrate the obtained first image blocks into a first image block set and perform average pooling processing and max pooling processing to obtain a pooled image. :
[0060]
[0061] Among them, represents performing average pooling processing, represents performing max pooling processing, and F represents the first image block set.
[0062] Next, perform a two-dimensional Fourier transform on the pooled image, and obtain the amplitude spectrum and phase spectrum of the pooled image. For the amplitude spectrum, then perform adaptive feature extraction, including sub-band feature extraction and calculation of spatial adaptive features and arranging and stacking them. Perform sub-band feature extraction, that is, first determine a frequency range, and use the amplitude spectrum within this frequency range as the frequency domain feature. , multiple non-overlapping frequency intervals are taken on the amplitude spectrum to obtain multiple amplitude spectra, so as to obtain multiple frequency-domain features. Subsequently, the spatial adaptive features of each frequency-domain feature are calculated :
[0063]
[0064] Among them, represents element-wise broadcast multiplication, and FC() represents activating the content within the parentheses using the LeakyReLU activation function. Then, all the obtained spatial adaptive features are arranged and stacked to obtain the target spectral amplitude , and finally the spatial features are calculated :
[0065]
[0066] Among them, MinMax() represents the min-max algorithm, represents performing Fourier transform and taking the reciprocal of the result of the Fourier transform, represents pooling the phase spectrum of the image. The spatial features include the spatial information contained in the image to be processed, that is, the distance, azimuth, etc. of the object in the image to be processed relative to the camera, which can be regarded as relative to the unmanned boat in this embodiment.
[0067] S2: Perform global grid filling on the image to be processed to obtain a filled image, perform defogging processing on the filled image to obtain an initial defogged image, and perform reverse filling on the initial defogged image to obtain a target defogged image;
[0068] Furthermore, the purpose of this stage is to perform defogging processing on the image to be processed to obtain an initial defogged image, and finally obtain a target defogged image for subsequent feature extraction of the image. Step S2 specifically includes:
[0069] S21: Perform edge filling and global feature grid filling on the image to be processed to obtain a filled image;
[0070] S22: Select an initial defogging algorithm, change the internal parameters of the defogging algorithm of the initial defogging algorithm and calculate the defogging loss of the initial defogging algorithm, retain the internal parameters of the defogging algorithm when the defogging loss is the smallest to obtain the defogging algorithm, and perform defogging processing on the filled image through the defogging algorithm to obtain the initial defogged image;
[0071] S23: Perform reverse edge filling and reverse global filling on the initial defogged image to obtain the target defogged image.
[0072] For the above steps, the specific implementation methods in this embodiment are as follows:
[0073] First, perform edge padding on the image to be processed. This is because the edges of the image to be processed are usually relatively blurred, which may affect the subsequent defogging process. Therefore, it is necessary to first perform image padding on the edges of the image to complete the edge padding. Then, perform global feature grid padding on the image after edge padding, that is, first determine the segmentation coefficient G of the image, and then segment the image according to the segmentation coefficient. The segmented image is divided into N second image blocks. The calculation method of N is as follows:
[0074]
[0075] where H is the height of the image after edge padding, and W is the width of the image after edge padding. Then, in each second image block, first set the pixel values of all pixels to zero, and then perform pixel swapping on the image within the block:
[0076]
[0077] where, is the pixel value of the pixel at the th row and th column in the second image block after pixel swapping, is the pixel value of the pixel at the th row and th column in the second image block before the pixel value is set to zero. In this way, the global feature grid padding is completed, and the padded image is obtained.
[0078] Then, select the initial defogging algorithm. In this embodiment, the initial defogging algorithm is the Uformer model. Subsequently, change the internal parameters of the defogging algorithm of the initial defogging algorithm multiple times. After each change of the internal parameters of the defogging algorithm, input the same training set including an equal number of images before defogging and images after defogging into the initial defogging algorithm, and defog the images before defogging through the initial defogging algorithm to generate test defogged images. Calculate the defogging loss L of the initial defogging algorithm at this time:
[0079]
[0080] where, is the mean square error between the test defogged image and the image after defogging, is the mean square error between the test defogged image and the image after defogging after calculation by the Laplacian pyramid, is a stability constant, which is taken as 10 -3 , where:
[0081]
[0082]
[0083] Here, n is the number of defogged images, denotes the feature vector of the k-th test defogged image generated by the initial defogging algorithm, denotes the feature vector of the k-th defogged image, denotes the product of the elements of the feature vector within the parentheses. Retain the Uformer model with the minimum defogging loss under the condition of a set of internal parameters of the defogging algorithm, and the defogging algorithm can be obtained. By using the defogging algorithm to defog the filled image, the initial defogged image can be obtained.
[0084] Next, since the initial defogged image has undergone edge padding and global feature grid padding, when obtaining the target defogged image, it is necessary to perform reverse padding on the initial defogged image, that is, remove the padded part in the initial defogged image to avoid its influence on the subsequent image to be extracted. The reverse padding includes reverse edge padding and reverse global padding. Reverse edge padding means removing the padding performed on the image edge during the edge padding process, and reverse global padding means returning the exchanged pixels in the global feature grid padding to their original positions. In this way, the target defogged image is obtained.
[0085] S3: Fuse the spatial feature with the target defogged image to obtain the image to be extracted, and perform feature extraction on the image to be extracted to obtain the target extraction image;
[0086] Furthermore, the purpose of this stage is to fuse the spatial feature with the defogged image to obtain the image to be extracted, and perform feature extraction on the image to be extracted to obtain the target extraction image. Step S3 specifically includes:
[0087] S31: Fuse the spatial feature with the target defogged image to obtain the image to be extracted, select the initial feature extraction algorithm, change the internal parameters of the feature extraction algorithm of the initial feature extraction algorithm, and calculate the classification error and bounding box regression error of the initial feature extraction algorithm;
[0088] S32: Calculate the loss error of the initial feature extraction algorithm through the classification error and the bounding box regression error of the initial feature extraction algorithm, retain the internal parameters of the feature extraction algorithm when the loss error is the smallest to obtain the feature extraction algorithm, and perform feature extraction on the image to be extracted through the feature extraction algorithm to obtain the target extraction image.
[0089] For the above steps, the specific implementation in this embodiment is as follows:
[0090] First, fuse the spatial features with the target dehazed image to obtain the image to be extracted including spatial features. Then, obtain the initial feature extraction algorithm to extract features from the image to be extracted. In this embodiment, the selected initial feature extraction algorithm is the YOLOv5 feature extraction algorithm. The initial feature extraction algorithm can extract features in the image and finally detect the objects in the image through the features of the image, generate calibration boxes to achieve the calibration and recognition of the objects in the image and the tracking target or fixed reference object. Next, we change the internal parameters of the feature extraction algorithm of the initial feature extraction algorithm multiple times, and after each change of the internal parameters of the feature extraction algorithm, input the same training picture set into the initial feature extraction algorithm and calculate the classification error of the initial feature extraction algorithm and the bounding box regression error , and the loss error of the initial feature extraction algorithm can be calculated through the classification error and the bounding box regression error of the initial feature extraction algorithm :
[0091]
[0092] Among them, the calculation method of the classification error is:
[0093]
[0094] Among them, B is the total number of classes into which the initial feature extraction algorithm may classify the content in the training picture set, A is the total number of targets in the training picture set, represents the expected probability of classifying the m-th target in the training picture set as the c-th class determined according to experience, represents the probability that the initial feature extraction algorithm actually classifies the m-th target in the training picture set as the c-th class, and E() represents taking the binary cross-entropy loss within the parentheses.
[0095] The calculation method of the bounding box regression error is:
[0096]
[0097] Among them, represents the area of the calibration box generated by the initial feature extraction algorithm, represents the theoretical area of the ideal calibration box determined according to experience, represents the area of the target judged by the initial feature extraction algorithm on the picture, A is the actual area of the target on the picture, represents obtaining the area of the intersection of the calibration box generated by the initial feature extraction algorithm and the ideal calibration box.
[0098] Calculate the loss error of the initial feature extraction algorithm under the internal parameters of multiple groups of feature extraction algorithms, and retain the internal parameters of the feature extraction algorithm of the group of initial feature extraction algorithms with the smallest loss error, then the feature extraction algorithm can be obtained. Use the feature extraction algorithm to extract features from the image to be extracted, that is, generate detection frames to calibrate and identify the tracking target or fixed reference object in the image, judge which object in the image to be extracted is the tracking target or fixed reference object and mark it with the detection frame, then the target extraction image can be obtained. Since the image to be extracted contains spatial features, the target extraction image also includes spatial features, and the tracking target or fixed reference object is marked in the target extraction image, so the distance, orientation, etc. between the camera and the tracking target or fixed reference object can be determined by combining the detection frame and the spatial information.
[0099] S4: Detect through the navigation radar to obtain the initial radar image, perform spatial folding on the initial radar image and calculate the channel weights to obtain the channel-weighted radar image, calculate the spatial weighting matrix of the channel-weighted radar image, and obtain the target radar image through the spatial weighting matrix;
[0100] Furthermore, the purpose of this stage is to perform channel weighting and spatial weighting on the initial radar image to obtain the target radar image. Step S4 specifically includes:
[0101] S41: Detect the shooting area of the camera through the navigation radar to obtain the initial radar image;
[0102] S42: Perform linear projection on the initial radar image to obtain the linearly projected radar image, fold the spatial dimension of the linearly projected radar image to obtain the spatially folded radar image, and calculate the channel weights of the spatially folded radar image to obtain the channel-weighted matrix feature map. Obtain the channel-weighted radar image through the channel information of the spatially folded radar image and the channel-weighted matrix feature map;
[0103] S43: Activate and convolve the channel-weighted radar image to obtain the spatial weighting matrix, multiply the spatial weighting matrix and the linearly projected radar image element by element to obtain the target radar image.
[0104] For the above steps, the specific implementation in this embodiment is as follows:
[0105] First, the navigation radar detects the area that has been photographed by the camera, i.e., the shooting area of the camera, and the initial radar image can be obtained. Here, the navigation radar can be a lidar, a millimeter-wave radar, etc. Then, the GELU activation function is used to perform a linear projection on the initial radar image, and the linearly projected radar image can be obtained. Next, the adaptive average pooling method is used to fold the spatial dimension of the linearly projected radar image in the spatial dimension to obtain the spatially folded radar image, so as to reduce the dimension of the initial radar image and enable the target radar image to be fused with the two-dimensional target extraction image.
[0106] Set the parameters of the softmax activation function as needed, and use the softmax activation function to calculate the weights of the spatially folded radar image on different channels, i.e., the channel weights, and the channel-weighted matrix feature map of the spatially folded radar image can be obtained. Multiply the vector related to the channel included in the spatially folded radar image, i.e., the channel information, by the channel-weighted matrix feature map, and the channel-weighted radar image can be obtained.
[0107] Then, the GELU activation function is used to activate the channel-weighted radar image, and a 1×1 convolutional kernel is used to perform convolution on the channel-weighted radar image, and the spatial weighted matrix for enhancing the spatial information of the linearly projected radar image can be obtained. Since the spatial weighted matrix is the same size as the linearly projected radar image, the spatial weighted matrix and the linearly projected radar image are multiplied element by element, and the target radar image can be obtained. The target radar image also contains spatial information.
[0108] S5: Fuse the target radar image with the target extraction image to obtain a navigation image, calculate the marker azimuth difference through the navigation image, and obtain a navigation control signal according to the marker azimuth difference. The unmanned boat completes the heading control according to the navigation control signal.
[0109] Furthermore, the purpose of this stage is to obtain a navigation image, calculate the marker azimuth difference through the navigation image, and control the unmanned boat. The specific implementation method is as follows: First, since both the target radar image and the target extraction image are images of the shooting area, the target radar image and the target extraction image can be fused to obtain the navigation image. Since both the target radar image and the target extraction image include spatial information, the navigation image also includes spatial information. Determine the azimuth relationship between the unmanned boat and the marker in the ideal state according to the task of the unmanned boat. According to the spatial information, the distance and azimuth between the camera or the navigation radar and the tracking target or the fixed reference object can be judged. Therefore, the marker azimuth difference can be calculated according to the navigation image, that is, the difference between the current azimuth relationship between the unmanned boat and the marker and the azimuth relationship between the unmanned boat and the marker in the ideal state.
[0110] For example, when the target of the unmanned boat is to track a tracking target, and the ideal azimuth is directly behind the tracking target and 100 m away from the tracking target, it can be obtained that the tracking target should be at the center of the navigation image, and the distance in the image, that is, the distance from the camera or the navigation radar is 100 m. Then, since the target extraction image marks the tracking target or the fixed reference object, the actual azimuth and distance of the tracking target in the navigation image can be judged. The difference between the actual azimuth and distance and the ideal azimuth and distance is obtained to get the marker azimuth difference. According to the marker azimuth difference, a navigation control signal can be generated to drive the unmanned boat to eliminate the marker azimuth difference until the marker azimuth difference is 0, thus completing the control of the unmanned boat. In addition, when there are obstacles in the navigation image and it is judged according to the spatial information that the unmanned boat is too close to the obstacle, a navigation control signal can also be generated to make the unmanned boat move away from the obstacle.
[0111] The present invention generates a navigation image through a navigation radar and a camera, and controls the unmanned boat by calculating the marker azimuth difference, so that the unmanned boat can still perform autonomous navigation control according to the tracking target or the fixed reference object without obtaining external navigation or control signals, so as to complete the task.
[0112] The following describes the image-based unmanned boat control device provided by the present invention. The image-based unmanned boat control device described below can be mutually corresponding and referred to the image-based unmanned boat control method described above.
[0113] Figure 2 The structural schematic diagram of the image-based unmanned boat control system is exemplified, as Figure 2 shown, for executing the image-based unmanned boat control method as described above, including:
[0114] Spatial feature acquisition module 100: used to arrange a camera and a navigation radar on the unmanned boat, determine the shooting area, shoot the shooting area through the camera, acquire the image to be processed, perform Fourier transform and adaptive feature extraction on the image to be processed, obtain the target amplitude spectrum, and obtain the spatial feature through the target amplitude spectrum;
[0115] Image dehazing module 200: used to perform global grid filling on the image to be processed to obtain a filled image, perform dehazing processing on the filled image to obtain an initial dehazed image, and perform reverse filling on the initial dehazed image to obtain a target dehazed image;
[0116] Feature extraction module 300: used to fuse the spatial feature with the target dehazed image to obtain an image to be extracted, perform feature extraction on the image to be extracted, and obtain a target extraction image;
[0117] Radar image acquisition module 400: It is used to detect through a marine radar to obtain an initial radar image, perform spatial folding on the initial radar image and calculate channel weights to obtain a channel-weighted radar image, calculate the spatial weighting matrix of the channel-weighted radar image, and obtain a target radar image through the spatial weighting matrix;
[0118] Unmanned boat control module 500: It is used to fuse the target radar image with the target extraction image to obtain a navigation image, calculate the marker azimuth difference through the navigation image, and obtain a navigation control signal according to the marker azimuth difference. The unmanned boat completes the heading control according to the navigation control signal.
[0119] On the other hand, Figure 3 An example of a schematic physical structure diagram of an electronic device is shown as Figure 3 shown. The electronic device may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communication interface 820, and the memory 830 complete mutual communication through the communication bus 840. The processor 810 can call the logical instructions in the memory 830 to execute the unmanned boat control method for images, and this method includes:
[0120] S1: Arrange a camera and a marine radar on the unmanned boat, determine the shooting area and shoot the shooting area through the camera to obtain an image to be processed, perform Fourier transform and adaptive feature extraction on the image to be processed to obtain a target amplitude spectrum, and obtain spatial features through the target amplitude spectrum;
[0121] S2: Perform global grid filling on the image to be processed to obtain a filled image, perform defogging processing on the filled image to obtain an initial defogged image, and perform reverse filling on the initial defogged image to obtain a target defogged image;
[0122] S3: Fuse the spatial features with the target defogged image to obtain an image to be extracted, perform feature extraction on the image to be extracted to obtain a target extraction image;
[0123] S4: Detect through a marine radar to obtain an initial radar image, perform spatial folding on the initial radar image and calculate channel weights to obtain a channel-weighted radar image, calculate the spatial weighting matrix of the channel-weighted radar image, and obtain a target radar image through the spatial weighting matrix;
[0124] S5: Fuse the target radar image with the target extraction image to obtain a navigation image, calculate the marker azimuth difference through the navigation image, and obtain a navigation control signal according to the marker azimuth difference. The unmanned boat completes the heading control according to the navigation control signal.
[0125] In addition, when the logic instructions in the above-mentioned memory 830 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0126] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the image-based unmanned boat control method provided by the above-mentioned various methods. The method includes:
[0127] S1: Arrange a camera and a navigation radar on the unmanned boat, determine the shooting area, and shoot the shooting area through the camera to obtain a to-be-processed image. Perform Fourier transform and adaptive feature extraction on the to-be-processed image to obtain a target amplitude spectrum, and obtain spatial features through the target amplitude spectrum;
[0128] S2: Perform global grid filling on the to-be-processed image to obtain a filled image. Perform defogging processing on the filled image to obtain an initial defogged image. Perform reverse filling on the initial defogged image to obtain a target defogged image;
[0129] S3: Fuse the spatial features with the target defogged image to obtain an image to be extracted. Perform feature extraction on the image to be extracted to obtain a target extracted image;
[0130] S4: Detect through the navigation radar to obtain an initial radar image. Perform spatial folding on the initial radar image and calculate the channel weight value to obtain a channel-weighted radar image. Calculate the spatial weighting matrix of the channel-weighted radar image, and obtain a target radar image through the spatial weighting matrix;
[0131] S5: Fuse the target radar image with the target extraction image to obtain a navigation image, calculate the azimuth difference of the landmark through the navigation image, and obtain a navigation control signal based on the azimuth difference of the landmark. The unmanned boat completes the heading control according to the navigation control signal.
[0132] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the image-based unmanned boat control method provided by the above-mentioned various methods. The method includes:
[0133] S1: Arrange a camera and a navigation radar on the unmanned boat, determine the shooting area and shoot the shooting area through the camera to obtain an image to be processed. Perform Fourier transform and adaptive feature extraction on the image to be processed to obtain a target amplitude spectrum, and obtain spatial features through the target amplitude spectrum;
[0134] S2: Perform global grid filling on the image to be processed to obtain a filled image, perform defogging processing on the filled image to obtain an initial defogged image, and perform reverse filling on the initial defogged image to obtain a target defogged image;
[0135] S3: Fuse the spatial features with the target defogged image to obtain an image to be extracted, perform feature extraction on the image to be extracted to obtain a target extraction image;
[0136] S4: Detect through the navigation radar to obtain an initial radar image, perform spatial folding on the initial radar image and calculate the channel weight to obtain a channel-weighted radar image, calculate the spatial weighting matrix of the channel-weighted radar image, and obtain a target radar image through the spatial weighting matrix;
[0137] S5: Fuse the target radar image with the target extraction image to obtain a navigation image, calculate the azimuth difference of the landmark through the navigation image, and obtain a navigation control signal based on the azimuth difference of the landmark. The unmanned boat completes the heading control according to the navigation control signal.
[0138] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.
[0139] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0140] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An image-based unmanned boat control method, characterized in that: include: S1: Arrange a camera and a navigation radar on the unmanned boat, determine a shooting area, shoot the shooting area through the camera, obtain an image to be processed, perform Fourier transform and adaptive feature extraction on the image to be processed, obtain a target amplitude spectrum, and obtain a spatial feature through the target amplitude spectrum; S2: performing global grid filling on the image to be processed to obtain a filled image, performing defogging on the filled image to obtain an initial defogged image, and performing reverse filling on the initial defogged image to obtain a target defogged image; S3: fusing the spatial features with the target defogging image to obtain an image to be extracted, and performing feature extraction on the image to be extracted to obtain a target extracted image; wherein step S3 specifically includes: S31: fusing the spatial features with the target defogging image to obtain an image to be extracted, selecting an initial feature extraction algorithm, changing internal parameters of the feature extraction algorithm of the initial feature extraction algorithm, and calculating a classification error and a bounding box regression error of the initial feature extraction algorithm; S32: calculating the loss error of the initial feature extraction algorithm through the classification error and the bounding box regression error of the initial feature extraction algorithm, retaining the internal parameters of the feature extraction algorithm when the loss error is minimum, obtaining a feature extraction algorithm, and performing feature extraction on the image to be extracted through the feature extraction algorithm to obtain the target extracted image; S4: Detecting by navigation radar to obtain an initial radar image, spatially folding the initial radar image and calculating channel weights to obtain a channel-weighted radar image, calculating a spatial weighting matrix of the channel-weighted radar image, and obtaining a target radar image by using the spatial weighting matrix; S5: The target radar image is merged with the target extraction image to obtain a navigation image, the marker azimuth difference is calculated through the navigation image, and a navigation control signal is obtained according to the marker azimuth difference, and the unmanned boat completes heading control according to the navigation control signal.
2. The image-based unmanned boat control method according to claim 1, characterized in that: Step S1 specifically includes: S11: Determine the installation conditions of the camera and the navigation radar and the equipment environment of the unmanned boat, install the camera and the navigation radar on the unmanned boat according to the installation conditions and the equipment environment, determine a shooting area, shoot the shooting area with the camera, and obtain an image to be processed; S12: Obtaining the channel dimension of the image to be processed, performing block processing on the image to be processed according to the channel dimension, and performing average pooling processing and maximum pooling processing on the image to be processed to obtain a pooled image; S13: Perform Fourier transform on the pooled image to obtain an amplitude spectrum, perform frequency band feature extraction on the amplitude spectrum to obtain spatial adaptive features, rearrange and stack the spatial adaptive features to obtain a target amplitude spectrum, use a minimax algorithm, and obtain the spatial features through the target amplitude spectrum.
3. The image-based unmanned boat control method according to claim 1, characterized in that: Step S2 specifically includes: S21: performing edge filling and global feature grid filling on the image to be processed to obtain a filled image; S22: Selecting an initial defogging algorithm, changing the internal parameters of the initial defogging algorithm and calculating the defogging loss of the initial defogging algorithm, retaining the internal parameters of the defogging algorithm when the defogging loss is minimum, obtaining the defogging algorithm, and performing defogging processing on the filled image by using the defogging algorithm to obtain the initial defogging image; S23: performing reverse edge filling and reverse global filling on the initial defogging image to obtain the target defogging image.
4. The image-based unmanned boat control method according to claim 1, characterized in that: Step S4 specifically includes: S41: detecting the shooting area of the camera by means of a navigation radar to obtain an initial radar image; S42: linearly projecting the initial radar image to obtain a linear projection radar image, folding the spatial dimension of the linear projection radar image to obtain a spatial folding radar image, calculating the channel weight of the spatial folding radar image to obtain a channel weighted matrix feature map, and obtaining the channel weighted radar image through the channel information of the spatial folding radar image and the channel weighted matrix feature map; S43: activating and convolving the channel weighted radar image to obtain the spatial weighted matrix, and multiplying the spatial weighted matrix by the linear projection radar image element by element to obtain the target radar image.
5. The image-based unmanned boat control method according to claim 1, characterized in that: The camera is an RGBD camera, and the image to be processed taken by the RGBD camera includes image depth information.
6. An image-based unmanned boat control system, used to execute the image-based unmanned boat control method according to any one of claims 1 to 5, characterized in that: include: Spatial feature acquisition module: used for arranging a camera and a navigation radar on the unmanned boat, determining a shooting area, and shooting the shooting area through the camera, acquiring an image to be processed, performing Fourier transform and adaptive feature extraction on the image to be processed, obtaining a target amplitude spectrum, and obtaining spatial features through the target amplitude spectrum; Image defogging module: used for performing global grid filling on the image to be processed to obtain a filled image, performing defogging on the filled image to obtain an initial defogging image, and performing reverse filling on the initial defogging image to obtain a target defogging image; Feature extraction module: used to fuse the spatial features with the target defogging image to obtain the image to be extracted, and perform feature extraction on the image to be extracted to obtain the target extracted image; wherein step S3 specifically includes: S31: fusing the spatial features with the target defogging image to obtain an image to be extracted, selecting an initial feature extraction algorithm, changing internal parameters of the feature extraction algorithm of the initial feature extraction algorithm, and calculating a classification error and a bounding box regression error of the initial feature extraction algorithm; S32: calculating the loss error of the initial feature extraction algorithm through the classification error and the bounding box regression error of the initial feature extraction algorithm, retaining the internal parameters of the feature extraction algorithm when the loss error is minimum, obtaining a feature extraction algorithm, and performing feature extraction on the image to be extracted through the feature extraction algorithm to obtain the target extracted image; Radar image acquisition module: used to detect through navigation radar to obtain an initial radar image, perform spatial folding on the initial radar image and calculate channel weights to obtain a channel-weighted radar image, calculate a spatial weighted matrix of the channel-weighted radar image, and obtain a target radar image through the spatial weighted matrix; Unmanned boat control module: used to fuse the target radar image with the target extraction image to obtain a navigation image, calculate the marker azimuth difference through the navigation image, and obtain a navigation control signal based on the marker azimuth difference. The unmanned boat completes heading control based on the navigation control signal.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the image-based unmanned boat control method as described in any one of claims 1 to 5 are implemented.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the image-based unmanned boat control method according to any one of claims 1 to 5 are implemented.
9. A computer program product, the computer program product comprising a computer program stored on a non-transitory computer-readable storage medium, the computer program comprising program instructions, characterized in that: When the program instructions are executed by a computer, the computer can execute the steps of the image-based unmanned boat control method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Target detection method for autonomous vehicle based on radar and vision fusion
CN116958934A
Image defogging processing method and device, electronic equipment and storage medium
CN118071645A
Frequency guidance space self-adaption-based camouflage target detection method
CN118379484A