Fish school detection preview method, device and equipment based on binocular camera image stitching
By using underwater binocular camera calibration and gradient weight stitching technology, the problems of brightness and texture discontinuity in underwater image stitching are solved, generating high-quality panoramic underwater images that support clear fish school detection previews.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN YINGZHI FUTURE TECHNOLOGY CO LTD
- Filing Date
- 2026-01-27
- Publication Date
- 2026-04-21
AI Technical Summary
In image stitching of underwater binocular cameras, images with overlapping fields of view may result in uneven color and brightness, leading to visual discontinuity.
By acquiring images from multiple discrete acquisition moments of an underwater binocular camera, the target calibration parameters are calibrated using auxiliary lines on the inner wall of a cube calibration device. Then, the corresponding pixels of the lens are fused using a gradient weighted mirror linear stitching method to eliminate brightness differences and texture breaks.
It achieves a natural transition of panoramic underwater images, improves image quality, provides clear and coherent previews of fish detection images, and enhances the realism of underwater scenes.
Smart Images

Figure CN121908120A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to a method, apparatus and device for fish school detection and preview based on image stitching from a binocular camera. Background Technology
[0002] Fish detection currently utilizes binocular cameras to capture underwater images, allowing for the determination of fish location and swimming direction, thus increasing fishing success rates. Typically, images captured simultaneously by the left and right lenses are stitched together to display the underwater scene. However, in the binocular panoramic imaging system of an underwater camera, there is an overlapping area of view between the left and right lenses. If the image from one camera in the overlapping area is discarded during stitching, the resulting image may not be smooth in terms of color and brightness, exhibiting visual discontinuities. Summary of the Invention
[0003] The purpose of this invention is to provide a method, apparatus, and device for fish detection and preview based on image stitching from a binocular camera, which aims to eliminate possible brightness differences or texture breaks at the stitching points, reduce stitching marks, and achieve a natural transition in the image.
[0004] To achieve the above objectives, a first aspect of this disclosure provides a fish school detection preview method based on binocular camera image stitching, the method comprising: Acquire underwater images captured by the two lenses of an underwater binocular camera at multiple discrete acquisition times; The underwater image is calibrated using pre-calibrated target calibration parameters to obtain the target underwater image at the corresponding acquisition time. The target calibration parameters are obtained by calibrating multiple core parameters sequentially based on multiple sets of auxiliary lines drawn on the inner wall of the cube calibration device as geometric references. The pixels in the target underwater image corresponding to either of the two lenses that are between a preset fusion percentage and a preset reference radius are mirrored and linearly stitched together with the pixels in the overlapping area of the target underwater image corresponding to the other lens, using a gradient weighting method, to obtain the panoramic underwater image at the time of acquisition. Based on the panoramic underwater images corresponding to the multiple discrete acquisition times and the preview mode selected by the user, a fish school detection image preview is displayed.
[0005] Optionally, the step of performing mirror linear stitching and fusion of pixels in the target underwater image corresponding to either of the two lenses that are located between a preset fusion percentage and a preset reference radius, and pixels in the overlapping area of the target underwater image corresponding to the other lens, using gradient weights, to obtain the panoramic underwater image at the time of acquisition, includes: Determine the true coordinates of the first pixel and the second pixel in the world coordinate system, wherein the first pixel is a pixel in the target underwater image corresponding to either of the two lenses that is located between a preset fusion percentage and a preset reference radius, and the second pixel is a pixel in the target underwater image corresponding to the other lens that is located in the overlapping area. Based on the true coordinates of each first pixel, a weight corresponding to the first pixel is determined, and based on the weight of the first pixel with the same true coordinates as each second pixel, a weight corresponding to the second pixel is determined, wherein the weight of the first pixel gradually decreases from the side of the overlapping region closer to the lens side to the side farther away from the lens side, and the weight of the second pixel gradually increases as the weight of the first pixel with the same true coordinates gradually decreases. Based on the weight of each first pixel and the weight of each second pixel, the first pixels and second pixels with the same real coordinates are stitched together and fused to obtain the panoramic underwater image at the time of acquisition.
[0006] Optionally, the weight of the first pixel gradually decreases from the side of the overlapping region closer to either lens side to the side farther away from either lens side, including: The region between the preset fusion percentage and the preset baseline radius is divided into a preset number of splicing and fusion regions of equal width; A first color difference value is determined based on the difference between the average pixel value of a first pixel in the first stitching and fusion region from the side closer to the lens side to the side farther away from the lens side of the overlapping region and a reference pixel value; and a second color difference value is determined based on the difference between the average pixel value of a second pixel in the first stitching and fusion region and the reference pixel value, wherein the reference pixel value is the average pixel value of the first pixel and the second pixel between the preset fusion percentage and the preset reference radius; The weight of the first pixel in the first splicing and fusion region is determined based on the color difference value and the preset tolerance parameter. Based on the preset regional weight difference and the weight of the first pixel in the first stitching and fusion region, the weight of the first pixel in each stitching and fusion region from the side of the overlapping region closest to any lens side to the side furthest from any lens side is determined.
[0007] Optionally, determining the true coordinates of the first pixel and the second pixel in the world coordinate system includes: Based on the first pixel coordinates in the first pixel coordinate system corresponding to any lens and the intrinsic and extrinsic parameter matrix of any lens, determine the true coordinates of the first pixel in the world coordinate system; Based on the second pixel coordinates in the second pixel coordinate system corresponding to the other lens and the intrinsic and extrinsic parameter matrix of the other lens, the true coordinates of the second pixel in the world coordinate system are determined.
[0008] Optionally, the step of previewing and displaying fish school detection images based on the panoramic underwater images corresponding to the multiple discrete acquisition times and the preview mode selected by the user includes: Based on the panoramic underwater images corresponding to the multiple discrete acquisition times, the panoramic underwater images are processed by OpenGL texture mapping and rendering pipeline, and mapped as textures onto the surface of a 3D panoramic model. The 3D panoramic model is a panoramic spherical or cylindrical model constructed based on vertex coordinates defined in three-dimensional space. When the user selects VR interactive preview mode, a VR panoramic preview based on viewpoint interaction is displayed to preview images of fish school detection, and the viewing perspective of the 3D panoramic model is switched in response to the user's viewpoint adjustment operation; or, When the user selects a planar panoramic preview mode, an equidistant cylindrical projection algorithm is used to map the texture information on the 3D panoramic model to a two-dimensional plane according to the equidistant cylindrical projection rule, and a fish school detection image preview is displayed based on the image on the two-dimensional plane.
[0009] Optionally, the target calibration parameters are obtained by calibration in the following manner: The geometric center of the lens is determined by drawing horizontal and vertical crosshairs on the original imaging image captured by the lens. Using the reference calibration center formed by the intersection of auxiliary lines on the inner wall of the cube calibration device as a reference, adjust the coordinates of the original image acquired by the lens until the position of the geometric center and the reference calibration center meets the preset position requirements, and use the translation amount of the coordinates as the corresponding center offset parameter of the lens. Given the center offset parameter, the original image is scaled to determine the effective radius of the lens's field of view. Given a fixed effective field of view radius, the original images captured by the two lenses are rendered onto a spherical circle model for Euler angle calibration to obtain the Euler angle parameters of the lenses. The target calibration parameters include the center offset parameter, the effective field of view radius, and the Euler angle parameter.
[0010] Optionally, calibrating the underwater image using pre-defined target calibration parameters to obtain the target underwater image at the corresponding acquisition time includes: Based on the center offset parameter, the underwater images corresponding to the two lenses are translated respectively so that the geometric center of the lens is aligned with the center of the preset canvas; Based on the effective field of view alarm, the translated underwater image is scaled, and invalid images that exceed the preset canvas are cropped out to obtain the effective area image; The effective area image is mapped onto the 3D semicircular model textures corresponding to the two lenses, and the attitude of the 3D semicircular model is adjusted according to the Euler angle parameters to obtain the target underwater image at the corresponding acquisition time.
[0011] A second aspect of this disclosure provides a fish school detection and preview device based on binocular camera image stitching, the device comprising: The acquisition module is configured to acquire underwater images captured by the two lenses of an underwater binocular camera at multiple discrete acquisition times. The calibration module is configured to calibrate the underwater image using pre-calibrated target calibration parameters to obtain the target underwater image at the corresponding acquisition time. The target calibration parameters are obtained by calibrating multiple core parameters sequentially based on multiple sets of auxiliary lines drawn on the inner wall of the cube calibration device as geometric references. The stitching and fusion module is configured to perform mirror linear stitching and fusion of pixels in the target underwater image corresponding to one of the two lenses that are located between the preset fusion percentage and the preset fusion radius, and pixels in the target underwater image corresponding to the other lens that are in the overlapping area, using a gradient weighting method, to obtain a panoramic underwater image at the time of acquisition. The preview module is configured to preview and display fish school detection images based on the panoramic underwater images corresponding to the multiple discrete acquisition times and the preview mode selected by the user.
[0012] A third aspect of this disclosure provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps of the method described in any of the first aspects.
[0013] A fourth aspect of this disclosure provides an electronic device, comprising: A memory on which computer programs are stored; A processor for executing the computer program in the memory to implement the steps of the method of any one of the first aspects.
[0014] This invention provides a method, apparatus, and device for fish school detection and preview based on image stitching from a binocular camera. Compared with existing technologies, it has the following advantages: By acquiring underwater images at multiple discrete acquisition times using an underwater binocular camera and calibrating them using target calibration parameters obtained from calibration based on multiple sets of auxiliary lines on the inner wall of a cube calibration device, the accuracy of the images is improved. In the stitching and fusion stage, a gradient weighted mirror linear stitching method is used to fuse pixels in specific areas of the target underwater images from the two lenses. This eliminates potential brightness differences or texture breaks at the stitching points, effectively reducing stitching artifacts and achieving a natural transition in the image. This significantly improves the quality of the generated panoramic underwater image, presenting the underwater scene more realistically. Finally, based on the panoramic underwater images from multiple discrete acquisition times and the user-selected preview mode, a fish school detection image preview is displayed, providing users with clear, coherent, and natural dynamic underwater fish school images.
[0015] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description
[0016] The accompanying drawings are provided to further illustrate the present disclosure and form part of the specification. They are used together with the following detailed description to explain the present disclosure, but do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart illustrating a fish school detection and preview method based on image stitching from a binocular camera, as shown in the embodiments of the specification.
[0017] Figure 2 A block diagram of a fish school detection and preview device based on binocular camera image stitching, as shown in the embodiment of the specification.
[0018] Figure 3 This is a block diagram of another fish school detection and preview device based on binocular camera image stitching, as shown in the embodiment of the specification. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0020] This application provides a method for previewing fish school detection based on image stitching from a binocular camera. Figure 1 This is a flowchart illustrating a fish school detection preview method based on binocular camera image stitching according to an embodiment. The method includes: In step S11, underwater images are acquired by the two lenses of the underwater binocular camera at multiple discrete acquisition times; Among them, the underwater binocular camera is a camera device used for underwater environment filming. Equipped with two lenses, it can capture images from different angles and obtain three-dimensional information of objects through the principle of binocular vision. Discrete acquisition moments are specific, discontinuous points in time on the timeline. The camera captures images at these points to obtain underwater scene information at different times. Underwater images are image data reflecting the underwater environment, objects, etc., captured by the two lenses of the underwater binocular camera at various discrete acquisition moments.
[0021] In this embodiment, an underwater binocular camera is deployed underwater and contains a precise timing module. The timing module triggers both lenses to acquire images at multiple discrete time points according to preset time intervals. Both lenses operate simultaneously, capturing underwater scene light information from their respective perspectives and converting it into digital image signals. These digital image signals undergo preliminary processing, such as noise reduction and enhancement, by the camera's internal image processing circuitry, and are then stored in the camera's storage device, thereby obtaining underwater images captured by the two lenses at multiple discrete acquisition times.
[0022] In step S12, the underwater image is calibrated using pre-calibrated target calibration parameters to obtain the target underwater image at the corresponding acquisition time. The target calibration parameters are obtained by calibrating multiple core parameters sequentially based on multiple sets of auxiliary lines drawn on the inner wall of the cube calibration device as geometric references. In this embodiment, the auxiliary lines on the inner wall of the cube calibration device are used as a geometric reference to calibrate core parameters such as lens distortion and camera intrinsic parameters to obtain target calibration parameters. Using these parameters to correct underwater images can eliminate the effects of lens distortion and other factors, making the images more accurately reflect the real scene. Targeted geometric transformation calibration can be performed on the acquired raw images to eliminate image distortion and positional deviations caused by errors.
[0023] The target calibration parameters may include: left camera parameters: left camera center x-coordinate, left camera center y-coordinate, left camera effective radius, left camera roll angle, left camera yaw angle, and left camera pitch angle; and right camera parameters: right camera center x-coordinate, right camera center y-coordinate, right camera effective radius, right camera roll angle, right camera yaw angle, and right camera pitch angle.
[0024] In step S13, pixels in the target underwater image corresponding to either of the two lenses that are located between a preset fusion percentage and a preset reference radius are mirrored and linearly stitched together with pixels in the target underwater image corresponding to the other lens in the overlapping area, using a gradient weighting method, to obtain the panoramic underwater image at the acquisition time. Gradient weighting assigns different weight values to pixels based on their positions during image stitching and fusion, ensuring a smooth transition between pixel values at the stitching point, avoiding obvious stitching artifacts, and improving image fusion quality. Mirrored linear stitching and fusion is a stitching method that linearly combines corresponding pixels from two images according to certain weights. Mirroring makes the stitching transition more natural, achieving seamless image fusion.
[0025] In this embodiment, the overlapping area of the underwater images of two lenses is determined, and pixels in either lens image between a preset fusion percentage and a preset reference radius are selected. These pixels are then compared with the pixels in the overlapping area of the other lens. Gradient weights are calculated based on the pixel positions, and the pixels are then mirrored linearly combined according to the weights to achieve a smooth transition of pixels and obtain a panoramic underwater image.
[0026] In step S14, a fish school detection image preview is displayed based on the panoramic underwater image corresponding to the multiple discrete acquisition times and the preview mode selected by the user.
[0027] In this embodiment of the disclosure, panoramic underwater images acquired at multiple discrete times are processed according to the preview mode selected by the user, such as dynamic display or static focus, such as adjusting the display speed or highlighting specific fish schools, and finally the fish school detection image is presented to the user on the display device.
[0028] The aforementioned technical solution improves image accuracy by acquiring underwater images from multiple discrete acquisition moments using an underwater binocular camera and calibrating them using target calibration parameters obtained from calibration with multiple sets of auxiliary lines on the inner wall of a cube calibration device. In the stitching and fusion stage, a gradient weighted mirror linear stitching method is used to fuse pixels in specific areas of the target underwater images from the two lenses. This eliminates potential brightness differences or texture breaks at the stitching points, effectively reducing stitching artifacts and achieving a natural transition in the image. This significantly improves the quality of the generated panoramic underwater image, presenting a more realistic underwater scene. Finally, a fish school detection image preview is displayed based on the panoramic underwater images from multiple discrete acquisition moments and the user-selected preview mode, providing users with clear, continuous, and natural dynamic underwater fish school images.
[0029] Optionally, in step S13, the step of performing mirror linear stitching and fusion of pixels in the target underwater image corresponding to either of the two lenses that are located between a preset fusion percentage and a preset reference radius, and pixels in the overlapping area of the target underwater image corresponding to the other lens, through gradient weighting, to obtain the panoramic underwater image at the acquisition time, includes: In step S131, the true coordinates of the first pixel and the second pixel in the world coordinate system are determined, wherein the first pixel is a pixel in the target underwater image corresponding to either of the two lenses that is located between a preset fusion percentage and a preset reference radius, and the second pixel is a pixel in the target underwater image corresponding to the other lens that is located in the overlapping area. In this embodiment of the disclosure, the first pixel (pixels between a preset fusion percentage and a preset reference radius) and the second pixel (pixels in the overlapping area) are converted from image coordinates to world coordinates by camera calibration parameters, thereby clarifying their positions in the real scene.
[0030] In step S132, the weight of the first pixel is determined according to the real coordinates of each first pixel, and the weight of the second pixel is determined according to the weight of the first pixel with the same real coordinates as each second pixel. The weight of the first pixel gradually decreases from the side of the overlapping region closer to the lens side to the side farther away from the lens side, and the weight of the second pixel gradually increases as the weight of the first pixel with the same real coordinates gradually decreases. In this embodiment, the weight of the first pixel gradually decreases from the side closest to the lens to the side furthest away. This is because pixels closer to the lens have a stronger influence on the stitching process, while those farther away have less impact. The weight of the second pixel increases as the weight of the first pixel with the same real coordinates decreases. This ensures a smooth change in pixel values at the stitching point, avoids abrupt changes, and achieves a natural transition. For example, when stitching architectural images, the weight of the first pixel on the left lens side is set to 0.8 - 0.2 (from near to far). If a second pixel has the same real coordinates as the first pixel and the corresponding weight of the first pixel is 0.5, then the weight of the second pixel is set to 0.5, making the pixel transition at the stitching point natural.
[0031] In step S133, based on the weight of each first pixel and the weight of each second pixel, the first pixels and the second pixels with the same real coordinates are stitched together and fused to obtain the panoramic underwater image at the acquisition time.
[0032] In this embodiment, pixels with the same real coordinates are linearly combined based on the determined weights of the first and second pixels. The first pixel value is multiplied by its weight, the second pixel value is multiplied by its weight, and then the two are added together to obtain the fused pixel value. All corresponding pixels are processed in this way to complete image stitching and fusion, resulting in a panoramic underwater image. For a first pixel value of 100 (weight 0.6) and a second pixel value of 120 (weight 0.4) at the same location in the image, the fused pixel value is 100 × 0.6 + 120 × 0.4 = 108. All pixels are processed in this way to obtain a panoramic image.
[0033] The aforementioned technical solution rationally allocates gradient weights based on location, resulting in a natural stitching transition, effectively eliminating stitching artifacts, and avoiding brightness differences and texture breaks. In underwater image stitching, it can generate high-quality panoramic underwater images, more realistically presenting underwater scenes. For fish detection, it can clearly show the distribution and dynamics of fish in the panorama, improving detection accuracy and comprehensiveness.
[0034] Optionally, in step S132, the weight of the first pixel gradually decreases from the side of the overlapping region closer to either lens side to the side farther away from either lens side, including: In step S1321, the area between the preset fusion percentage and the preset reference radius is divided into a preset number of splicing and fusion areas of equal width; The stitching and fusion region is a sub-region obtained by dividing the area between the preset fusion percentage and the preset baseline radius on an average basis. It is used to more finely control the distribution of pixel weights and the stitching transition.
[0035] In this embodiment, the region between a preset fusion percentage and a preset baseline radius is divided into multiple stitching and fusion regions of equal width. This allows for precise control of the weights of pixels at different positions during the stitching process, resulting in a smoother and more natural stitching transition and avoiding obvious stitching artifacts. By subdividing the region, weight adjustments can be made based on the characteristics of each small region.
[0036] For example, when stitching two images of underwater fish, the area between the preset blending percentage and the preset baseline radius is divided into five stitching regions. This allows for more precise adjustment of pixel weights based on the characteristics of each region, such as fish distribution and lighting conditions, resulting in a more natural stitched image.
[0037] In step S1322, a first color difference value is determined based on the difference between the average pixel value of the first pixel in the first stitching and fusion region from the side of the overlapping region closest to the lens side to the side furthest from the lens side and a reference pixel value, and a second color difference value is determined based on the difference between the average pixel value of the second pixel in the first stitching and fusion region and the reference pixel value, wherein the reference pixel value is the average pixel value of the first pixel and the second pixel between the preset fusion percentage and the preset reference radius; In this embodiment of the disclosure, the difference between the average pixel value of the first pixel and the second pixel in the first stitching and fusion region and the reference pixel value is calculated to obtain the first color difference value and the second color difference value. The reference pixel value is the average value of the entire stitching region. By comparison, the degree of deviation of the pixels in this region from the overall average level can be understood. For example, for the first stitching and fusion region, the average pixel value of the first pixel is 120, the average pixel value of the second pixel is 130, and the reference pixel value is 125. Then the first color difference value is |120 - 125| = 5, and the second color difference value is |130 - 125| = 5.
[0038] In step S1323, the weight of the first pixel in the first splicing and fusion region is determined based on the color difference value and the preset tolerance parameter value. The preset tolerance parameter can be a pre-defined range of allowable color differences, used to determine whether the color difference is within an acceptable range, thereby determining the weight of the pixel.
[0039] In this embodiment, the weight of the first pixel in the first stitching and fusion region is determined based on the color difference value and a preset tolerance parameter. If the color difference value is within the preset tolerance parameter range, it indicates that the pixel in that region is relatively harmonious with the overall composition, and a relatively appropriate weight can be given; if it exceeds the range, the weight is adjusted appropriately. For example, if the preset tolerance parameter is 8, since the previously calculated color difference value of 5 is within the tolerance range, the first pixel in the first stitching and fusion region can be assigned a weight of 0.7.
[0040] In step S1324, the weight of the first pixel in each of the stitching and fusion regions from the side closer to any lens side to the side farther away from any lens side is determined based on the preset region weight difference and the weight of the first pixel in the first stitching and fusion region.
[0041] The preset region weight difference is the difference in the weight of the first pixel between adjacent stitching and fusion regions, which is used to ensure that the weight gradually decreases from the side closer to the lens to the side farther away.
[0042] In this embodiment, the weights of the first pixels in other stitching and fusion regions are determined sequentially based on the preset region weight difference and the weight of the first pixel in the first stitching and fusion region. The preset region weight difference ensures that the weights gradually decrease from the side closer to the lens to the side farther away. For example, if the preset region weight difference is 0.05 and the weight of the first pixel in the first stitching and fusion region is 0.7, then the weight of the first pixel in the second stitching and fusion region is 0.7 - 0.05 = 0.65, and so on, thus achieving a gradual decrease in weights.
[0043] The aforementioned technical solution effectively eliminates brightness differences and texture breaks at the stitching points of underwater images by subdividing the stitching area and determining weights based on color differences. Gradually decreasing the weight of the first pixel from the side closer to the lens to the side farther away results in a natural and smooth stitching transition, avoiding obvious stitching artifacts and improving the quality of panoramic underwater images. In fish detection, it can more realistically and clearly present the distribution and dynamics of fish underwater.
[0044] Optionally, in step S131, determining the true coordinates of the first pixel and the second pixel in the world coordinate system includes: In step S1311, the true coordinates of the first pixel in the world coordinate system are determined based on the first pixel coordinates of the first pixel in the first pixel coordinate system corresponding to any lens and the intrinsic and extrinsic parameter matrix of any lens. In this embodiment, the first pixel has specific coordinate values in the first pixel coordinate system. Using the intrinsic parameter matrix of any lens, the first pixel can be transformed from the pixel coordinate system to the camera coordinate system, taking into account the camera's internal optical characteristics, such as focal length. Then, using the extrinsic parameter matrix of the lens, the coordinates in the camera coordinate system are transformed to the world coordinate system. The extrinsic parameter matrix contains the camera's position and orientation information in the world coordinate system. Through these two transformations, the true coordinates of the first pixel in the world coordinate system can be accurately determined.
[0045] In step S1312, the true coordinates of the second pixel in the world coordinate system are determined based on the second pixel coordinates of the second pixel in the second pixel coordinate system corresponding to the other lens and the intrinsic and extrinsic parameter matrix of the other lens.
[0046] In this embodiment, the second pixel is first transformed from the pixel coordinate system to the camera coordinate system corresponding to the other lens based on the intrinsic parameter matrix of the other lens, taking into account the internal optical characteristics of the lens. Then, the coordinates in the camera coordinate system are transformed to the world coordinate system using the extrinsic parameter matrix of the lens, thereby determining the true coordinates of the second pixel in the world coordinate system.
[0047] The above technical solution determines the true coordinates of the first and second pixels in the world coordinate system. When stitching underwater images, it can accurately find the positional relationship of corresponding pixels in the two lens images, making the stitching more precise and effectively avoiding problems such as stitching misalignment and ghosting caused by inaccurate positioning, thereby generating high-quality panoramic underwater images.
[0048] Optionally, in step S14, the step of previewing and displaying the fish school detection image based on the panoramic underwater image corresponding to the multiple discrete acquisition times and the preview mode selected by the user includes: In step S141, based on the panoramic underwater image corresponding to the multiple discrete acquisition times, the panoramic underwater image is processed by OpenGL texture mapping and rendering pipeline to be mapped onto the surface of a 3D panoramic model as a texture. The 3D panoramic model is a panoramic spherical or cylindrical model constructed based on vertex coordinates defined in three-dimensional space. Texture mapping is a technique that applies 2D image data (textures) to the surface of a 3D model, giving the 3D model a more realistic appearance and detail, enhancing visual effects. The rendering pipeline is the stage in OpenGL that processes graphics data, including vertex processing, primitive assembly, rasterization, etc., ultimately converting the 3D model into a 2D image displayed on the screen. A 3D panoramic model is a model constructed based on vertex coordinates defined in 3D space to display a panorama; for example, a panoramic spherical or cylindrical model, thus providing a comprehensive visual experience.
[0049] In this embodiment, the OpenGL texture mapping and rendering pipeline is crucial. First, panoramic underwater images corresponding to multiple discrete acquisition moments are loaded as texture data. Then, based on the vertex coordinates of the 3D panoramic model (panoramic sphere or cylinder), the mapping method of the texture on the model surface is determined. In the rendering pipeline, the vertex processing stage transforms and performs lighting calculations on the model vertices; the primitive assembly stage combines vertices into primitives; the rasterization stage converts primitives into pixels, and according to the texture mapping relationship, assigns the pixel values of the panoramic underwater image to the corresponding screen pixels, ultimately mapping the panoramic underwater image onto the surface of the 3D panoramic model.
[0050] In step S142, when the user selects VR interactive preview mode, the VR panoramic preview based on perspective interaction is used to preview and display the fish school detection image, and in response to the user's perspective adjustment operation, the observation perspective of the 3D panoramic model is switched. Among them, the VR interactive preview mode uses virtual reality technology to allow users to interact with panoramic images through interactive devices (such as head-mounted displays, controllers, etc.), freely adjust the viewing angle, and enhance the sense of immersion.
[0051] In this embodiment of the disclosure, when the user selects the VR interactive preview mode, the VR panoramic preview mechanism based on viewpoint interaction is activated. Sensors acquire the user's viewpoint adjustment information, such as the direction and angle of head rotation. The new viewing angle of the 3D panoramic model is then calculated in real time, the rendering pipeline is reprocessed, and the displayed content is updated, allowing the user to observe the fish detection image from different angles and achieve an immersive interactive experience.
[0052] In step S143, when the user selects a planar panoramic preview mode, an equidistant cylindrical projection algorithm is used to map the texture information on the 3D panoramic model to a two-dimensional plane according to the equidistant cylindrical projection rule, and a fish school detection image preview is displayed based on the image on the two-dimensional plane.
[0053] Among them, the equidistant cylindrical projection algorithm projects points on a sphere onto a plane to convert panoramic images from a three-dimensional spherical model into a two-dimensional planar image.
[0054] In this embodiment, an equidistant cylindrical projection algorithm is used in the planar panoramic preview mode. This algorithm projects the latitude and longitude coordinates of each point on the 3D panoramic model (sphere) onto a two-dimensional plane according to a specific mathematical formula. During the projection process, the equidistant characteristics of the image are maintained, meaning that the distances corresponding to the same angular intervals are equal before and after projection. After mapping the texture information on the 3D panoramic model onto the two-dimensional plane according to this rule, a planar panoramic image is generated for previewing and displaying fish swarm detection images.
[0055] The aforementioned technical solutions offer users diverse viewing experiences of fish school detection images through different preview display methods. In VR interactive preview mode, users can observe fish schools from different angles in an immersive way, enhancing their perception and understanding of the underwater scene and helping them accurately analyze fish behavior and distribution. The planar panoramic preview mode facilitates quick viewing of panoramic images on ordinary devices, improving the efficiency of information dissemination.
[0056] Optionally, the target calibration parameters are obtained by calibration in the following manner: The geometric center of the lens is determined by drawing horizontal and vertical crosshairs on the original imaging image captured by the lens. The raw image is the unprocessed image directly captured by the lens, containing the original information of the scene as captured. Crosshairs are two perpendicular lines drawn on the image to aid in positioning and measurement. The geometric center is the symmetrical center point in the image; for regularly shaped image areas, it reflects the symmetrical characteristics of the lens's imaging.
[0057] In this embodiment, horizontal and vertical crosshairs are drawn on the original image captured by the lens, and these two lines intersect each other perpendicularly. Since the image is ideally symmetrical, the intersection of the crosshairs is the geometric center of the image. This allows for accurate positioning of the geometric center of the lens image, and calibration operations are performed using this geometric center as a reference for coordinate adjustments and parameter calculations.
[0058] For example, when photographing a regular square object, the original image captured by the lens may have some deviation. By drawing horizontal and vertical crosshairs on the image, and assuming the square image is basically symmetrical in the horizontal and vertical directions, the intersection of these crosshairs is the geometric center of the square image. This center point reflects the position of the center of symmetry on the image plane when the lens is forming the image.
[0059] Using the reference calibration center formed by the intersection of auxiliary lines on the inner wall of the cube calibration device as a reference, adjust the coordinates of the original image acquired by the lens until the position of the geometric center and the reference calibration center meets the preset position requirements, and use the translation amount of the coordinates as the corresponding center offset parameter of the lens. The cube calibration device is a specific device used for lens calibration. Its inner wall has auxiliary lines that intersect to form a reference calibration center, providing a standard reference for lens calibration. The reference calibration center is the point formed by the intersection of the auxiliary lines on the inner wall of the cube calibration device, serving as the benchmark for adjusting the lens image coordinates and ensuring that the lens image is aligned with the standard position. The center offset parameter is the numerical value of the coordinate translation when the position of the lens's geometric center and the reference calibration center does not meet the preset requirements; it describes the offset of the lens center relative to the standard position.
[0060] In this embodiment, the coordinates of the original image acquired by the lens are adjusted based on the reference calibration center formed by the intersection of auxiliary lines on the inner wall of the cube calibration device. By continuously changing the coordinate position of the image, the geometric center determined by the lens is gradually moved closer to the reference calibration center until the positions of the two meet the preset position requirements. During this process, the translation amount of the coordinates is recorded. The translation amount reflects the degree of offset of the lens center relative to the reference calibration center and is used as the center offset parameter of the corresponding lens for correcting the lens imaging position.
[0061] For example, within a cube-shaped calibration device, the intersecting auxiliary lines on its inner walls form a clearly defined reference calibration center. The geometric center of the image captured by the lens deviates slightly from this reference calibration center; suppose it's offset by 5 pixels horizontally and 3 pixels vertically. By adjusting the image coordinates to make the geometric center coincide with the reference calibration center, the image is translated by -5 pixels horizontally and -3 pixels vertically. These two translation amounts are the center offset parameters.
[0062] For example, the two center points are aligned by translation adjustment: using the real center of the scene formed by the intersection of the auxiliary lines on the inner wall of the cube as a reference, the x-axis and y-axis translation of the original image are gradually adjusted until the intersection of the crosshair auxiliary lines of the original image completely coincides with the real center of the scene. The coordinates corresponding to the x and y translation adjustments at this time are the real center (x, y) coordinates of the shot.
[0063] Given the center offset parameter, the original image is scaled to determine the effective radius of the lens's field of view. Scaling refers to the operation of enlarging or reducing the size of an image according to a certain ratio to change the image size and adapt to different calibration requirements and display scenarios. The effective field of view radius is the radius length corresponding to the area from the lens's geometric center to the image edge that can be effectively imaged, reflecting the effective range of the lens's imaging.
[0064] In this embodiment, the original image is scaled. By gradually changing the image size, the imaging performance at different sizes is observed. When the image size is adjusted to a suitable level, the radius corresponding to the area from the lens's geometric center to the image edge that can be effectively imaged in sharp focus is measured, using the lens's geometric center as the center. This radius determines the effective field of view radius of the lens. The effective field of view radius determines the range that the lens can effectively cover in actual imaging.
[0065] Given a fixed effective field of view radius, the original images captured by the two lenses are rendered onto a spherical circle model for Euler angle calibration to obtain the Euler angle parameters of the lenses. The target calibration parameters include the center offset parameter, the effective field of view radius, and the Euler angle parameter.
[0066] The spherical circle model is a mathematical model used to describe image rendering in 3D space. It renders the image onto a sphere to simulate human visual perception and present the scene more realistically. Euler angle calibration determines the rotation angles of an object around three coordinate axes in 3D space (Euler angles). It describes the object's spatial attitude and orientation, and is used to accurately determine the spatial position and angle of the lens. Euler angle parameters are the numerical values of the angles by which the lens rotates around three coordinate axes in 3D space, used to describe the lens's spatial attitude and orientation.
[0067] In this embodiment, the original images captured by the two lenses are rendered into a spherical circular model. Within the spherical model, Euler angles are used to describe the lens's attitude in three-dimensional space by analyzing the relative position and orientation of the images. The images from the two lenses are matched and adjusted within the spherical model, and the rotation angle of each lens around three coordinate axes (typically roll, pitch, and yaw axes) is calculated. Euler angle parameters accurately describe the lens's attitude in space.
[0068] For example, after rendering the images onto a spherical model, differences were found in the positions and angles of the two images on the sphere. Through calculation and analysis, it was determined that one lens rotated 10 degrees around the roll axis, 5 degrees around the pitch axis, and -3 degrees around the yaw axis; these angles are the Euler angle parameters of that lens. The other lens also has corresponding Euler angle parameters. These parameters can be used to adjust the lens attitude, ensuring that the two images are correctly matched on the sphere.
[0069] The target calibration parameters obtained from the above technical solution, including center offset parameters, effective field of view radius, and Euler angle parameters, provide a comprehensive and precise calibration for lens imaging. The center offset parameter corrects the geometric center position of the lens imaging, making the image more accurate in coordinates; the effective field of view radius determines the effective imaging range of the lens, avoiding interference from invalid information; and the Euler angle parameter accurately describes the lens's attitude in three-dimensional space, ensuring the synergy and consistency of images acquired by the binocular lenses. This improves the quality and accuracy of lens imaging, enabling the acquired images to more realistically and accurately reflect the actual scene, providing reliable basic data for subsequent image processing, 3D reconstruction, and other applications, and enhancing the performance and reliability of the entire system.
[0070] Optionally, in step S12, calibrating the underwater image using pre-calibrated target calibration parameters to obtain the target underwater image at the corresponding acquisition time includes: In step S121, the underwater images corresponding to the two lenses are translated according to the center offset parameter so that the geometric center of the lens is aligned with the center of the preset canvas. In this embodiment, a pre-calibrated center offset parameter is used to define the offset between the lens's geometric center and the center of a preset canvas in both the horizontal and vertical directions. For underwater images corresponding to two lenses, a translation operation is performed according to this offset. In the horizontal direction, if the center offset parameter indicates that the image center has shifted to the right by a certain number of pixels, the image is moved to the left by the corresponding number of pixels; the same applies to the vertical direction. This translation aligns the lens's geometric center with the center of the preset canvas, eliminating image position deviations caused by lens mounting or imaging characteristics. This provides an accurate positional basis for subsequent image processing, ensuring that the image is correctly displayed within the preset frame.
[0071] For example, when shooting underwater scenes, due to lens mounting issues, the center of the image captured by one lens is offset 10 pixels horizontally to the right and 5 pixels vertically downwards relative to the center of the preset canvas. Based on the center offset parameters, a translation operation is performed on the image, moving it 10 pixels to the left and 5 pixels upwards. This aligns the geometric center of the image with the center of the preset canvas, ensuring the image is in the correct position and preventing positional deviations from affecting the overall effect.
[0072] In step S122, based on the effective field of view alarm, the translated underwater image is scaled, and the invalid image that exceeds the preset canvas is cropped out to obtain the effective area image; In this embodiment, the translated underwater image is scaled according to the effective radius of the field of view. If the image size corresponding to the effective radius of the field of view is smaller than the preset canvas size, the image is enlarged; if it is larger than the preset canvas size, the image is reduced so that the effective portion of the image can fit the preset canvas. Then, it is checked whether the image exceeds the preset canvas range, and the excess portion is cropped out. Because the excess portion cannot be effectively displayed within the preset canvas and may contain interference information, the cropped effective area image meets the size requirements of the preset canvas while ensuring image quality.
[0073] For example, if the effective field of view radius of the translated underwater image is large, exceeding the size of the preset canvas, the image is scaled down proportionally based on the effective field of view radius to make it roughly fit the preset canvas. However, after scaling down, it is found that parts of the four corners of the image still extend beyond the preset canvas; these excess, invalid image portions are cropped out. The final effective area image is exactly within the preset canvas, retaining the image's effective information and clearly displaying the main content of the underwater scene.
[0074] In step S123, the effective area image is mapped onto the 3D semicircular model textures corresponding to the two lenses, and the attitude of the 3D semicircular model is adjusted according to the Euler angle parameter to obtain the target underwater image at the corresponding acquisition time.
[0075] Among them, the 3D semi-circular model texture is the texture information used to display images on the surface of the 3D semi-circular model. It maps the image onto the model surface, so that the model presents the corresponding visual effect.
[0076] In this embodiment, the effective area image is mapped onto the textures of the 3D semicircular models corresponding to the two lenses, allowing the image to fit the surface of the 3D semicircular models. Then, the pose of the 3D semicircular models is adjusted according to the Euler angle parameters. The Euler angle parameters define the rotation angle of the lens in three-dimensional space. By applying these angles to the 3D semicircular models, the rotation angle of the models is made consistent with the actual pose of the lenses. Thus, after mapping and pose adjustment, the image presented by the 3D semicircular models can accurately reflect the actual underwater scene at the corresponding acquisition time, obtaining the target underwater image.
[0077] In this embodiment, to eliminate lens center offset and uncertainty in the effective image range based on calibration parameters, the original images of the left and right lenses undergo the same operational process. First, based on the true center (x, y) coordinates of the left / right lenses, the corresponding original images are translated to ensure that the true center of the lenses is perfectly aligned with the center of the preset canvas, thus ensuring that the imaging reference of the two lenses remains consistent. Subsequently, based on the effective radius parameters of the calibrated left / right lenses, the translated images are scaled and adjusted. Since the effective radius defines the effective range of lens imaging, invalid image content exceeding the canvas after scaling is naturally discarded, retaining only the effective imaging area.
[0078] Furthermore, even after translation and scaling, images still exhibit geometric distortion due to lens pose deviations. This distortion is corrected through 3D rendering using calibrated pose parameters, restoring the image to a spatial pose consistent with the real scene. For example, a semi-circular 3D model is first constructed as the rendering medium to simulate the imaging projection characteristics of a binocular panoramic lens, matching the lens's wide field-of-view imaging effect. Then, the left and right camera images are mapped onto their respective 3D semi-circular model textures. The calibrated pose parameters (yaw, pitch, and roll) are then used to adjust the model's pose: adjusting the yaw angle corrects the horizontal offset, adjusting the pitch angle corrects the vertical offset, and adjusting the roll angle corrects the rotational offset. Ultimately, the rendered left and right camera images perfectly match the spatial geometry of the real scene, resulting in a corrected standard image.
[0079] The above technical solution translates the image based on the center offset parameter, eliminating image position deviations caused by lens installation or imaging characteristics, ensuring accurate image display within the preset canvas. Secondly, by scaling and cropping the image according to the effective field of view radius, the resulting effective area image meets processing size requirements while maintaining image quality and removing interference from invalid information. Finally, the effective area image is mapped onto a 3D semi-circular model and its attitude is adjusted. Combined with Euler angle parameters, the image presented by the model accurately reflects the attitude and content of the actual underwater scene, improving image realism and accuracy, and enhancing the performance and reliability of the entire underwater image processing system.
[0080] This disclosure also provides a fish school detection and preview device based on binocular camera image stitching, see [link to relevant documentation]. Figure 2 As shown, the device includes: The acquisition module 210 is configured to acquire underwater images captured by the two lenses of the underwater binocular camera at multiple discrete acquisition times. The calibration module 220 is configured to calibrate the underwater image using pre-calibrated target calibration parameters to obtain the target underwater image at the corresponding acquisition time. The target calibration parameters are obtained by calibrating multiple core parameters sequentially based on multiple sets of auxiliary lines drawn on the inner wall of the cube calibration device as geometric references. The stitching and fusion module 230 is configured to perform mirror linear stitching and fusion of pixels in the target underwater image corresponding to one of the two lenses that are located between the preset fusion percentage and the preset fusion percentage and the reference radius, and pixels in the target underwater image corresponding to the other lens that are in the overlapping area, through gradient weighting, to obtain the panoramic underwater image at the acquisition time. The preview module 240 is configured to preview and display fish school detection images based on the panoramic underwater images corresponding to the multiple discrete acquisition times and the preview mode selected by the user.
[0081] Optionally, the splicing and fusion module 230 is configured as follows: Determine the true coordinates of the first pixel and the second pixel in the world coordinate system, wherein the first pixel is a pixel in the target underwater image corresponding to either of the two lenses that is located between a preset fusion percentage and a preset reference radius, and the second pixel is a pixel in the target underwater image corresponding to the other lens that is located in the overlapping area. Based on the true coordinates of each first pixel, a weight corresponding to the first pixel is determined, and based on the weight of the first pixel with the same true coordinates as each second pixel, a weight corresponding to the second pixel is determined, wherein the weight of the first pixel gradually decreases from the side of the overlapping region closer to the lens side to the side farther away from the lens side, and the weight of the second pixel gradually increases as the weight of the first pixel with the same true coordinates gradually decreases. Based on the weight of each first pixel and the weight of each second pixel, the first pixels and second pixels with the same real coordinates are stitched together and fused to obtain the panoramic underwater image at the time of acquisition.
[0082] Optionally, the splicing and fusion module 230 is configured as follows: The region between the preset fusion percentage and the preset baseline radius is divided into a preset number of splicing and fusion regions of equal width; A first color difference value is determined based on the difference between the average pixel value of a first pixel in the first stitching and fusion region from the side closer to the lens side to the side farther away from the lens side of the overlapping region and a reference pixel value; and a second color difference value is determined based on the difference between the average pixel value of a second pixel in the first stitching and fusion region and the reference pixel value, wherein the reference pixel value is the average pixel value of the first pixel and the second pixel between the preset fusion percentage and the preset reference radius; The weight of the first pixel in the first splicing and fusion region is determined based on the color difference value and the preset tolerance parameter. Based on the preset regional weight difference and the weight of the first pixel in the first stitching and fusion region, the weight of the first pixel in each stitching and fusion region from the side of the overlapping region closest to any lens side to the side furthest from any lens side is determined.
[0083] Optionally, the splicing and fusion module 230 is configured as follows: Based on the first pixel coordinates in the first pixel coordinate system corresponding to any lens and the intrinsic and extrinsic parameter matrix of any lens, determine the true coordinates of the first pixel in the world coordinate system; Based on the second pixel coordinates in the second pixel coordinate system corresponding to the other lens and the intrinsic and extrinsic parameter matrix of the other lens, the true coordinates of the second pixel in the world coordinate system are determined.
[0084] Optionally, the preview module 240 is configured to: Based on the panoramic underwater images corresponding to the multiple discrete acquisition times, the panoramic underwater images are processed by OpenGL texture mapping and rendering pipeline, and mapped as textures onto the surface of a 3D panoramic model. The 3D panoramic model is a panoramic spherical or cylindrical model constructed based on vertex coordinates defined in three-dimensional space. When the user selects VR interactive preview mode, a VR panoramic preview based on viewpoint interaction is displayed to preview images of fish school detection, and the viewing perspective of the 3D panoramic model is switched in response to the user's viewpoint adjustment operation; or, When the user selects a planar panoramic preview mode, an equidistant cylindrical projection algorithm is used to map the texture information on the 3D panoramic model to a two-dimensional plane according to the equidistant cylindrical projection rule, and a fish school detection image preview is displayed based on the image on the two-dimensional plane.
[0085] Optionally, the device includes a calibration module configured to calibrate the target calibration parameters in the following manner: The geometric center of the lens is determined by drawing horizontal and vertical crosshairs on the original imaging image captured by the lens. Using the reference calibration center formed by the intersection of auxiliary lines on the inner wall of the cube calibration device as a reference, adjust the coordinates of the original image acquired by the lens until the position of the geometric center and the reference calibration center meets the preset position requirements, and use the translation amount of the coordinates as the corresponding center offset parameter of the lens. Given the center offset parameter, the original image is scaled to determine the effective radius of the lens's field of view. Given a fixed effective field of view radius, the original images captured by the two lenses are rendered onto a spherical circle model for Euler angle calibration to obtain the Euler angle parameters of the lenses. The target calibration parameters include the center offset parameter, the effective field of view radius, and the Euler angle parameter.
[0086] Optionally, the calibration module 220 is configured to: Based on the center offset parameter, the underwater images corresponding to the two lenses are translated respectively so that the geometric center of the lens is aligned with the center of the preset canvas; Based on the effective field of view alarm, the translated underwater image is scaled, and invalid images that exceed the preset canvas are cropped out to obtain the effective area image; The effective area image is mapped onto the 3D semicircular model textures corresponding to the two lenses, and the attitude of the 3D semicircular model is adjusted according to the Euler angle parameters to obtain the target underwater image at the corresponding acquisition time.
[0087] Specific limitations regarding the fish school detection preview device based on binocular camera image stitching can be found in the limitations of the fish school detection preview method based on binocular camera image stitching mentioned above, and will not be repeated here. Each module in the aforementioned fish school detection preview device based on binocular camera image stitching can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0088] This disclosure also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the methods described in the foregoing embodiments.
[0089] This disclosure also provides an electronic device, including: A memory on which computer programs are stored; A processor for executing the computer program in the memory to implement the steps of any of the methods described in the foregoing embodiments.
[0090] Figure 3 The fish school detection and preview device 100 based on binocular camera image stitching, as shown, includes a processor 1001 and a memory 1003. The processor 1001 and the memory 1003 are connected, for example, via a bus 1002. Optionally, the fish school detection and preview device 100 based on binocular camera image stitching may further include a communication component, which can be used for data interaction between the device 100 and other devices, such as sending or receiving data. It should be noted that in actual scheduling, the communication component is not limited to one, and the structure of this fish school detection and preview device 100 based on binocular camera image stitching does not constitute a limitation on the embodiments of this application.
[0091] Processor 1001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 1001 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.
[0092] Bus 1002 may include a pathway for transmitting information between the aforementioned components. Bus 1002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Bus 1002 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 3 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0093] The memory 1003 may be ROM (Read Only Memory) or other types of static storage devices capable of storing static information and instructions, RAM (Random Access Memory) or other types of dynamic storage devices capable of storing information and instructions, or EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium capable of carrying or storing program code and capable of being read by a computer, without limitation herein.
[0094] The memory 1003 is used to store program code for executing embodiments of the present disclosure, and its execution is controlled by the processor 1001. The processor 1001 is used to execute the program code stored in the memory 1003 to implement the steps shown in the aforementioned embodiments of the fish swarm detection preview method based on binocular camera image stitching.
[0095] This disclosure also provides a computer-readable storage medium storing program code. When the program code is executed by a processor, it can implement the steps and corresponding content of the aforementioned embodiment of the fish swarm detection and preview method based on binocular camera image stitching.
[0096] The preferred embodiments of the present disclosure have been described in detail above with reference to the accompanying drawings. However, the present disclosure is not limited to the specific details of the above embodiments. Within the scope of the technical concept of the present disclosure, various changes, modifications, substitutions and variations can be made to these embodiments, and all such changes, modifications, substitutions and variations fall within the protection scope of the present disclosure.
[0097] It should also be noted that the various specific technical features described in the above embodiments can be combined in any suitable manner without contradiction, and such combinations should also be considered as part of this disclosure. To avoid unnecessary repetition, this disclosure will not further describe the various possible combinations. The technical scope of this application is not limited to the contents of the specification, but must be determined according to the scope of the claims.
Claims
1. A method for fish school detection and preview based on image stitching from a binocular camera, characterized in that, The method includes: Acquire underwater images captured by the two lenses of an underwater binocular camera at multiple discrete acquisition times; The underwater image is calibrated using pre-calibrated target calibration parameters to obtain the target underwater image at the corresponding acquisition time. The target calibration parameters are obtained by calibrating multiple core parameters sequentially based on multiple sets of auxiliary lines drawn on the inner wall of the cube calibration device as geometric references. The pixels in the target underwater image corresponding to either of the two lenses that are between a preset fusion percentage and a preset reference radius are mirrored and linearly stitched together with the pixels in the overlapping area of the target underwater image corresponding to the other lens, using a gradient weighting method, to obtain the panoramic underwater image at the time of acquisition. Based on the panoramic underwater images corresponding to the multiple discrete acquisition times and the preview mode selected by the user, a fish school detection image preview is displayed.
2. The method according to claim 1, characterized in that, The step of taking pixels in the target underwater image corresponding to either of the two lenses that are located between a preset fusion percentage and a preset reference radius, and pixels in the overlapping area of the target underwater image corresponding to the other lens, and performing mirror linear stitching fusion using gradient weights to obtain the panoramic underwater image at the time of acquisition includes: Determine the true coordinates of the first pixel and the second pixel in the world coordinate system, wherein the first pixel is a pixel in the target underwater image corresponding to either of the two lenses that is located between a preset fusion percentage and a preset reference radius, and the second pixel is a pixel in the target underwater image corresponding to the other lens that is located in the overlapping area. Based on the true coordinates of each first pixel, a weight corresponding to the first pixel is determined, and based on the weight of the first pixel with the same true coordinates as each second pixel, a weight corresponding to the second pixel is determined, wherein the weight of the first pixel gradually decreases from the side of the overlapping region closer to the lens side to the side farther away from the lens side, and the weight of the second pixel gradually increases as the weight of the first pixel with the same true coordinates gradually decreases. Based on the weight of each first pixel and the weight of each second pixel, the first pixels and second pixels with the same real coordinates are stitched together and fused to obtain the panoramic underwater image at the time of acquisition.
3. The method according to claim 2, characterized in that, The weight of the first pixel gradually decreases from the side of the overlapping region closer to either lens side to the side farther away from either lens side, including: The region between the preset fusion percentage and the preset baseline radius is divided into a preset number of splicing and fusion regions of equal width; A first color difference value is determined based on the difference between the average pixel value of a first pixel in the first stitching and fusion region from the side closer to the lens side to the side farther away from the lens side of the overlapping region and a reference pixel value; and a second color difference value is determined based on the difference between the average pixel value of a second pixel in the first stitching and fusion region and the reference pixel value, wherein the reference pixel value is the average pixel value of the first pixel and the second pixel between the preset fusion percentage and the preset reference radius; The weight of the first pixel in the first splicing and fusion region is determined based on the color difference value and the preset tolerance parameter. Based on the preset regional weight difference and the weight of the first pixel in the first stitching and fusion region, the weight of the first pixel in each stitching and fusion region from the side of the overlapping region closest to any lens side to the side furthest from any lens side is determined.
4. The method according to claim 2, characterized in that, Determining the true coordinates of the first pixel and the second pixel in the world coordinate system includes: Based on the first pixel coordinates in the first pixel coordinate system corresponding to any lens and the intrinsic and extrinsic parameter matrix of any lens, determine the true coordinates of the first pixel in the world coordinate system; Based on the second pixel coordinates in the second pixel coordinate system corresponding to the other lens and the intrinsic and extrinsic parameter matrix of the other lens, the true coordinates of the second pixel in the world coordinate system are determined.
5. The method according to claim 1, characterized in that, The step of previewing and displaying fish school detection images based on the panoramic underwater images corresponding to the multiple discrete acquisition times and the preview mode selected by the user includes: Based on the panoramic underwater images corresponding to the multiple discrete acquisition times, the panoramic underwater images are processed by OpenGL texture mapping and rendering pipeline, and mapped as textures onto the surface of a 3D panoramic model. The 3D panoramic model is a panoramic spherical or cylindrical model constructed based on vertex coordinates defined in three-dimensional space. When the user selects VR interactive preview mode, a VR panoramic preview based on viewpoint interaction is displayed to preview images of fish school detection, and the viewing perspective of the 3D panoramic model is switched in response to the user's viewpoint adjustment operation; or, When the user selects a planar panoramic preview mode, an equidistant cylindrical projection algorithm is used to map the texture information on the 3D panoramic model to a two-dimensional plane according to the equidistant cylindrical projection rule, and a fish school detection image preview is displayed based on the image on the two-dimensional plane.
6. The method according to any one of claims 1-5, characterized in that, The target calibration parameters were obtained by calibration in the following manner: The geometric center of the lens is determined by drawing horizontal and vertical crosshairs on the original imaging image captured by the lens. Using the reference calibration center formed by the intersection of auxiliary lines on the inner wall of the cube calibration device as a reference, adjust the coordinates of the original image acquired by the lens until the position of the geometric center and the reference calibration center meets the preset position requirements, and use the translation amount of the coordinates as the corresponding center offset parameter of the lens. Given the center offset parameter, the original image is scaled to determine the effective radius of the lens's field of view. Given a fixed effective field of view radius, the original images captured by the two lenses are rendered onto a spherical circle model for Euler angle calibration to obtain the Euler angle parameters of the lenses. The target calibration parameters include the center offset parameter, the effective field of view radius, and the Euler angle parameter.
7. The method according to claim 6, characterized in that, The process of calibrating the underwater image using pre-defined target calibration parameters to obtain the target underwater image at the corresponding acquisition time includes: Based on the center offset parameter, the underwater images corresponding to the two lenses are translated respectively so that the geometric center of the lens is aligned with the center of the preset canvas; Based on the effective field of view alarm, the translated underwater image is scaled, and invalid images that exceed the preset canvas are cropped out to obtain the effective area image; The effective area image is mapped onto the 3D semicircular model textures corresponding to the two lenses, and the attitude of the 3D semicircular model is adjusted according to the Euler angle parameters to obtain the target underwater image at the corresponding acquisition time.
8. A fish school detection and preview device based on binocular camera image stitching, characterized in that, The device includes: The acquisition module is configured to acquire underwater images captured by the two lenses of an underwater binocular camera at multiple discrete acquisition times. The calibration module is configured to calibrate the underwater image using pre-calibrated target calibration parameters to obtain the target underwater image at the corresponding acquisition time. The target calibration parameters are obtained by calibrating multiple core parameters sequentially based on multiple sets of auxiliary lines drawn on the inner wall of the cube calibration device as geometric references. The stitching and fusion module is configured to perform mirror linear stitching and fusion of pixels in the target underwater image corresponding to one of the two lenses that are located between the preset fusion percentage and the preset fusion radius, and pixels in the target underwater image corresponding to the other lens that are in the overlapping area, using a gradient weighting method, to obtain a panoramic underwater image at the time of acquisition. The preview module is configured to preview and display fish school detection images based on the panoramic underwater images corresponding to the multiple discrete acquisition times and the preview mode selected by the user.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the method described in any one of claims 1-7.
10. A binocular camera, characterized in that, include: A memory on which computer programs are stored; A processor for executing the computer program in the memory to implement the steps of the method according to any one of claims 1-7.