Fish school detection preview method, device and equipment for improving anti-shake performance of underwater binocular camera

By acquiring inertial measurement unit data and synchronization timestamps from underwater binocular cameras, and combining this with parameters from a cube calibration device to generate an anti-shake matrix, the underwater camera shake problem was solved, enabling stable display of panoramic underwater images and high-quality fish detection.

CN121908122APending Publication Date: 2026-04-21SHENZHEN YINGZHI FUTURE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN YINGZHI FUTURE TECHNOLOGY CO LTD
Filing Date
2026-01-27
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Underwater binocular cameras are prone to shaking when acquiring images due to factors such as water flow, resulting in poor image quality and fish detection. Existing image stabilization matrix generation methods are not adaptable enough and cannot meet user experience requirements.

Method used

By acquiring inertial measurement unit data from an underwater binocular camera, adding synchronization timestamps, calibrating parameters using a cube calibration device, generating an image stabilization matrix for image calibration and stitching fusion, and combining this with user perspective adjustment operations, image stabilization processing for panoramic underwater images is achieved.

Benefits of technology

It improves image stability and realism, provides a more complete and clear underwater panoramic view, and enhances the user's fish detection effect and experience in VR interactive preview.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121908122A_ABST
    Figure CN121908122A_ABST
Patent Text Reader

Abstract

The invention relates to a fish school detection preview method, device and equipment for improving the anti-shake performance of an underwater binocular camera, and the method comprises the steps: obtaining the angular velocity data and acceleration data of an inertial measurement unit disposed at the underwater binocular camera when the underwater binocular camera collects an underwater image; synchronous timestamps are added to the underwater image, the angular velocity data and the acceleration data at the same acquisition moment; calibrating the underwater image through a pre-calibrated target calibration parameter to obtain a corresponding target underwater image; according to a preset fusion percentage and a preset reference radius, splicing and fusing the target underwater images of the two lenses at the same acquisition moment through a gradient weight to obtain a corresponding panoramic underwater image; generating an anti-shake matrix according to the angular velocity data and the acceleration data corresponding to each synchronization timestamp; and in a VR interactive preview mode, in response to a visual angle adjustment operation of a user, performing preview display of the fish school detection image according to the synchronization timestamp corresponding to the anti-shake matrix and the panoramic underwater image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a method, apparatus, and device for fish detection and preview that improves the image stabilization of underwater binocular cameras. Background Technology

[0002] In the fields of underwater detection and VR interactive preview, underwater binocular cameras are prone to shaking when acquiring images due to factors such as water flow, which seriously affects image quality and subsequent fish detection results. Currently, the methods for generating image stabilization matrices are relatively conventional and lack adaptability to camera shaking caused by complex underwater environments. They also struggle to accurately quantify the amplitude and direction of shaking, resulting in limited stabilization effects and failing to meet the user experience requirements for previewing and displaying fish detection images. Summary of the Invention

[0003] The purpose of this invention is to provide a method, apparatus, and device for previewing fish school detection that improves the image stabilization of underwater binocular cameras. The aim is to balance the stability of the image preview field of view and the image stabilization of underwater fish school images, thereby improving the user experience of fish school detection image preview display.

[0004] To achieve the above objectives, a first aspect of this disclosure provides a fish school detection preview method to improve the image stabilization of an underwater binocular camera, the method comprising: When an underwater binocular camera acquires an underwater image, it sets the angular velocity and acceleration data in the inertial measurement unit of the underwater binocular camera, and adds a synchronization timestamp to the underwater image, the angular velocity data, and the acceleration data at the same acquisition time. The underwater image is calibrated using pre-calibrated target calibration parameters to obtain the target underwater image at the corresponding acquisition time. The target calibration parameters are obtained by calibrating multiple core parameters sequentially based on multiple sets of auxiliary lines drawn on the inner wall of the cube calibration device as geometric references. Based on a preset fusion percentage and a preset reference radius, the target underwater images of the two lenses at the same acquisition time are mirrored and linearly stitched together using a gradient weighting method to obtain a panoramic underwater image at that acquisition time. The preset fusion percentage is determined based on the field of view of the lens. Based on the angular velocity data and acceleration data corresponding to each synchronization timestamp, an anti-shake matrix is ​​generated to counteract the angular motion of the underwater binocular camera. The anti-shake matrix is ​​a geometric transformation parameter used to quantify the amplitude and direction of the shaking of the underwater binocular camera when capturing the underwater image. In VR interactive preview mode, in response to the user's perspective adjustment operation, a fish school detection image preview is displayed according to the synchronization timestamp corresponding to the anti-shake matrix and the panoramic underwater image. The anti-shake matrix is ​​used to reverse the perspective of the panoramic underwater image.

[0005] Optionally, the step of responding to the user's perspective adjustment operation by performing a fish detection image preview display based on the synchronization timestamp corresponding to the image stabilization matrix and the panoramic underwater image includes: In response to user interaction, generate corresponding interaction commands; Based on the synchronization timestamp, the image stabilization matrix corresponding to the panoramic underwater image is invoked, and a rendering matrix is ​​generated based on the image stabilization matrix and the interactive operation instructions. Based on the rendering matrix, the viewing angle of the corresponding panoramic underwater image is adjusted, and the panoramic underwater image after the viewing angle adjustment is rendered, and a fish school detection image preview is displayed.

[0006] Optionally, generating the rendering matrix based on the stabilization matrix and the interactive operation command includes: The interactive operation instructions are parsed to determine the operation type and operation amount corresponding to the interactive operation. The operation type includes one or more of the following: screen panning, image scaling, and rotation around the world coordinate principal axis. Based on the operation amount, determine an adjustment value relative to the initial state corresponding to the operation type; The adjustment value is converted into a preview matrix corresponding to this interactive operation using a matrix transformation function; The rendering matrix is ​​determined by multiplying the stabilization matrix and the preview matrix.

[0007] Optionally, generating a stabilization matrix to counteract the angular motion of the underwater binocular camera based on the angular velocity data and acceleration data corresponding to each synchronization timestamp includes: The angular velocity data is transformed from the body coordinate system of the inertial measurement unit to the optical coordinate system of the underwater binocular camera, and the acceleration data is used to assist in correcting the attitude estimation error and correcting the alignment deviation between the body coordinate system and the optical coordinate system. The angular velocity data in the optical coordinate system is discretized into rotational changes over a continuous time period. The quaternion state is updated based on the rotational changes to characterize the three-dimensional rotational attitude of the camera. The accelerometer data is fused with the gravity vector to estimate and correct the attitude drift error generated by the integration of the inertial measurement unit. The acquisition time with the smallest motion amplitude is selected as the reference time. The relative rotation matrix of each acquisition time relative to the reference time is calculated. The relative rotation matrix is ​​used to describe the image plane rotation change caused by the angular motion of the underwater binocular camera. The image stabilization matrix corresponding to each lens in the underwater binocular camera is determined based on the inverse of the relative rotation matrix.

[0008] Optionally, the step of performing mirror linear stitching and fusion of the target underwater images of the two lenses at the same acquisition time according to a preset fusion percentage and a preset reference radius, using a gradual weighting method, to obtain a panoramic underwater image at that acquisition time, includes: The pixels in the target underwater image corresponding to either of the two lenses that are located between the preset fusion percentage and the preset reference radius are mirrored and linearly stitched together with the pixels in the target underwater image corresponding to the other lens in the overlapping area, using a gradient weighting method, to obtain the panoramic underwater image at the time of acquisition. The weight of the pixel corresponding to any one of the lenses gradually decreases from the side of the overlapping region closer to the lens to the side farther away from the lens, while the weight of the pixel corresponding to the other lens gradually increases as the weight of the pixel in any one of the lenses gradually decreases.

[0009] Optionally, the target calibration parameters are obtained by calibration in the following manner: The geometric center of the lens is determined by drawing horizontal and vertical crosshairs on the original imaging image captured by the lens. Using the reference calibration center formed by the intersection of auxiliary lines on the inner wall of the cube calibration device as a reference, adjust the coordinates of the original image acquired by the lens until the position of the geometric center and the reference calibration center meets the preset position requirements, and use the translation amount of the coordinates as the corresponding center offset parameter of the lens. Given the center offset parameter, the original image is scaled to determine the effective radius of the lens's field of view. Given a fixed effective field of view radius, the original images captured by the two lenses are rendered onto a spherical circle model for Euler angle calibration to obtain the Euler angle parameters of the lenses. The target calibration parameters include the center offset parameter, the effective field of view radius, and the Euler angle parameter.

[0010] Optionally, calibrating the underwater image using pre-defined target calibration parameters to obtain the target underwater image at the corresponding acquisition time includes: Based on the center offset parameter, the underwater images corresponding to the two lenses are translated respectively so that the geometric center of the lens is aligned with the center of the preset canvas; Based on the effective field of view alarm, the translated underwater image is scaled, and invalid images that exceed the preset canvas are cropped out to obtain the effective area image; The effective area image is mapped onto the 3D semicircular model textures corresponding to the two lenses, and the attitude of the 3D semicircular model is adjusted according to the Euler angle parameters to obtain the target underwater image at the corresponding acquisition time.

[0011] A second aspect of this disclosure provides a fish detection and preview device for improving the image stabilization of an underwater binocular camera, the device comprising: The acquisition module is configured to acquire angular velocity and acceleration data set in the inertial measurement unit of the underwater binocular camera when the underwater binocular camera acquires an underwater image, and to add a synchronization timestamp to the underwater image, the angular velocity data and the acceleration data at the same acquisition time. The calibration module is configured to calibrate the underwater image using pre-calibrated target calibration parameters to obtain the target underwater image at the corresponding acquisition time. The target calibration parameters are obtained by calibrating multiple core parameters sequentially based on multiple sets of auxiliary lines drawn on the inner wall of the cube calibration device as geometric references. The stitching and fusion module is configured to perform mirror linear stitching and fusion of the target underwater images of the two lenses at the same acquisition time according to a preset fusion percentage and a preset reference radius, through gradual weighting, to obtain a panoramic underwater image at that acquisition time. The preset fusion percentage is determined according to the field of view of the lens. The generation module is configured to generate an anti-shake matrix to counteract the angular motion of the underwater binocular camera based on the angular velocity data and acceleration data corresponding to each synchronization timestamp. The anti-shake matrix is ​​a geometric transformation parameter used to quantify the amplitude and direction of the shaking of the underwater binocular camera when capturing the underwater image. The preview module is configured to, in VR interactive preview mode, respond to the user's perspective adjustment operation and display a fish detection image preview based on the synchronization timestamp corresponding to the anti-shake matrix and the panoramic underwater image. The anti-shake matrix is ​​used to reverse the perspective of the panoramic underwater image.

[0012] A third aspect of this disclosure provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps of the method described in any of the first aspects.

[0013] A fourth aspect of this disclosure provides an electronic device, comprising: A memory on which computer programs are stored; A processor for executing the computer program in the memory to implement the steps of the method of any one of the first aspects.

[0014] This invention provides a method, apparatus, and device for fish detection and preview to improve the image stabilization performance of underwater binocular cameras. Compared with existing technologies, it has the following advantages: By acquiring angular velocity and acceleration data from underwater binocular cameras during image acquisition and adding synchronization timestamps, precise temporal correspondence between images and motion data is ensured, improving the accuracy of image stabilization. In the image calibration stage, target calibration parameters obtained using auxiliary lines on the inner wall of a cube calibration device enable more accurate underwater image calibration, effectively eliminating image distortion caused by the complex underwater environment, improving image quality, and making the target underwater image more realistically reflect the underwater scene. Furthermore, in image stitching and fusion, a preset fusion percentage is determined based on the lens's field of view, combined with a preset reference radius and gradient weights for mirror linear stitching and fusion, reducing stitching artifacts and improving the realism of the generated panoramic underwater image, providing users with a more complete and clear underwater panoramic view. For image stabilization, an image stabilization matrix is ​​generated based on the motion data corresponding to the synchronization timestamps, accurately quantifying the amplitude and direction of camera shake. During VR interactive preview, the anti-shake matrix is ​​used to reverse the viewing angle of the panoramic underwater image, effectively counteracting the camera's angular motion and greatly improving image stability. This allows users to smoothly and clearly preview and display fish detection images, enhancing the user experience and underwater detection effect.

[0015] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description

[0016] The accompanying drawings are provided to further illustrate the present disclosure and form part of the specification. They are used together with the following detailed description to explain the present disclosure, but do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart illustrating a method for improving the image stabilization of an underwater binocular camera for fish detection and preview, according to an embodiment of the instruction manual.

[0017] Figure 2 A block diagram of a fish detection preview device for improving the image stabilization of an underwater binocular camera, as shown in the embodiments of the specification.

[0018] Figure 3 This is a block diagram of another fish detection and preview device for improving the image stabilization of an underwater binocular camera, as shown in the embodiments of the specification. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0020] This disclosure provides a fish school detection preview method to improve the image stabilization of an underwater binocular camera. The method aims to use an image stabilization matrix to reverse the viewing angle of the panoramic underwater image during VR interactive preview, effectively counteracting camera angular motion, greatly improving image stability, and allowing users to smoothly and clearly preview and display fish school detection images, thereby enhancing user experience and underwater detection effectiveness. Figure 1 This is a flowchart illustrating a method for improving the image stabilization of an underwater binocular camera for fish detection and preview, according to one embodiment. The method includes: In step S11, the angular velocity data and acceleration data set in the inertial measurement unit of the underwater binocular camera are obtained when the underwater binocular camera acquires an underwater image, and a synchronization timestamp is added to the underwater image, the angular velocity data and the acceleration data at the same acquisition time. The synchronization timestamp is a unified time identifier added to underwater images, angular velocity data, and acceleration data acquired at the same time. The time deviation between IMU data and video frames is controlled within milliseconds to avoid image stabilization failure or image distortion due to data misalignment. In this embodiment of the disclosure, when the underwater binocular camera acquires images, the IMU simultaneously measures angular velocity and acceleration. Adding a synchronization timestamp establishes a temporal correlation between the underwater images, angular velocity data, and acceleration data. Through the synchronization timestamp, the camera motion state at the corresponding moment in each underwater image can be associated.

[0021] In this disclosure, a built-in 6-axis IMU sensor can simultaneously acquire three-axis angular velocity data and three-axis acceleration data. The angular velocity data reflects the camera's rotational motion around the x, y, and z axes, while the acceleration data helps correct attitude estimation errors. By capturing camera motion details at a high video frame rate, it provides accurate and continuous raw data support for attitude calculation, ensuring effective response to instantaneous shake. The 6-axis IMU sensor achieves an optimal balance between cost and performance and is the mainstream hardware choice for panoramic camera image stabilization solutions.

[0022] It can be explained that underwater binocular cameras capture images of schools of fish, while the IMU records motion data. Adding synchronization timestamps clarifies the correspondence between images and motion data, such as knowing the camera's rotation angle and acceleration at the moment a particular image is captured.

[0023] In this step, the underwater binocular camera can be a panoramic camera. This allows the panoramic camera to simultaneously acquire video frames and, via a 6-axis IMU, collect real-time angular velocity and acceleration data. A synchronization timestamp is added to each set of IMU data and its corresponding video frame. This timestamp is crucial for subsequent data correlation, ensuring that each video frame matches the camera's motion data at the moment of capture, providing a time reference for precise image stabilization.

[0024] In step S12, the underwater image is calibrated using pre-calibrated target calibration parameters to obtain the target underwater image at the corresponding acquisition time. The target calibration parameters are obtained by calibrating multiple core parameters sequentially based on multiple sets of auxiliary lines drawn on the inner wall of the cube calibration device as geometric references. In this embodiment, the auxiliary lines on the inner wall of the cube calibration device are used as a geometric reference to calibrate core parameters such as lens distortion and camera intrinsic parameters to obtain target calibration parameters. Using these parameters to correct underwater images can eliminate the effects of lens distortion and other factors, making the images more accurately reflect the real scene. Targeted geometric transformation calibration can be performed on the acquired raw images to eliminate image distortion and positional deviations caused by errors.

[0025] The target calibration parameters may include: left camera parameters: left camera center x-coordinate, left camera center y-coordinate, left camera effective radius, left camera roll angle, left camera yaw angle, and left camera pitch angle; and right camera parameters: right camera center x-coordinate, right camera center y-coordinate, right camera effective radius, right camera roll angle, right camera yaw angle, and right camera pitch angle.

[0026] In step S13, based on the preset fusion percentage and the preset reference radius, the target underwater images of the two lenses at the same acquisition time are mirrored and linearly stitched together using a gradient weighting method to obtain a panoramic underwater image at that acquisition time. The preset fusion percentage is determined based on the field of view of the lens. In this embodiment, a preset fusion percentage is determined based on the lens's field of view. Combined with a preset reference radius, a gradual weighting method is used to linearly stitch and fuse the target underwater images acquired by the two lenses at the same acquisition time. The gradual weighting ensures a natural transition at the stitching point, avoiding obvious stitching marks, resulting in a panoramic underwater image and expanding the observation range.

[0027] In step S14, based on the angular velocity data and acceleration data corresponding to each synchronization timestamp, a stabilization matrix is ​​generated to counteract the angular motion of the underwater binocular camera. The stabilization matrix is ​​a geometric transformation parameter used to quantify the amplitude and direction of the shaking of the underwater binocular camera when capturing the underwater image. In this embodiment, the motion state of the underwater binocular camera is analyzed based on the angular velocity and acceleration data corresponding to the synchronization timestamps. A stabilization matrix is ​​generated through mathematical model calculations, which quantifies the amplitude and direction of camera shake during recording.

[0028] In step S15, in VR interactive preview mode, in response to the user's perspective adjustment operation, a fish school detection image preview is displayed according to the synchronization timestamp corresponding to the anti-shake matrix and the panoramic underwater image. The anti-shake matrix is ​​used to reverse the perspective of the panoramic underwater image.

[0029] In this embodiment, the IMU data corresponding to each timestamp is first processed using an IMU attitude estimation algorithm to generate a stabilization matrix to counteract camera angular motion. This stabilization matrix is ​​a set of geometric transformation parameters that quantifies the amplitude and direction of camera shake during shooting. Before video frame rendering, the stabilization matrix can be invoked to reverse the rendering perspective of the current frame, that is, to counteract the camera's own rotational shake through perspective shift, resulting in a stable visual effect in the playback image. In this way, panoramic stabilization can counteract the camera's own angular motion (rotational shake).

[0030] In this embodiment of the disclosure, when the user adjusts the viewing angle in VR interactive preview mode, the corresponding anti-shake matrix and panoramic underwater image are found based on the synchronization timestamp. The anti-shake matrix is ​​used to reverse the viewing angle of the panoramic underwater image, offsetting the effect of camera shake, so that the user can see a stable and clear image of fish detection, thus improving the interactive experience.

[0031] The aforementioned technical solution involves three stages: synchronous data acquisition, anti-shake matrix estimation, and real-time viewpoint adjustment. A high-frequency IMU sensor provides shake data input, an attitude estimation algorithm generates the basis for anti-shake control, and matrix fusion during the rendering stage achieves synergy between anti-shake and interaction. These three elements together constitute a complete anti-shake implementation system. Unlike traditional monocular anti-shake, panoramic anti-shake needs to balance the stability of the 360° field of view while ensuring the effectiveness of user interaction. Therefore, the technical implementation must balance the compatibility between anti-shake effect and viewpoint control.

[0032] The aforementioned technical solution acquires angular velocity and acceleration data from underwater binocular cameras during image acquisition and adds synchronization timestamps, ensuring precise temporal correspondence between images and motion data and improving the accuracy of image stabilization. In the image calibration stage, target calibration parameters obtained using auxiliary lines on the inner wall of a cube calibration device enable more accurate underwater image calibration, effectively eliminating image distortion caused by the complex underwater environment, improving image quality, and making the target underwater image more realistically reflect the underwater scene. Furthermore, in image stitching and fusion, a preset fusion percentage is determined based on the lens's field of view, combined with a preset reference radius and gradient weights for mirror linear stitching and fusion, reducing stitching artifacts and improving the realism of the generated panoramic underwater image, providing users with a more complete and clear underwater panoramic view. For image stabilization, an image stabilization matrix is ​​generated based on the motion data corresponding to the synchronization timestamps, accurately quantifying the amplitude and direction of camera shake. During VR interactive preview, the anti-shake matrix is ​​used to reverse the viewing angle of the panoramic underwater image, effectively counteracting the camera's angular motion and greatly improving image stability. This allows users to smoothly and clearly preview and display fish detection images, enhancing the user experience and underwater detection effect.

[0033] Optionally, in step S15, the step of responding to the user's perspective adjustment operation and performing a fish detection image preview display based on the synchronization timestamp corresponding to the image stabilization matrix and the panoramic underwater image includes: In step S151, in response to the user's interactive operation, a corresponding interactive operation instruction is generated; In this embodiment of the disclosure, the user performs interactive operations in VR interactive preview mode through an input device (such as a controller, head-mounted device, etc.). The interactive operations are monitored in real time and converted into specific electrical signals or data. Then, according to preset rules and protocols, these signals or data are parsed into corresponding interactive operation instructions.

[0034] The interactive operations can include one or more of the following: touch to rotate, touch to slide, touch to zoom.

[0035] For example, when viewing a panoramic image of an underwater school of fish in VR, the user rotates the controller to simulate head rotation. The controller's rotation signal can be detected and converted into a viewpoint rotation command according to preset rules. For instance, if the user rotates the controller clockwise by a certain angle, a corresponding clockwise viewpoint rotation command is generated, instructing the user to adjust the image viewpoint according to that direction and angle.

[0036] In step S152, the image stabilization matrix corresponding to the panoramic underwater image is called according to the synchronization timestamp, and a rendering matrix is ​​generated according to the image stabilization matrix and the interactive operation command. The rendering matrix is ​​a mathematical matrix used for viewpoint transformation and rendering of images, and it contains information such as image rotation, translation, and scaling.

[0037] In this embodiment, based on the synchronization timestamp, the image stabilization matrix corresponding to the current panoramic underwater image is searched in the stored data. The image stabilization matrix records the camera's shaking information when capturing the image. Then, combined with the viewpoint adjustment information in the interactive operation commands, the image stabilization matrix and the viewpoint adjustment information are fused through mathematical operations to generate a rendering matrix. The rendering matrix can both counteract camera shaking and achieve the user's desired viewpoint adjustment.

[0038] For example, if a user wants to turn the camera to the left to view a school of fish, the corresponding stabilization matrix can be found based on the synchronization timestamp. Suppose the stabilization matrix indicates that the camera is shaking to the right, and the user interaction command requests to turn the camera to the left. Through mathematical calculations, the actions that cancel out the rightward shaking and the leftward camera rotation are combined to generate a rendering matrix.

[0039] In step S153, the viewpoint of the corresponding panoramic underwater image is adjusted according to the rendering matrix, and the panoramic underwater image after the viewpoint adjustment is rendered, and a fish school detection image preview is displayed.

[0040] Viewpoint adjustment allows you to change the angle and direction of viewing an image to achieve different visual effects. Rendering processes image data according to certain rules and algorithms to generate the final image that can be displayed on a display device.

[0041] In this embodiment, the panoramic underwater image is mathematically transformed based on the generated rendering matrix. Operations such as matrix multiplication are used to perform image rotation and translation, adjusting the viewing angle. Then, a graphics rendering engine is used to render the adjusted image data, incorporating information such as lighting and texture, to generate a realistic image of fish detection. Finally, the rendered image is displayed on a display device for user preview.

[0042] For example, after adjusting the viewing angle of the panoramic underwater image based on the rendering matrix, the rendering stage begins. The graphics rendering engine considers the lighting effects of the underwater environment, such as the refraction of sunlight through the water surface, and the texture details of the fish. After calculation and processing, a brightly colored and richly detailed image of the fish is generated and displayed on the screen of the VR device, making the user feel as if they are observing the fish underwater.

[0043] The aforementioned technical solution generates commands in response to interactive operations, promptly meeting users' needs for adjusting their perspective and enhancing interactivity and immersion. By invoking the anti-shake matrix and generating a rendering matrix based on the synchronization timestamp, camera shake is effectively counteracted, ensuring image stability and preventing image blurring or misalignment caused by shake, thus improving image quality. Adjusting and rendering the panoramic underwater image based on the rendering matrix allows users to clearly observe fish schools from different angles and obtain more comprehensive information. This improves the accuracy, stability, and interactivity of fish school detection image previews, providing users with a higher quality and more realistic observation experience, and contributing to a deeper study of underwater fish ecology and behavior.

[0044] Optionally, in step S152, generating the rendering matrix based on the image stabilization matrix and the interactive operation command includes: In step S1521, the interactive operation instruction is parsed to determine the operation type and operation amount corresponding to the interactive operation. The operation type includes one or more of the following: screen translation, image scaling, and rotation around the world coordinate axis. Among them, the operation quantity is a numerical value used to measure the degree of execution of an operation type, such as the translation distance, scaling ratio, rotation angle, etc.

[0045] In this embodiment of the disclosure, after receiving an interactive operation command, it is parsed. Key information included in the command is identified through preset command parsing rules and protocols to determine the operation type. Simultaneously, numerical values ​​related to the operation type, i.e., operation quantities, are extracted from the command. Different operation types correspond to different parsing methods; for example, for a screen panning command, the panning direction and distance are parsed; for an image zoom command, the zoom ratio is parsed; and for a rotation command around the world coordinate principal axis, the rotation axis and rotation angle are parsed.

[0046] For example, when viewing a panoramic image of an underwater school of fish in VR, the user presses a specific button on the controller and moves the controller to pan the image. Upon receiving the command, the system interprets the operation type as panning and the amount of movement as a horizontal movement to the right a certain distance. If the user rotates the controller to zoom in or out, the system interprets the operation type as zooming and the amount of movement as doubling the zoom level.

[0047] In step S1522, an adjustment value relative to the initial state corresponding to the operation type is determined based on the operation amount; The initial state refers to the original state of the image before any interactive operation is performed, including parameters such as position, size, and angle. The adjustment value is calculated based on the operation amount and represents the numerical change required relative to the initial state, used to achieve the image changes required by the operation type.

[0048] In this embodiment, adjustment values ​​are calculated based on relevant parameters of the initial state. For image translation, the adjustment value is the sum of the initial position coordinates and the coordinate changes corresponding to the operation amount; for image scaling, the adjustment value is the product of the initial scaling ratio and the scaling ratio of the operation amount; for rotation around the world coordinate principal axis, the adjustment value is the sum of the initial rotation angle and the rotation angle of the operation amount. These adjustment values ​​accurately reflect the degree of image change.

[0049] For example, if the image is initially centered on the screen, and the operation is a pan, moving 100 pixels to the right, then the adjustment value is the initial x-coordinate plus 100. If the initial image scaling is 1, and the operation is a 1.5x zoom, then the adjustment value is 1 x 1.5 = 1.5.

[0050] In step S1523, the adjustment value is converted into a preview matrix corresponding to this interactive operation using a matrix transformation function; The preview matrix records viewpoint adjustment parameters (such as drag direction and rotation angle) generated by the user through interactive operations. The matrix transformation function converts the adjustment values ​​into mathematical functions in matrix form, and performs image transformation operations through matrix operations.

[0051] In this embodiment, matrix transformation functions are used to calculate the adjustment value as an input parameter. Different operation types correspond to different matrix transformation functions; for example, translation operations use translation matrix transformation functions, scaling operations use scaling matrix transformation functions, and rotation operations use rotation matrix transformation functions. These functions generate corresponding transformation matrices, i.e., preview matrices, based on the adjustment values. The preview matrix contains all the information required for image transformation, such as the translation vector, scaling factor, rotation angle, and rotation axis.

[0052] For example, for a screen panning operation, if the adjustment value is to move 100 pixels to the right and 50 pixels up, the translation matrix transformation function will generate a translation matrix, where the translation amount in the x-direction is 100 and the translation amount in the y-direction is 50. This matrix is ​​the preview matrix, used to describe the translation transformation of the image.

[0053] In step S1524, the rendering matrix is ​​determined based on the product of the stabilization matrix and the preview matrix.

[0054] Before rendering the video frame, the stabilization matrix and the preview matrix are multiplied to obtain the rendering matrix. For example, during playback, the corresponding IMU data is matched according to the timestamp of the current video frame, and a stabilization matrix M1 is generated through a pose estimation algorithm; user interaction operations (such as sliding and rotating) are captured in real time, and the operation commands are converted into a preview matrix M2; the rendering matrix M = M1 × M2 (matrix multiplication to ensure the correct logical order) is calculated; the perspective of the current video frame is adjusted and rendered based on matrix M, outputting a picture that combines stabilization effect and interactive response.

[0055] In real-world VR applications, users typically need to adjust their viewing angle through swiping and dragging. Applying only a stabilization matrix can overwrite these interactions, rendering swiping ineffective. By fusing the stabilization matrix with the preview matrix, synergistic compatibility between stabilization and user interaction is achieved.

[0056] The above technical solution accurately parses interactive operation commands, clarifies the operation type and quantity, and calculates adjustment values ​​to ensure that image changes meet user expectations. The adjustment values ​​are converted into a preview matrix, efficiently describing image transformations in matrix form. A rendering matrix is ​​generated, integrating image stabilization and transformation functions. It can respond promptly to user interactions, generating a stable and accurate rendering matrix that effectively counteracts camera shake, allowing users to clearly and stably observe underwater fish from different angles in VR, greatly enhancing the user experience and observation effect.

[0057] Optionally, in step S14, generating a stabilization matrix to counteract the angular motion of the underwater binocular camera based on the angular velocity data and acceleration data corresponding to each synchronization timestamp includes: In step S141, the angular velocity data is transformed from the body coordinate system of the inertial measurement unit to the optical coordinate system of the underwater binocular camera, and the acceleration data is used to assist in correcting the attitude estimation error and correcting the alignment deviation between the body coordinate system and the optical coordinate system. The underwater binocular camera's optical coordinate system is a coordinate system established with the camera's optical center as the origin and the optical axis as a certain coordinate axis, used to describe the geometric relationship of the camera's imaging. Attitude estimation error is the deviation between the estimated value and the true value caused by various factors (such as sensor noise, coordinate system transformation errors, etc.) during the estimation of the camera's attitude.

[0058] In this embodiment, angular velocity data is converted to the optical coordinate system using mathematical transformations based on known transformation relationships between the body coordinate system and the optical coordinate system (such as rotation matrices and Euler angles). Simultaneously, the acceleration data contains gravity information, which can help correct attitude estimation errors. Because accelerometer measurements are affected by gravity, analyzing the components of acceleration data in different coordinate systems can identify and correct alignment deviations between the body coordinate system and the optical coordinate system, improving data accuracy.

[0059] In step S142, the angular velocity corresponding to the angular velocity data in the optical coordinate system is discretized into the rotational change in a continuous time, and the quaternion state is updated according to the rotational change to characterize the three-dimensional rotational attitude of the camera. The accelerometer data is fused with the gravity vector to estimate and correct the attitude drift error generated by the integration of the inertial measurement unit. Among these methods, angular velocity discretization transforms continuous angular velocity data into discrete rotational changes, with quaternion states used to describe the rotational attitude. Attitude drift error is the phenomenon where, during the process of integrating angular velocities to calculate camera attitude, the estimated attitude gradually deviates from the true value over time due to integration errors, sensor noise, and other factors.

[0060] In this embodiment, the angular velocity data in the optical coordinate system is discretized, converting the angular velocity over a continuous time interval into a rotational change. The quaternion state is updated using this rotational change, and the quaternion can accurately describe the camera's rotational attitude in three-dimensional space. Simultaneously, accelerometer data is fused. Since the gravity vector is fixed in the inertial coordinate system, by analyzing the changes in the gravity vector measured by the accelerometer at different times, the attitude drift error generated by the inertial measurement unit integration can be estimated and corrected, thereby improving the accuracy of attitude estimation.

[0061] For example, when monitoring the dynamics of an underwater fish school, the IMU of an underwater binocular camera continuously measures angular velocity. The angular velocity over a period of time is discretized into multiple small rotational values, such as one rotational value every 0.01 seconds. The quaternion is updated using these rotational values ​​to obtain the camera's rotational attitude at different times. Since the integration process introduces attitude drift errors, gravity data measured by the accelerometer can be used for correction. If the component of the gravity vector in a certain direction should remain constant, but the actual measurement shows a change, it indicates attitude drift. Through analysis, the error can be calculated and corrected, making the attitude estimation more accurate.

[0062] In step S143, the acquisition time with the smallest motion amplitude is selected as the reference time, and the relative rotation matrix of each acquisition time relative to the reference time is calculated. The relative rotation matrix is ​​used to describe the image plane rotation change caused by the angular motion of the underwater binocular camera. The reference time is a selected moment during data acquisition used to calculate the changes at other times relative to it. The relative rotation matrix describes the rotation relationship between one coordinate system and another, representing the rotational changes of the image plane caused by the angular motion of the camera at different times.

[0063] In this embodiment, the moment with the smallest motion amplitude is selected as the reference moment because the camera is relatively stable at this time, and the attitude estimation is more accurate. The relative rotation matrix of each acquisition moment relative to the reference moment is calculated. By comparing the camera's rotational attitude at different moments, the relative rotation matrix is ​​obtained using mathematical methods (such as rotation matrix multiplication). The relative rotation matrix can accurately describe the image plane rotation changes caused by the camera's angular motion.

[0064] For example, suppose the camera is almost stationary at a certain moment, and this moment is selected as the reference moment. At each subsequent moment, based on the camera's pose estimation, the rotation relative to the reference moment is calculated. For instance, if the camera rotates a certain angle around the optical axis at a certain moment, the corresponding relative rotation matrix can be calculated. This matrix reflects the effect of this rotation on the image plane, meaning the image plane also rotates around the center by the same angle.

[0065] In step S144, the image stabilization matrix corresponding to each lens in the underwater binocular camera is determined based on the inverse matrix of the relative rotation matrix.

[0066] In this embodiment, the inverse of the relative rotation matrix is ​​applied to image processing to eliminate image rotation caused by camera angular motion, thus canceling out camera angular motion during image display. For each lens of the underwater binocular camera, the inverse matrix is ​​calculated based on the corresponding relative rotation matrix to obtain its respective image stabilization matrix.

[0067] For example, when filming an underwater school of fish, the camera undergoes angular motion, causing the image to rotate. If the relative rotation matrix corresponding to a certain lens rotates the image 10 degrees clockwise, then its inverse matrix will rotate the image 10 degrees counterclockwise. In image processing, the image stabilization matrix generated by applying this inverse matrix can counteract the camera's angular motion, allowing the observer to see a stable image of the fish school.

[0068] The above technical solution achieves data coordinate system transformation and deviation correction to ensure data accuracy; it discretizes angular velocity and updates quaternions, fuses acceleration data to correct attitude drift, and improves attitude estimation accuracy; it selects a reference moment to calculate the relative rotation matrix, clearly describing the impact of camera angular motion on the image; and it determines the stabilization matrix based on the inverse matrix to effectively counteract angular motion. In this way, an accurate stabilization matrix can be generated, reducing image shake and making the captured underwater fish images clear and stable.

[0069] Optionally, in step S13, the step of performing mirror linear stitching and fusion of the target underwater images of the two lenses at the same acquisition time according to a preset fusion percentage and a preset reference radius, through gradual weighting, to obtain a panoramic underwater image at that acquisition time, includes: The pixels in the target underwater image corresponding to either of the two lenses that are located between the preset fusion percentage and the preset reference radius are mirrored and linearly stitched together with the pixels in the target underwater image corresponding to the other lens in the overlapping area, using a gradient weighting method, to obtain the panoramic underwater image at the time of acquisition. The weight of the pixel corresponding to any one of the lenses gradually decreases from the side of the overlapping region closer to the lens to the side farther away from the lens, while the weight of the pixel corresponding to the other lens gradually increases as the weight of the pixel in any one of the lenses gradually decreases.

[0070] In this embodiment, since the field of view (FOV) of the lens used is approximately 200° (greater than 180°), there is a fixed overlapping area between the left and right lens images. This area is the core object of the fusion processing. In actual operation, the fusion percentage can be defined first based on the field of view. The fusion percentage can be determined based on the empirical value of the lens field of view. For example, the empirical value of the fusion percentage corresponding to a field of view of 200° is 0.9754, and the reference radius is set to 1.0.

[0071] When the fusion algorithm is executed, the fusion region is divided based on the radius range. For example, pixels in the left shot image with a radius of 0.9754 to 1.0 are mirrored and linearly fused with the corresponding overlapping area pixels in the right shot image. Linear fusion is achieved through gradual weight allocation, that is, the weight of the left shot pixels gradually decreases from the inside to the outside of the fusion region, while the weight of the right shot pixels gradually increases synchronously, ultimately making the pixel transition at the stitching point smooth and eliminating visual discontinuities.

[0072] Optionally, the target calibration parameters are obtained by calibration in the following manner: The geometric center of the lens is determined by drawing horizontal and vertical crosshairs on the original imaging image captured by the lens. The raw image is the unprocessed image directly captured by the lens, containing the original information of the scene as captured. Crosshairs are two perpendicular lines drawn on the image to aid in positioning and measurement. The geometric center is the symmetrical center point in the image; for regularly shaped image areas, it reflects the symmetrical characteristics of the lens's imaging.

[0073] In this embodiment, horizontal and vertical crosshair auxiliary lines are drawn on the original image captured by the lens. These two lines intersect each other perpendicularly. Since the image is ideally symmetrical, the intersection of the crosshair auxiliary lines is the geometric center of the image. The geometric center of the lens imaging can be accurately located, and the calibration operation is performed based on the geometric center for coordinate adjustment and parameter calculation.

[0074] For example, when photographing a regular square object, the original image captured by the lens may have some deviation. By drawing horizontal and vertical crosshairs on the image, and assuming the square image is basically symmetrical in the horizontal and vertical directions, the intersection of these crosshairs is the geometric center of the square image. This center point reflects the position of the center of symmetry on the image plane when the lens is forming the image.

[0075] Using the reference calibration center formed by the intersection of auxiliary lines on the inner wall of the cube calibration device as a reference, adjust the coordinates of the original image acquired by the lens until the position of the geometric center and the reference calibration center meets the preset position requirements, and use the translation amount of the coordinates as the corresponding center offset parameter of the lens. The cube calibration device is a specific device used for lens calibration. Its inner wall has auxiliary lines that intersect to form a reference calibration center, providing a standard reference for lens calibration. The reference calibration center is the point formed by the intersection of the auxiliary lines on the inner wall of the cube calibration device, serving as the benchmark for adjusting the lens image coordinates and ensuring that the lens image is aligned with the standard position. The center offset parameter is the numerical value of the coordinate translation when the position of the lens's geometric center and the reference calibration center does not meet the preset requirements; it describes the offset of the lens center relative to the standard position.

[0076] In this embodiment, the coordinates of the original image acquired by the lens are adjusted based on the reference calibration center formed by the intersection of auxiliary lines on the inner wall of the cube calibration device. By continuously changing the coordinate position of the image, the geometric center determined by the lens is gradually moved closer to the reference calibration center until the positions of the two meet the preset position requirements. During this process, the translation amount of the coordinates is recorded. The translation amount reflects the degree of offset of the lens center relative to the reference calibration center and is used as the center offset parameter of the corresponding lens for correcting the lens imaging position.

[0077] For example, within a cube-shaped calibration device, the intersecting auxiliary lines on its inner walls form a clearly defined reference calibration center. The geometric center of the image captured by the lens deviates slightly from this reference calibration center; suppose it's offset by 5 pixels horizontally and 3 pixels vertically. By adjusting the image coordinates to make the geometric center coincide with the reference calibration center, the image is translated by -5 pixels horizontally and -3 pixels vertically. These two translation amounts are the center offset parameters.

[0078] For example, the two center points are aligned by translation adjustment: using the real center of the scene formed by the intersection of the auxiliary lines on the inner wall of the cube as a reference, the x-axis and y-axis translation of the original image are gradually adjusted until the intersection of the crosshair auxiliary lines of the original image completely coincides with the real center of the scene. The coordinates corresponding to the x and y translation adjustments at this time are the real center (x, y) coordinates of the shot.

[0079] Given the center offset parameter, the original image is scaled to determine the effective radius of the lens's field of view. Scaling refers to the operation of enlarging or reducing the size of an image according to a certain ratio to change the image size and adapt to different calibration requirements and display scenarios. The effective field of view radius is the radius length corresponding to the area from the lens's geometric center to the image edge that can be effectively imaged, reflecting the effective range of the lens's imaging.

[0080] In this embodiment, the original image is scaled. By gradually changing the image size, the imaging performance at different sizes is observed. When the image size is adjusted to a suitable level, the radius corresponding to the area from the lens's geometric center to the image edge that can be effectively imaged in sharp focus is measured, using the lens's geometric center as the center. This radius determines the effective field of view radius of the lens. The effective field of view radius determines the range that the lens can effectively cover in actual imaging.

[0081] For example, since there is an offset between the true center of the lens and the center of the original image, in order to ensure the effective overlap area matching during subsequent binocular image stitching, it is necessary to define the effective field of view range of the lens, i.e., the "effective radius". This can be based on the wide field of view characteristic of binocular lenses—the field of view (FOV) of the selected lens is greater than 180°, and the actual field of view is about 200°. Therefore, there is a natural overlap between the right side of the left lens image and the left side of the right lens image, and this overlapping area corresponds to the same physical scene of the cube.

[0082] During calibration, after completing the first step of true center calibration, the imaging of the overlapping area is observed, and the overlapping area of ​​one lens (left or right lens) is selected as the reference. Then, the original image of the other lens is scaled and adjusted so that the size of the imaging content in the overlapping area of ​​the two lenses is exactly the same. The range parameter calculated by the image scaling ratio and the original image size is the effective image radius of each of the left and right lenses.

[0083] For example, when an image is scaled down, the edges begin to blur or become distorted as the image shrinks to a certain size, while the image from the geometric center to a certain point remains clear. Measuring the distance from this point to the geometric center, let's say 100 pixels, gives us the effective radius of the lens's field of view.

[0084] Given a fixed effective field of view radius, the original images captured by the two lenses are rendered onto a spherical circle model for Euler angle calibration to obtain the Euler angle parameters of the lenses. The target calibration parameters include the center offset parameter, the effective field of view radius, and the Euler angle parameter.

[0085] The spherical circle model is a mathematical model used to describe image rendering in 3D space. It renders the image onto a sphere to simulate human visual perception and present the scene more realistically. Euler angle calibration determines the rotation angles of an object around three coordinate axes in 3D space (Euler angles). It describes the object's spatial attitude and orientation, and is used to accurately determine the spatial position and angle of the lens. Euler angle parameters are the numerical values ​​of the angles by which the lens rotates around three coordinate axes in 3D space, used to describe the lens's spatial attitude and orientation.

[0086] In this embodiment, the original images captured by the two lenses are rendered into a spherical circular model. Within the spherical model, Euler angles are used to describe the lens's attitude in three-dimensional space by analyzing the relative position and orientation of the images. The images from the two lenses are matched and adjusted within the spherical model, and the rotation angle of each lens around three coordinate axes (typically roll, pitch, and yaw axes) is calculated. Euler angle parameters accurately describe the lens's attitude in space.

[0087] For example, after rendering the images onto a spherical model, differences were found in the positions and angles of the two images on the sphere. Through calculation and analysis, it was determined that one lens rotated 10 degrees around the roll axis, 5 degrees around the pitch axis, and -3 degrees around the yaw axis; these angles are the Euler angle parameters of that lens. The other lens also has corresponding Euler angle parameters. These parameters can be used to adjust the lens attitude, ensuring that the two images are correctly matched on the sphere.

[0088] The target calibration parameters obtained from the above technical solution, including center offset parameters, effective field of view radius, and Euler angle parameters, provide a comprehensive and precise calibration for lens imaging. The center offset parameter corrects the geometric center position of the lens imaging, making the image more accurate in coordinates; the effective field of view radius determines the effective imaging range of the lens, avoiding interference from invalid information; and the Euler angle parameter accurately describes the lens's attitude in three-dimensional space, ensuring the synergy and consistency of images acquired by the binocular lenses. This improves the quality and accuracy of lens imaging, enabling the acquired images to more realistically and accurately reflect the actual scene, providing reliable basic data for subsequent image processing, 3D reconstruction, and other applications, and enhancing the performance and reliability of the entire system.

[0089] Optionally, in step S12, calibrating the underwater image using pre-calibrated target calibration parameters to obtain the target underwater image at the corresponding acquisition time includes: In step S121, the underwater images corresponding to the two lenses are translated according to the center offset parameter so that the geometric center of the lens is aligned with the center of the preset canvas. In this embodiment, a pre-calibrated center offset parameter is used to define the offset between the lens's geometric center and the center of a preset canvas in both the horizontal and vertical directions. For underwater images corresponding to two lenses, a translation operation is performed according to this offset. In the horizontal direction, if the center offset parameter indicates that the image center has shifted to the right by a certain number of pixels, the image is moved to the left by the corresponding number of pixels; the same applies to the vertical direction. This translation aligns the lens's geometric center with the center of the preset canvas, eliminating image position deviations caused by lens mounting or imaging characteristics. This provides an accurate positional basis for subsequent image processing, ensuring that the image is correctly displayed within the preset frame.

[0090] For example, when shooting underwater scenes, due to lens mounting issues, the center of the image captured by one lens is offset 10 pixels horizontally to the right and 5 pixels vertically downwards relative to the center of the preset canvas. Based on the center offset parameters, a translation operation is performed on the image, moving it 10 pixels to the left and 5 pixels upwards. This aligns the geometric center of the image with the center of the preset canvas, ensuring the image is in the correct position and preventing positional deviations from affecting the overall effect.

[0091] In step S122, based on the effective field of view alarm, the translated underwater image is scaled, and the invalid image that exceeds the preset canvas is cropped out to obtain the effective area image; In this embodiment, the translated underwater image is scaled according to the effective radius of the field of view. If the image size corresponding to the effective radius of the field of view is smaller than the preset canvas size, the image is enlarged; if it is larger than the preset canvas size, the image is reduced so that the effective portion of the image can fit the preset canvas. Then, it is checked whether the image exceeds the preset canvas range, and the excess portion is cropped out. Because the excess portion cannot be effectively displayed within the preset canvas and may contain interference information, the cropped effective area image meets the size requirements of the preset canvas while ensuring image quality.

[0092] For example, if the effective field of view radius of the translated underwater image is large, exceeding the size of the preset canvas, the image is scaled down proportionally based on the effective field of view radius to make it roughly fit the preset canvas. However, after scaling down, it is found that parts of the four corners of the image still extend beyond the preset canvas; these excess, invalid image portions are cropped out. The final effective area image is exactly within the preset canvas, retaining the image's effective information and clearly displaying the main content of the underwater scene.

[0093] In step S123, the effective area image is mapped onto the 3D semicircular model textures corresponding to the two lenses, and the attitude of the 3D semicircular model is adjusted according to the Euler angle parameter to obtain the target underwater image at the corresponding acquisition time.

[0094] Among them, the 3D semi-circular model texture is the texture information used to display images on the surface of the 3D semi-circular model. It maps the image onto the model surface, so that the model presents the corresponding visual effect.

[0095] In this embodiment, the effective area image is mapped onto the textures of the 3D semicircular models corresponding to the two lenses, allowing the image to fit the surface of the 3D semicircular models. Then, the pose of the 3D semicircular models is adjusted according to the Euler angle parameters. The Euler angle parameters define the rotation angle of the lens in three-dimensional space. By applying these angles to the 3D semicircular models, the rotation angle of the models is made consistent with the actual pose of the lenses. Thus, after mapping and pose adjustment, the image presented by the 3D semicircular models can accurately reflect the actual underwater scene at the corresponding acquisition time, obtaining the target underwater image.

[0096] In this embodiment, to eliminate lens center offset and uncertainty in the effective image range based on calibration parameters, the original images of the left and right lenses undergo the same operational process. First, based on the true center (x, y) coordinates of the left / right lenses, the corresponding original images are translated to ensure that the true center of the lenses is perfectly aligned with the center of the preset canvas, thus ensuring that the imaging reference of the two lenses remains consistent. Subsequently, based on the effective radius parameters of the calibrated left / right lenses, the translated images are scaled and adjusted. Since the effective radius defines the effective range of lens imaging, invalid image content exceeding the canvas after scaling is naturally discarded, retaining only the effective imaging area.

[0097] Furthermore, even after translation and scaling, images still exhibit geometric distortion due to lens pose deviations. This distortion is corrected through 3D rendering using calibrated pose parameters, restoring the image to a spatial pose consistent with the real scene. For example, a semi-circular 3D model is first constructed as the rendering medium to simulate the imaging projection characteristics of a binocular panoramic lens, matching the lens's wide field-of-view imaging effect. Then, the left and right camera images are mapped onto their respective 3D semi-circular model textures. The calibrated pose parameters (yaw, pitch, and roll) are then used to adjust the model's pose: adjusting the yaw angle corrects the horizontal offset, adjusting the pitch angle corrects the vertical offset, and adjusting the roll angle corrects the rotational offset. Ultimately, the rendered left and right camera images perfectly match the spatial geometry of the real scene, resulting in a corrected standard image.

[0098] The above technical solution translates the image based on the center offset parameter, eliminating image position deviations caused by lens installation or imaging characteristics, ensuring accurate image display within the preset canvas. Secondly, by scaling and cropping the image according to the effective field of view radius, the resulting effective area image meets processing size requirements while maintaining image quality and removing interference from invalid information. Finally, the effective area image is mapped onto a 3D semi-circular model and its attitude is adjusted. Combined with Euler angle parameters, the image presented by the model accurately reflects the attitude and content of the actual underwater scene, improving image realism and accuracy, and enhancing the performance and reliability of the entire underwater image processing system.

[0099] This disclosure also provides a fish detection and preview device to improve the image stabilization of underwater binocular cameras. See [link to relevant documentation]. Figure 2 As shown, the device includes: The acquisition module 210 is configured to acquire angular velocity data and acceleration data set in the inertial measurement unit of the underwater binocular camera when the underwater binocular camera acquires an underwater image, and to add a synchronization timestamp to the underwater image, the angular velocity data and the acceleration data at the same acquisition time. The calibration module 220 is configured to calibrate the underwater image using pre-calibrated target calibration parameters to obtain the target underwater image at the corresponding acquisition time. The target calibration parameters are obtained by calibrating multiple core parameters sequentially based on multiple sets of auxiliary lines drawn on the inner wall of the cube calibration device as geometric references. The stitching and fusion module 230 is configured to perform mirror linear stitching and fusion of the target underwater images of the two lenses at the same acquisition time according to a preset fusion percentage and a preset reference radius, through gradual weighting, to obtain a panoramic underwater image at that acquisition time. The preset fusion percentage is determined according to the field of view of the lens. The generation module 240 is configured to generate an anti-shake matrix to counteract the angular motion of the underwater binocular camera based on the angular velocity data and acceleration data corresponding to each synchronization timestamp. The anti-shake matrix is ​​a geometric transformation parameter used to quantify the amplitude and direction of the shaking of the underwater binocular camera when capturing the underwater image. The preview module 250 is configured to, in VR interactive preview mode, respond to the user's perspective adjustment operation and display a fish detection image preview based on the synchronization timestamp corresponding to the anti-shake matrix and the panoramic underwater image. The anti-shake matrix is ​​used to reverse the perspective of the panoramic underwater image.

[0100] Optionally, the preview module 250 is configured to: In response to user interaction, generate corresponding interaction commands; Based on the synchronization timestamp, the image stabilization matrix corresponding to the panoramic underwater image is invoked, and a rendering matrix is ​​generated based on the image stabilization matrix and the interactive operation instructions. Based on the rendering matrix, the viewing angle of the corresponding panoramic underwater image is adjusted, and the panoramic underwater image after the viewing angle adjustment is rendered, and a fish school detection image preview is displayed.

[0101] Optionally, the preview module 250 is configured to: The interactive operation instructions are parsed to determine the operation type and operation amount corresponding to the interactive operation. The operation type includes one or more of the following: screen panning, image scaling, and rotation around the world coordinate principal axis. Based on the operation amount, determine an adjustment value relative to the initial state corresponding to the operation type; The adjustment value is converted into a preview matrix corresponding to this interactive operation using a matrix transformation function; The rendering matrix is ​​determined by multiplying the stabilization matrix and the preview matrix.

[0102] Optionally, the generation module 240 is configured to: The angular velocity data is transformed from the body coordinate system of the inertial measurement unit to the optical coordinate system of the underwater binocular camera, and the acceleration data is used to assist in correcting the attitude estimation error and correcting the alignment deviation between the body coordinate system and the optical coordinate system. The angular velocity data in the optical coordinate system is discretized into rotational changes over a continuous time period. The quaternion state is updated based on the rotational changes to characterize the three-dimensional rotational attitude of the camera. The accelerometer data is fused with the gravity vector to estimate and correct the attitude drift error generated by the integration of the inertial measurement unit. The acquisition time with the smallest motion amplitude is selected as the reference time. The relative rotation matrix of each acquisition time relative to the reference time is calculated. The relative rotation matrix is ​​used to describe the image plane rotation change caused by the angular motion of the underwater binocular camera. The image stabilization matrix corresponding to each lens in the underwater binocular camera is determined based on the inverse of the relative rotation matrix.

[0103] Optionally, the splicing and fusion module 230 is configured as follows: The pixels in the target underwater image corresponding to either of the two lenses that are located between the preset fusion percentage and the preset reference radius are mirrored and linearly stitched together with the pixels in the target underwater image corresponding to the other lens in the overlapping area, using a gradient weighting method, to obtain the panoramic underwater image at the time of acquisition. The weight of the pixel corresponding to any one of the lenses gradually decreases from the side of the overlapping region closer to the lens to the side farther away from the lens, while the weight of the pixel corresponding to the other lens gradually increases as the weight of the pixel in any one of the lenses gradually decreases.

[0104] Optionally, the device includes a calibration module configured to calibrate the target calibration parameters in the following manner: The geometric center of the lens is determined by drawing horizontal and vertical crosshairs on the original imaging image captured by the lens. Using the reference calibration center formed by the intersection of auxiliary lines on the inner wall of the cube calibration device as a reference, adjust the coordinates of the original image acquired by the lens until the position of the geometric center and the reference calibration center meets the preset position requirements, and use the translation amount of the coordinates as the corresponding center offset parameter of the lens. Given the center offset parameter, the original image is scaled to determine the effective radius of the lens's field of view. Given a fixed effective field of view radius, the original images captured by the two lenses are rendered onto a spherical circle model for Euler angle calibration to obtain the Euler angle parameters of the lenses. The target calibration parameters include the center offset parameter, the effective field of view radius, and the Euler angle parameter.

[0105] Optionally, the calibration module 220 is configured to: Based on the center offset parameter, the underwater images corresponding to the two lenses are translated respectively so that the geometric center of the lens is aligned with the center of the preset canvas; Based on the effective field of view alarm, the translated underwater image is scaled, and invalid images that exceed the preset canvas are cropped out to obtain the effective area image; The effective area image is mapped onto the 3D semicircular model textures corresponding to the two lenses, and the attitude of the 3D semicircular model is adjusted according to the Euler angle parameters to obtain the target underwater image at the corresponding acquisition time.

[0106] This disclosure also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the methods described in the foregoing embodiments.

[0107] This disclosure also provides an electronic device, including: A memory on which computer programs are stored; A processor for executing the computer program in the memory to implement the steps of any of the methods described in the foregoing embodiments.

[0108] Figure 3 The fish detection and preview device 100 for improving the image stabilization of an underwater binocular camera, as shown, includes a processor 1001 and a memory 1003. The processor 1001 and the memory 1003 are connected, for example, via a bus 1002. Optionally, the fish detection and preview device 100 may further include a communication component, which can be used for data interaction between the device 100 and other devices, such as sending or receiving data. It should be noted that in actual operation, the communication component is not limited to one, and the structure of this fish detection and preview device 100 for improving the image stabilization of an underwater binocular camera does not constitute a limitation on the embodiments of this application.

[0109] Processor 1001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 1001 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.

[0110] Bus 1002 may include a pathway for transmitting information between the aforementioned components. Bus 1002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Bus 1002 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 3 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0111] The memory 1003 may be ROM (Read Only Memory) or other types of static storage devices capable of storing static information and instructions, RAM (Random Access Memory) or other types of dynamic storage devices capable of storing information and instructions, or EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium capable of carrying or storing program code and capable of being read by a computer, without limitation herein.

[0112] The memory 1003 is used to store program code for executing embodiments of this disclosure, and its execution is controlled by the processor 1001. The processor 1001 is used to execute the program code stored in the memory 1003 to implement the steps shown in the aforementioned embodiments of the fish school detection preview method for improving the image stabilization of underwater binocular cameras.

[0113] This disclosure also provides a computer-readable storage medium storing program code. When the program code is executed by a processor, it can implement the steps and corresponding content of the aforementioned embodiment of the fish school detection and preview method for improving the image stabilization of an underwater binocular camera.

[0114] The preferred embodiments of the present disclosure have been described in detail above with reference to the accompanying drawings. However, the present disclosure is not limited to the specific details of the above embodiments. Within the scope of the technical concept of the present disclosure, various changes, modifications, substitutions and variations can be made to these embodiments, and all such changes, modifications, substitutions and variations fall within the protection scope of the present disclosure.

[0115] It should also be noted that the various specific technical features described in the above embodiments can be combined in any suitable manner without contradiction, and such combinations should also be considered as part of this disclosure. To avoid unnecessary repetition, this disclosure will not further describe the various possible combinations. The technical scope of this application is not limited to the contents of the specification, but must be determined according to the scope of the claims.

Claims

1. A method for fish school detection and preview to improve the image stabilization of an underwater binocular camera, characterized in that, The method includes: When an underwater binocular camera acquires an underwater image, it sets the angular velocity and acceleration data in the inertial measurement unit of the underwater binocular camera, and adds a synchronization timestamp to the underwater image, the angular velocity data, and the acceleration data at the same acquisition time. The underwater image is calibrated using pre-calibrated target calibration parameters to obtain the target underwater image at the corresponding acquisition time. The target calibration parameters are obtained by calibrating multiple core parameters sequentially based on multiple sets of auxiliary lines drawn on the inner wall of the cube calibration device as geometric references. Based on a preset fusion percentage and a preset reference radius, the target underwater images of the two lenses at the same acquisition time are mirrored and linearly stitched together using a gradient weighting method to obtain a panoramic underwater image at that acquisition time. The preset fusion percentage is determined based on the field of view of the lens. Based on the angular velocity data and acceleration data corresponding to each synchronization timestamp, an anti-shake matrix is ​​generated to counteract the angular motion of the underwater binocular camera. The anti-shake matrix is ​​a geometric transformation parameter used to quantify the amplitude and direction of the shaking of the underwater binocular camera when capturing the underwater image. In VR interactive preview mode, in response to the user's perspective adjustment operation, a fish school detection image preview is displayed according to the synchronization timestamp corresponding to the anti-shake matrix and the panoramic underwater image. The anti-shake matrix is ​​used to reverse the perspective of the panoramic underwater image.

2. The method according to claim 1, characterized in that, The process of responding to the user's perspective adjustment operation by performing a fish detection image preview display based on the synchronization timestamp corresponding to the image stabilization matrix and the panoramic underwater image includes: In response to user interaction, generate corresponding interaction commands; Based on the synchronization timestamp, the image stabilization matrix corresponding to the panoramic underwater image is invoked, and a rendering matrix is ​​generated based on the image stabilization matrix and the interactive operation instructions. Based on the rendering matrix, the viewing angle of the corresponding panoramic underwater image is adjusted, and the panoramic underwater image after the viewing angle adjustment is rendered, and a fish school detection image preview is displayed.

3. The method according to claim 2, characterized in that, The step of generating a rendering matrix based on the stabilization matrix and the interactive operation instructions includes: The interactive operation instructions are parsed to determine the operation type and operation amount corresponding to the interactive operation. The operation type includes one or more of the following: screen panning, image scaling, and rotation around the world coordinate principal axis. Based on the operation amount, determine an adjustment value relative to the initial state corresponding to the operation type; The adjustment value is converted into a preview matrix corresponding to this interactive operation using a matrix transformation function; The rendering matrix is ​​determined by multiplying the stabilization matrix and the preview matrix.

4. The method according to claim 1, characterized in that, The step of generating a stabilization matrix to counteract the angular motion of the underwater binocular camera based on the angular velocity data and acceleration data corresponding to each synchronization timestamp includes: The angular velocity data is transformed from the body coordinate system of the inertial measurement unit to the optical coordinate system of the underwater binocular camera, and the acceleration data is used to assist in correcting the attitude estimation error and correcting the alignment deviation between the body coordinate system and the optical coordinate system. The angular velocity data in the optical coordinate system is discretized into rotational changes over a continuous time period. The quaternion state is updated based on the rotational changes to characterize the three-dimensional rotational attitude of the camera. The accelerometer data is fused with the gravity vector to estimate and correct the attitude drift error generated by the integration of the inertial measurement unit. The acquisition time with the smallest motion amplitude is selected as the reference time. The relative rotation matrix of each acquisition time relative to the reference time is calculated. The relative rotation matrix is ​​used to describe the image plane rotation change caused by the angular motion of the underwater binocular camera. The image stabilization matrix corresponding to each lens in the underwater binocular camera is determined based on the inverse of the relative rotation matrix.

5. The method according to claim 1, characterized in that, The step of performing mirror-linear stitching and fusion of the target underwater images from the two lenses at the same acquisition time, based on a preset fusion percentage and a preset reference radius, using a gradual weighting method, to obtain a panoramic underwater image at that acquisition time includes: The pixels in the target underwater image corresponding to either of the two lenses that are located between the preset fusion percentage and the preset reference radius are mirrored and linearly stitched together with the pixels in the target underwater image corresponding to the other lens in the overlapping area, using a gradient weighting method, to obtain the panoramic underwater image at the time of acquisition. The weight of the pixel corresponding to any one of the lenses gradually decreases from the side of the overlapping region closer to the lens to the side farther away from the lens, while the weight of the pixel corresponding to the other lens gradually increases as the weight of the pixel in any one of the lenses gradually decreases.

6. The method according to any one of claims 1-5, characterized in that, The target calibration parameters were obtained by calibration in the following manner: The geometric center of the lens is determined by drawing horizontal and vertical crosshairs on the original imaging image captured by the lens. Using the reference calibration center formed by the intersection of auxiliary lines on the inner wall of the cube calibration device as a reference, adjust the coordinates of the original image acquired by the lens until the position of the geometric center and the reference calibration center meets the preset position requirements, and use the translation amount of the coordinates as the corresponding center offset parameter of the lens. Given the center offset parameter, the original image is scaled to determine the effective radius of the lens's field of view. Given a fixed effective field of view radius, the original images captured by the two lenses are rendered onto a spherical circle model for Euler angle calibration to obtain the Euler angle parameters of the lenses. The target calibration parameters include the center offset parameter, the effective field of view radius, and the Euler angle parameter.

7. The method according to claim 6, characterized in that, The process of calibrating the underwater image using pre-defined target calibration parameters to obtain the target underwater image at the corresponding acquisition time includes: Based on the center offset parameter, the underwater images corresponding to the two lenses are translated respectively so that the geometric center of the lens is aligned with the center of the preset canvas; Based on the effective field of view alarm, the translated underwater image is scaled, and invalid images that exceed the preset canvas are cropped out to obtain the effective area image; The effective area image is mapped onto the 3D semicircular model textures corresponding to the two lenses, and the attitude of the 3D semicircular model is adjusted according to the Euler angle parameters to obtain the target underwater image at the corresponding acquisition time.

8. A fish school detection and preview device for improving the image stabilization of an underwater binocular camera, characterized in that, The device includes: The acquisition module is configured to acquire angular velocity and acceleration data set in the inertial measurement unit of the underwater binocular camera when the underwater binocular camera acquires an underwater image, and to add a synchronization timestamp to the underwater image, the angular velocity data and the acceleration data at the same acquisition time. The calibration module is configured to calibrate the underwater image using pre-calibrated target calibration parameters to obtain the target underwater image at the corresponding acquisition time. The target calibration parameters are obtained by calibrating multiple core parameters sequentially based on multiple sets of auxiliary lines drawn on the inner wall of the cube calibration device as geometric references. The stitching and fusion module is configured to perform mirror linear stitching and fusion of the target underwater images of the two lenses at the same acquisition time according to a preset fusion percentage and a preset reference radius, through gradual weighting, to obtain a panoramic underwater image at that acquisition time. The preset fusion percentage is determined according to the field of view of the lens. The generation module is configured to generate an anti-shake matrix to counteract the angular motion of the underwater binocular camera based on the angular velocity data and acceleration data corresponding to each synchronization timestamp. The anti-shake matrix is ​​a geometric transformation parameter used to quantify the amplitude and direction of the shaking of the underwater binocular camera when capturing the underwater image. The preview module is configured to, in VR interactive preview mode, respond to the user's perspective adjustment operation and display a fish detection image preview based on the synchronization timestamp corresponding to the anti-shake matrix and the panoramic underwater image. The anti-shake matrix is ​​used to reverse the perspective of the panoramic underwater image.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the method described in any one of claims 1-7.

10. A binocular camera, characterized in that, include: A memory on which computer programs are stored; A processor for executing the computer program in the memory to implement the steps of the method according to any one of claims 1-7.