Biomass estimation using monocular underwater cameras

EP4670120A1Pending Publication Date: 2025-12-31TIDALX AI INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
EP2024717538
Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-03-30
Filing Date
2024-03-19
Publication Date
2025-12-31

AI Technical Summary

Technical Problem

Current methods for biomass estimation in aquaculture are time-intensive and potentially harmful, relying on manual fish removal and weighing, and often only measure a small portion of the fish population, leaving true characteristics unknown.

Method used

A method using a monocular underwater camera to estimate fish biomass by processing multiple frames, identifying key points, generating 3D models, and determining biomass estimates, which reduces reliance on stereo camera hardware and allows for real-time estimation without exact calibration settings.

Benefits of technology

This approach enables efficient and cost-effective biomass estimation in aquatic environments, improving the accuracy and availability of monitoring systems while reducing the need for complex stereo camera setups.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024020513_03102024_PF_FP_ABST
    Figure US2024020513_03102024_PF_FP_ABST
Patent Text Reader

Abstract

Methods, systems, and apparatus, including computer programs encoded on computer-storage media, for estimating fish biomass using a monocular camera including obtaining multiple frames captured by the monocular camera over a period of time, where at least one frame of the multiple frames includes a fish, identifying a first set of key points in a two-dimensional space for the fish in the multiple frames, receiving motion data for the monocular camera over the period of time, generating, from the first set of key points in the two-dimensional space and the motion data for the monocular camera, a fish model including a second set of key points in a three-dimensional space for the fish, generating a biomass estimate of the fish based on the second set of key points, and determining an action based on one or more biomass estimates including the biomass estimate of the fish.
Need to check novelty before this filing date? Find Prior Art

Description

BIOMASS ESTIMATION USING MONOCULAR UNDERWATER CAMERASCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit under 35 U.S.C. § 119(e) of the filing date of U.S. Provisional Patent Application No. 63 / 493,196, entitled “BIOMASS ESTIMATION USING MONOCULAR UNDERWATER CAMERAS,” which was filed on March 30, 2023, and which is incorporated here by reference.FIELD

[0002] This specification generally relates to cameras that are used for biomass estimation, and particularly to underwater cameras that are used for aquatic life.BACKGROUND

[0003] A population of fish may include fish of varying sizes, shapes, and health conditions. In the aquaculture context, prior to harvesting, a worker may remove some fish from the fish pen and weigh them. The manual process of removing the fish from the fish pen and weighing them is both time intensive and potentially harmful to the fish. In addition, because only a small portion of a fish population may be effectively measured in this way, the true characteristics of the population remain unknown.SUMMARY

[0004] In general, innovative aspects of the subject matter described in this specification relate to methods for biomass estimation, particularly that are used for aquatic life.

[0005] One innovative aspect of the subject matter described in this specification is embodied in a method for estimating fish biomass using a monocular camera including obtaining multiple frames captured by the monocular camera over a period of time, where at least one frame of the multiple frames includes a fish, identifying a first set of key points in a two-dimensional space for the fish in the multiple frames, receivingmotion data for the monocular camera over the period of time, generating, from the first set of key points in the two-dimensional space and the motion data for the monocular camera, a fish model including a second set of key points in a three-dimensional space for the fish, generating a biomass estimate of the fish based on the second set of key points, and determining an action based on one or more biomass estimates including the biomass estimate of the fish.

[0006] Other implementations of this and other aspects include corresponding systems, apparatus, and computer programs, configured to perform the actions of the methods, encoded on computer storage devices. A system of one or more computers can be so configured by virtue of software, firmware, hardware, or a combination of them installed on the system that in operation causes the system to perform the actions. One or more computer programs can be so configured by virtue of having instructions that, when executed by data processing apparatus, cause the apparatus to perform the actions.

[0007] An advantage of the methods, systems, and apparatuses described herein includes reducing a reliance on stereo camera hardware for performing biomass estimations in aquatic environments. Monocular cameras can, in general, be less expensive to produce, calibrate, and maintain than stereo camera equivalents. The ease and relative low-cost of deployment of the monocular cameras can increase availability of deployed hardware for monitoring of aquatic life, e.g., biomass estimations. Additionally, monocular cameras can be more efficient by requiring less image data to be transferred to processing elements when compared to stereo camera setups. At times, a commercial aquatic vision system can include stereo camera(s) that may be out of calibration or have unknown calibration settings, making them difficult to use for stereovision-based biomass estimates. The technology of this specification allows for real-time estimation of calibration settings for monocular and stereo cameras such that biomass estimates can be performed without having access to exact calibration settings and / or without knowledge of mis-calibration of stereo cameras.

[0008] The details of one or more embodiments of the invention are set forth in the accompanying drawings and the description below. Other features and advantages of the invention will become apparent from the description, the drawings, and the claims.BRIEF DESCRIPTION OF THE DRAWINGS

[0009] FIG. 1 is a diagram showing an example of a system that is used for monocular underwater camera biomass estimation.

[0010] FIGS. 2A-2B depict various aspects of an example process for monocular underwater camera biomass estimation.

[0011] FIGS. 3A-3B depict various aspects of an example process for monocular underwater camera biomass estimation.

[0012] FIGS. 4A-4B depict various aspects of an example process for monocular underwater camera biomass estimation.

[0013] FIG. 5 is a flow diagram showing an example of a system for monocular underwater camera biomass estimation.

[0014] FIG. 6 is a diagram showing an example of a truss network.

[0015] FIG. 7 is a diagram illustrating an example of a computing system used for monocular underwater camera biomass estimation.

[0016] Like reference numbers and designations in the various drawings indicate like elements.DETAILED DESCRIPTION

[0017] The technology of this specification describes methods for biomass estimation, particularly that are used for aquatic livestock. More particularly, the technology includes methods for estimating biomass of fish using monocular camera image data. Individual fish can be photographed using an underwater camera, e.g., a monocular camera generating two-dimensional (2D) imaging data. An image from the underwater camera can be processed using computer vision and machine learning-based techniques to identify fish within the imaging data. A Bayesian framework can be implemented including multiple modules to generate a set of three-dimensional (3D) key points (e.g., key points in a 3D space) for a fish given a set of 2D key points (e.g., key points in a 2D space) captured in 2D image data and the outputs of one or more of the multiple modules. For example, the multiple modules may implement one or more ofthe following functionalities: (i) structure from motion, (ii) direct regression, (iii) eye sizebased calibration, (iv) truss length ratios, and (v) observed fiducial marker calibration.

[0018] In some embodiments, an output of the Bayesian framework can be provided to a trained model (e.g., neural network, Random Forest Regressor, Support Vector Regressor, or Gaussian Process Regressor, among others) that is trained to generate predicted biomass estimations of individual fish based on the output of the Bayesian framework, for example, the truss lengths (e.g., which can be extracted from the set of 3D key points). The biomass of fish populations may be used to perform an action, for example, to control the amount of feed given to a fish population, e.g., by controlling a feed distribution system, as well as to identify and isolate runt, diseased, or other subpopulations.

[0019] As used in this specification, a truss measurement is determined based on locations of 3D key points (e.g., key points in 3D space) of a fish. For example, a truss length is a distance between particular morphological areas of interest of fish anatomy. Generally, truss measurements can be used as a taxonomic tool for stock identification (e.g., to identify different species of fish). A set of truss lengths can be used to estimate a volume of a fish, and, assuming a density of fish for a species of fish, a biomass estimate for the fish.

[0020] FIG. 1 is a diagram showing an example of a system 100 that is used for underwater camera biomass estimation in an aquatic environment. The system 100 includes a control unit 116 and an underwater camera device 102. The control unit 116 obtains images captured by a camera of the camera device 102 and processes the images to generate biomass estimations for one or more fish 106. In some embodiments, fish 106 can be constrained by a pen 104 within an aquatic environment, e.g., in a livestock tank or a net. In some embodiments, fish 106 can be free to move with respect to camera 102, e.g., in an open ocean, river, lake, or another aquatic environment. The biomass estimations for one or more fish can be processed to determine actions such as population monitoring, sorting, model training, and user report feedback, among others.

[0021] In some implementations, the control unit 116 includes one or more components for processing data. For example, the control unit 116 can include an object detector 118, module(s) 120, Bayesian Framework 122, biomass engine 124, and a decision engine 128. Components can include one or more processes that are executed by the control unit 116.

[0022] At least one of the one or more cameras of the camera device 102 includes a camera that captures images from a single viewpoint at a time. This type of camera may be referred to as a monocular camera. Where a stereo camera setup can include multiple cameras each capturing a unique viewpoint at a particular time, a monocular camera captures one viewpoint at a particular time. A computer processing output of a stereo camera setup can determine, based on differences in the appearance of objects in one viewpoint compared to another viewpoint at a particular time, depth information of the objects.

[0023] In some implementations, the camera device 102 has one or more cameras in a stereo camera setup that are non-functional or obscured, e.g., by debris or other objects, including fish or other animals, in an environment. In some implementations, the control unit 116 can process one or more images from the camera device 102, or obtain a signal from the camera device 102, indicating that one or more cameras of the camera device 102 are obscured or non-functional or the camera device 102 is operating as a monocular camera. The control unit 116 can adjust a processing method based on the status of the cameras of the camera device 102, such as a stereo setup status or a monocular camera setup status. In some implementations, a monocular underwater camera includes stereo camera setups that have become, operationally, a monocular camera setup based on non-functioning elements, debris, among other causes. In some cases, stereo camera setups can obtain images from a single viewpoint at a given time and be considered, operationally, a monocular camera.

[0024] In some implementations, the camera device 102 includes a single camera. For example, the single camera can be a monocular camera that captures a single viewpoint at a time. In some implementations, the camera device 102 with a singlecamera is more efficient to produce and maintain, more economical, and can be more robust with fewer components prone to failure.

[0025] In some embodiments, e.g., as depicted in FIG. 1 , the camera device 102 includes a propulsion mechanism, e.g., propellers, to move the camera device 102 around an aquatic environment, e.g., an ocean, a river, a lake, a livestock tank, etc. In general, the camera device 102 may use any method of movement including ropes and winches, waterjets, thrusters, tethers, buoyancy control apparatus, chains, among others.

[0026] In some embodiments, the camera device 102 may include a floatation device, e.g., a buoy, an inner tube, etc. The camera device can be free-floating, where propulsion of the camera 102 through the aquatic environment may be directed by a movement of the current / waves surrounding the camera device 102.

[0027] In some implementations, the camera device 102 is equipped with the control unit 116 as an onboard component, while in other implementations, the control unit 116 is not affixed to the camera device 102 and is external to the camera device 102. For example, the camera device 102 may provide image data 110, e.g., including image 112, over a network to the control unit 116. Similarly, the control unit 116 can provide return data, including movement commands to the camera device 102 over the network.

[0028] In some implementations, camera device 102 includes an onboard inertial measurement unit (IMU) 105 configured to collect a pose of the camera device 102, an angular and linear velocity of the camera device 102, etc. Control unit 116 can receive IMU data from the IMU and associate the IMU data with one or more images captured by the camera device 102, e.g., as metadata to enrich the image data 110 captured by the camera.

[0029] Stages A through C of FIG. 1 , depict image data 110, including multiple frames captured by the camera 102, e.g., including image 112, that are processed by the control unit 116. The image 112 includes representations of the fish 113 and 115. Images of fish obtained by the camera device 102 may include fish in any conceivable pose including head on, reverse head on, or skewed. The image 112 can additionallyinclude one or more additional components, e.g., food pellets, fiducial markers, robotic or pre-tagged fish, etc.

[0030] In stage A, the camera device 102 obtains the image data 110 including image 112 of the fish 113 and 115 within the aquatic environment. The camera device 102 provides the data 110 to the control unit 116. In some implementations, the camera device 102 obtains multiple images and provides the multiple images, including the image 112, to the control unit 116.

[0031] In stage B, the control unit 116 processes the images of the data 110, including the image 112. The control unit 116 provides the data 110 to the object detector 118. The object detector 118 can run on the control unit 116 or be communicably connected to the control unit 116. The object detector 118 detects one or more objects in the images of the data 110. For example, object detector 118 can use computer vision and machine learning-based techniques to detect the one or more objects within the image data 110. The one or more objects can include large scale objects, such as fish, as well as smaller objects, such as key points of the fish in two-dimensional (2D) space, e.g., 2D key points. In the example of FIG. 1 , the object detector 118 detects the fish 113 and the fish 115 in the image 112.

[0032] The object detector 118 can provide image data 110 as output to the Bayesian framework 122 including module(s) 120, where the image data 110 includes at least one fish detected within the images, e.g. , fish 113, 115 in image 112. Module(s) 120 can include one or more of: (i) structure from motion, (ii) direct regression, (iii) eye sizebased calibration, (iv) truss length ratios, (v) observed fiducial marker calibration, the operations of which are described in further detail below.

[0033] In some embodiments, an output of the Bayesian framework 122 includes a biomass estimate for a fish. The control unit 116 can use biomass estimates generated by the Bayesian Framework 122 to generate a biomass distribution 126 including one or more biomasses, corresponding to one or more fish.

[0034] In some embodiments, an output of the Bayesian framework 122 includes truss lengths, a set of 3D key points, or the like, for a fish, from which a biomass estimate of the fish can be generated. In such instances, the output of the Bayesian framework 122can be provided to a biomass engine 124. Biomass engine 124 can include a trained model (e.g., neural network, Random Forest Regressor, Support Vector Regressor, or Gaussian Process Regressor, among others) that is trained to generate predicted biomass estimations of individual fish based on the output, for example, the truss lengths (e.g., which can be extracted from the set of 3D key points).

[0035] The biomass distribution 126 generated by the biomass engine 124 includes one or more biomasses, corresponding to one or more fish, of the fish 106. In the example of FIG. 1 , a number of fish corresponding to ranges of biomass is determined by the biomass engine 124. For example, the biomass engine 124 can determine that two fish correspond to biomass range 3.5 kilogram to 3.7 kilogram and three fish correspond to the biomass range 3.7 kilogram to 3.8 kilogram. In some implementations, the ranges are predetermined. In some implementations, the ranges are dynamically chosen by the control unit 116 based on the number of fish and the distribution of biomasses. In some implementations, the biomass distribution 126 is a histogram.

[0036] The biomass engine 124 generates the biomass distribution 126 and provides the biomass distribution 126 to the decision engine 128. The decision engine 128 obtains the biomass distribution 126. In stage C, the control unit 116 determines an action based on the biomass distribution 126. In some implementations, the control unit 116 provides the biomass distribution 126, including one or more biomass estimation values, to the decision engine 128.

[0037] In some implementations, the biomass engine 124 generates the biomass distribution 126 that includes likelihoods that a number of fish of the fish 106 are a particular biomass or within a range of biomass. For example, weight ranges for the biomass distribution 126 can include ranges from 3 to 3.1 kilograms (kg), 3.1 to 3.2 kg, and 3.2 to 3.3 kg. A likelihood that a number of the fish 106 are within the first range, 3 to 3.1 kg, can be 10 percent. A likelihood that the number of the fish 106 are within the second or third range, 3.1 to 3.2 kg or 3.2 to 3.3 kg, respectively, can be 15 percent and 13 percent. In general, the sum of all likelihoods across all weight ranges can be normalized (e.g., equal to a value, such as 1 , or percent such as 100 percent).

[0038] In some implementations, the decision engine 128 determines a portion of the fish 106, based on data generated by the biomass engine 124, are below an expected weight or below weights of others of the fish 106. For example, the decision engine 128 can determine subpopulations within the fish 106 and determine one or more actions based on the determined subpopulations, such as actions to mitigate or correct for issues (e.g., runting, health issues, infections, disfiguration, among others). Actions can include feed adjustment, sorting, model training, and user report feedback, among others.

[0039] In some implementations, the control unit 116 includes multiple computer processors. For example, the control unit 116 can include a first and a second computer processor communicably connected to one another. The first and the second computer processor can be connected by a wired or wireless connection. The first and second computer processors can each perform one or more of the operations described in this specification. Operations not performed by the first computer processor can be performed by the second computer processor or an additional computer processor. Operations not performed by the second computer processor can be performed by the first computer processor or an additional computer processor.

[0040] In some implementations, the control unit 116 operates one or more processing components described in this specification. In some implementations, the control unit 116 communicates with an external processor that operates one or more of the processing components. The control unit 116 can store training data, or other data used to train one or more models of the processing components, or can communicate with an external storage device that stores data including training data.

[0041] In general, the control unit 116 can process one or more images of the data 110 in aggregate or process each image of the data 110 individually. In some implementations, one or more components of the control unit 116 process items of the data individually and one or more components of the control unit 116 process the data 110 in aggregate.

[0042] In general, processing individually as discussed herein, can include processing one or more items sequentially or in parallel using one or more processors. In someimplementations, the decision engine 128 processes data in aggregate. For example, the decision engine 128 can determine, based on one or more data values indicating biomass estimates provided by the biomass engine 124, one or more decisions and related actions, as described herein. In general, processing in aggregate as discussed herein, can include processing data corresponding to two or more items of the data 110 to generate a single result. In some implementations, an item of the data 110 includes the image 112.

[0043] In some implementations, the control unit 116 obtains data from a monocular underwater camera indicating a current operation status of the monocular underwater camera and, in response to obtaining an image of a fish, provides the image of the fish to a trained model. For example, the camera 102 may include a dysfunctional stereo camera pair. The dysfunction can result in images without the depth data provided by the stereo effect of the cameras. To mitigate this situation, the camera device 102 can send a signal to the control unit 116 indicating a camera of a stereo pair is dysfunctional after the camera device 102 determines a camera of a stereo pair has become dysfunctional. In response to the signal, the control unit 116 can process images obtained from the camera device 102 as discussed herein.

[0044] In some implementations, the control unit 116 determines, based on the data 110 or other data provided by the camera device 102, the images from the camera device 102 do not include depth data. For example, the control unit 116 can process the data 110 and determine that images lack a stereo feature. In response, the control unit 116 can process the images to determine a biomass distribution without this additional depth data.

[0045] As described above with reference to Stage B, the Bayesian framework 122 can integrate one or more of the multiple modules 120 to support multiple aspects of the Bayesian framework. The operations of modules 120 and the Bayesian framework 122 are described in further detail below. Though described below as operations performed by the modules 120 and the Bayesian framework 122, the operations can be performed by a control unit 116, where control unit 116 is programmed to perform the operations ofthe modules 120 and Bayesian framework 122.Structure from Motion Module

[0046] In some embodiments, a structure from motion (SFM) module can receive 2D images captured by a monocular camera as input and generate, as output, estimated 3D truss lengths. The SFM module can disambiguate the scale of fish captured in 2D images by tracking salient key points of a fish over several image frames (e.g., 10 frames) to infer a size of the fish. In other words, the SFM module can reduce an ambiguity due to scale, for example, where an apparent size of a fish is 2x when the fish is located at a 2x distance from the camera, by constraining a size of a fish to a set range of sizes. The SFM module can track a fish moving through several sequential frames and / or the camera moving through several sequential frames (e.g., such that the camera sees the fish from one or more angles) while assuming that a shape / size of the fish does not change from frame-to-frame as well as the motion of the fish is smooth from frame-to-frame. The trajectory of the 2D key points of the fish can be tracked through the multiple frames. From the observations of the 2D key points over time, the SFM module can fit a hypothesis of fish motion, including 3D key points of the fish and rotation and translation of the fish over time.

[0047] In some embodiments, the SFM module receives IMU data from an IMU associated with the camera setup. IMU data can include information related to camera motion including position in 3D space (e.g., pose), linear and angular velocity, and linear and angular acceleration.

[0048] In some embodiments, SFM methods by the SFM module can include denoting 3D world coordinates for points on a fish as X_world(i) where i is an index of a key point on a fish. The 2D screen coordinates of the same points on the fish are denoted as X_screen(i). Using a pinhole camera system assumption, the module applies a relationship between the 3D world coordinates and the 2D screen coordinates using a camera projection matrix P as:

[0049] X_world(i) = P*X_screen(i)

[0050] A structure of the projection matrix P is defined as

[0051]

[0052] Where K is an intrinsic matrix and [R|t] is an extrinsic matrix. The intrinsic matrix K includes a 2D translation component, a 2D scaling component, and a 2D shear component. The intrinsic matrix K includes point (xO, yO) as the optical center of the image in pixels, fx, fy as the focal lengths, s as the shear value. The intrinsic matrix values of K can be obtained from known camera calibration, e.g., using a known calibration target. The extrinsic matrix [R|t] includes a rotation component and a translation component. The SFM module is configured to extract the rotation and translation components of the extrinsic matrix using tracking of a fish through multiple frames captured by the camera. In some embodiments, the camera is non-stationary (e.g., free floating) such that two matrices are needed over each time step t for both the fish, E_fish, and the camera, E_camera.

[0053] The module generates a system of equations such that E_fish(t) * X_world(i) = K * E_camera(t) * X_screen(i, t), which is solved for the 3D points of the fish as:

[0054] X_world(i) = K * E_camera(t) * X_screen(i, t) * E_fish(t)A-1 , for instances where at least (A) a threshold number of key points on a fish, (B) a threshold number of fish, or (C) a threshold static background is larger than a set of unknown parameters for the camera system. Using the X_world(i) values of the fish, the module can use truss lengths to determine a volume of the fish and / or a biomass of the fish.

[0055] In some embodiments, the SFM module measures lens distortion parameters for the camera using the captured images using a lens distortion model for the pinhole camera model to relate X_screen(i) = undistort(X_raw(i)) , where X_raw(i) are the key points originally measured in the image and X_screen is an idealized undistorted screen space 2D coordinates. The lens distortion model can be written as:

[0056]

[0057] where X_screen = (u, v) and X_raw = (x’, y’) and X_camera = E_camera * X_world = (X_c, Y_c, Z_c).

[0058] In some embodiments, at least parts of the calibration procedures performed by the SFM module can be used in one or more of the other modules described herein, in order to compensate for different camera configurations (e.g., for cameras where a prior calibration is not available or where in-field calibration is required). For example, the direct regression module can use the camera calibration generated by the SFM module as input to perform the methods of the direct regression module.Direct Regression Module

[0059] In some embodiments, a direct regression module includes a deep learning model trained on visual cues extracted from images captured by the monocular camera to make predictions of a weight of a fish in an image. Visual cues can include, for example, (i) a relative size of a fish’s eye to other body parts, (ii) relative 2D distances between key points on the fish, (iii) reference objects in the environment that have recognizable approximate sizes, or (iv) any combination thereof. The deep learning model approach for the direct regression module does not require humans to identify / extract appropriate features from a domain knowledge of fish appearance.Additionally, if the appropriate features vary between fish species, these features can be learned automatically for each species without repeating the effort of feature discovery.

[0060] In some embodiments, a direct regression approach by the direct regression module includes the methods of receiving an image captured by the camera and applying a pre-processing to generate an undistorted image. For example, undistorting the image can be performed using a measured camera calibration or a camera calibration learned from another module (e.g., the structure from motion module). The methods include applying an object detection or instance segmentation model to the undistorted image to select a region of the undistorted image including a single fish and a number of additional visual cues. For example, the selected region can be a padded crop including a threshold border surrounding the fish. A number of additional visual cues can include at least a threshold number of additional visual cues. In some embodiments, the selected region is not resized in order to preserve original scale information for the image. Alternatively, scale information can be encoded as a separate feature input to a deep regression model.

[0061] In some embodiments, if a threshold number of visual cues can be identified from the shape and appearance of a single fish to infer weight, the selected region of the undistorted image can be tight around that fish, e.g., include a minimal threshold border surrounding the fish. If a threshold number of visual cues cannot be identified from the shape and appearance of the single fish to infer weight, an additional border region surrounding the fish may be necessary, for example, to include nearby fish, feed pellets, etc., that can provide additional cues for estimating the weight of the primary fish.

[0062] The methods further include applying a deep regression model to the selected region which receives the selected region (e.g., and encoded scale information) as input and provides, as output, a prediction including an estimate of the weight of the fish in the selected region of the image.

[0063] In some embodiments, the deep regression model is trained to discover a set of invariant characteristics for a fish, e.g., for a species of fish. Training a deep regression model can include supervised learning, e.g., providing a supervised signal to the model. Supervised signals can be generated, for example, by supervising controlled sets of fishpopulations using harvest distributions. In some embodiments, supervised signals can be generated from collected images, where each image includes single fishes of known weight.

[0064] In some embodiments, supervised signals for training the deep regression model can be generated using a stereo-based camera assembly combining two monocular cameras, where fish biomass can be estimated from key points in 3D space extracted from stereo images captured by the stereo camera assembly. The trained model can then perform inference on images captured by a (single) monocular camera assembly without requiring a stereo-based camera assembly.Eye Size-based Calibration Module

[0065] In some embodiments, an eye size-based calibration module for estimating fish biomass from 2D images can be performed based on assumptions about one or more known and detectable features of a fish. Although described in this section as using eye size for performing the estimates of biomass, one or more other features may be used. Statistically across a large (e.g., millions) population of fish, the size of a fish eye is relatively constant across individual fish, such that a size (e.g., diameter) of a fish eye can be used to disambiguate a scale of the fish.

[0066] In some embodiments, as depicted in FIG. 2A, for a species of fish, an eye area corresponding to the fish eye can be for example, on average about 0.87 cm2for fish having biomasses ranging between about 2.2 kg to 4.1 kg. The consistency of the average area of the eye of the fish can be used as a calibration point for the camera, e.g., to understand a scaling factor for fish dimensions. For example, if an image includes a fish having an eye area of 42 pixels and a fish body of 42880 pixels, and the eye has an expected average eye area (e.g., from population / species studies) of 1 cm2,42 vix6 Is a scaling factor of — — — can be used to scale other fish dimensions (e.g., trusslengths), accordingly.

[0067] In some embodiments, algorithms described with reference to the SFM module can be used to infer camera parameters, e.g., lens distortion and other intrinsic camera properties. Calibrating the camera parameters can reduce a likelihood of distortion resulting from a fish being close to the camera lens, which can lead to significant distortion of the image of the fish, e.g., such that the fish may appear larger than it actually is. In some embodiments, the structure from motion module can be used to integrate information over several frames captured by the camera.

[0068] In some embodiments, as depicted in FIG. 2B, an eye size-based calibration by the eye size-based calibration module includes the methods of detecting the fish eye. A program detects the fish’s eye. This can be detected from a computer vision algorithm based on one 2D image or a series of images. The detection algorithm may use other features on the fish to confirm this is an eye (e.g., relationship to fins, shape, color, size, or texture). For example, parts of a (healthy) fish eye, e.g., a pupil can be assumed tobe approximately circular and planar. A fish eye can include contrast between an outline region (lighter color) and a pupil (darker color).

[0069] The methods include fitting a shape of the fish eye. The program fits a geometric shape (such as an ellipse) to the detected eye, and for the fish’s body outline.

[0070] The methods include projecting the shapes of the fish to a flat plane. If the fish is at an angle to the camera, then the eye shape can be projected onto the known flat plane of the image. The fish’s body can be assumed to be approximately on the same plane, and the program projects the fish body outline onto the same flat plane.

[0071] The methods include computing an area of the fish body. Using the assumed area of the eye, the program extrapolates to compute the approximate area of the fish’s body. From this area, the program can estimate biomass of the fish.Truss Length Ratio Module

[0072] In some embodiments, a 3D size of a fish is inferred based on 2D truss lengths in units of pixels of the images captured by the camera. The truss length ratio module may additionally infer the 3D size of the fish using quantities available, e.g., eye size, fiducial markers, and direct-from-pixel extraction, using perception models applied to monocular images or from outputs of one or more of the other modules described herein. The inference by the truss length ratio module relies on an assumption that small fish and big fish are not proportional copies of each other. In other words, small fish of a first range of dimensions (e.g., biomass) have a first set of truss length ratios and big fish of a second range of dimensions (e.g., biomass) have a second set of different truss length ratios. For example, as depicted in FIG. 3A, the three fish of different sizes have different ratios of trusses T 13 / T24, for a given species of fish. For a single species of fish having a consistent pattern of different ratios of trusses, 2D images of these fish (with corresponding bounding boxes and key points) can be used to differentiate between their weights.

[0073] In some embodiments, a fish can include K key points and N trusses with a variety of selectable truss ratios. For example, a salmon includes K = 10 key points andN = 45 trusses. As depicted in FIG. 3B, a truss length ratio module includes a neural network that receives, as input, the 2D sizes of the trusses for a fish along with supporting variables (e.g., the bounding box area, the explicit values of all truss ratios, the explicit values of all trusses squared, etc.) that are output from applying a deep learning perception model on the images of a monocular camera. The neural network architecture can take in inputs from the monocular image and output a weight of the fish, e.g., in grams. To train it with ground truth, the module can either use weights from single fish in a specialized rig or use a distribution-based approach. With this data driven approach, the neural network can learn all appropriate correlations from the data directly and come up with the best possible representation of the actual weight given the 2D inputs.

[0074] In some embodiments, an output of the truss length ratio module can be used to physically constrain solutions (e.g., height vs length ratio) for possible fish dimension in the fusion methodology described below.Observed Fiducial Marker-based Calibration Module

[0075] In some embodiments, a reference object of known dimensions, for example, an observed fiducial marker as depicted in FIG. 4A, can be inserted into a field of view of the camera in order to calibrate nearby fish and / or accurately estimate depth distance for applying a stereo-based truss model. The observed fiducial marker can be used to calibrate a set of camera intrinsic properties and reduce (e.g., remove) lens distortion. For example, as depicted in FIG. 4B, a live or robotic fish of known dimension that can insert itself next to one or more fish, e.g., into a school of fish, can be used as a visual size reference. In another example, a fish identification (fish-ID) system that identifies fish that have been previously weighed and measured can be used as a size reference in images captured by a monocular camera to calibrate other nearby fish in the images. In another example, a QR code can be used as a size reference in images where the QR code is visible alongside fish.

[0076] In some embodiments, a distance to an observed fiducial marker can be calculated using a method of similar triangles. For example, an image can include afiducial marker with a known dimension, e.g., 10 cm length, where the dimension of the fiducial marker in pixels within the image of 20 pixels. The module can calibrate a distance to the fiducial marker as 10 cm / distance to fiducial = 20 pixels / fx(focal length of the camera), such that the distance to the fiducial = fx * 10 / 20.

[0077] In some embodiments, an output of the observed fiducial marker-based calibration module can be used to provide a physical means to constrain a solution for the fusion methodology described below.Bayesian Framework for Biomass Estimation

[0078] As described with reference to FIG. 1 , a system for fish biomass estimation in an aquatic environment includes a Bayesian framework engine. The Bayesian framework engine can incorporate one or more of the modules described herein, where an input to the framework is a set of key points in two-dimensional (2D) space extracted from captured 2D images including a fish, and an output from the framework is a set of key points in three-dimensional (3D) space that correspond to at least a threshold consistency with the outcomes of the one or more modules. In some embodiments, the modules include (A) structure from motion, (B) direct regression, (C) eye size-based calibration, (D) truss length ratio, (E) observed fiducial marker-based calibration.

[0079] The Bayesian framework of the Bayesian framework module can be implemented as:

[0080] P(hypothesis | data) = P(data | hypothesis) * P(hypothesis) I P(data)

[0081] where P(hypothesis | data) is a probability of a hypothesis for given data where a higher probability outcome implies that the hypothesis has a greater likelihood of being true. As used in this specification, P(hypothesis | data) is a posterior probability, P(hypothesis) is a prior probability, and P(data | hypothesis) is an evidence term. In some embodiments where the outcome of the Bayesian framework is a maximum value (e.g., optimized value) for P(hypothesis | data), P(data) is a constant value which can be discarded.

[0082] In some embodiments, a hypothesis can be P(estimated biomass of the fish). In other words, the hypothesis can be the estimated mass (or density and volume estimates) of the fish. In some embodiments, a hypothesis can be P(fish model). For example, the hypothesis can be a model representation of a fish, which can include, for example, a biomass, eye size, 2D or 3D truss length ratios, etc. In some embodiments, a hypothesis can be P(truss) (e.g., 3D structure of the fish connecting each 3D key point to a dorsal or tail fin). As used in this specification, truss is a 3D structure of the fish connecting each key point in 3D space, for example, eye-to-dorsal fin-to-tail fin. For example, 10 key points in 3D space corresponds to 30 x, y, z coordinates of the fish in cartesian coordinates.

[0083] In some embodiments, the one or more modules can each contribute independently to a posterior term. In some embodiments, one or more of the multiple modules can be used to reinforce or modify the estimates determined by one or more other modules. For example, two or more modules can contribute to the posterior term by multiplying the data generated from each together, e.g., P(hypothesis | data_1 , data_2) = P(hypothesis | data_1 )*P(hypothesis | data_2), where data from each module represent respective measurements is assumed to correspond to independent observations.

[0084] In some embodiments, one or more modules can be incorporated to extend an evidence term, P(data | hypothesis), e.g., can be used as evidence supportive of the fish model. For example, eye size-based calibration and / or fiducial marker calibration can be used as constraints on modeled truss lengths, e.g., P(truss length ratios | hypothesis). In some embodiments, at least one of the modules described below can be implemented independently to generate fish biomass estimates.

[0085] Monte Carlo, Gibbs sampling, Hamiltonian Monte Carlo, or another similar approach, can be used to propose parameters that maximize the posterior probability, P(hypothesis | data). The model parameters (hypothesis) can be, for example, (i) the key points in 3D space of the fish, (ii) the angular and linear velocity of the fish, (iii) the pose of the camera, and (iv) the angular and linear velocity of the camera. The data in this case are the 2D key points projected onto the screen that we observe. Forexample, the Bayesian framework includes a predictor step, e.g., a forward pass, followed by a corrector state which uses back projection to validate if a predicted location of a fish in a subsequent frame matches a prior frame, e.g., based on the model parameters. A search step can be implemented to correct defects from the predictor step, based on an outcome of the comparison, where the Monte Carlo, Gibbs sampling, or Hamiltonian Monte Carlo algorithms can be used to correct these defects. In another example, a gradient descent (or another gradient method) can be used as a lower computational requirement method. In some embodiments, the initial 3D model is perturbed, e.g., using finite difference methods.

[0086] An output of the Bayesian framework can include a set of 3D key points for a given fish for an input including a set of 2D key points. The set of 3D key points can be provided, as input, to a biomass engine which can generate, from the set of 3D key points, an estimation for a biomass of the fish.Example Bayesian Framework Method

[0087] In one example, a hypothesis of the Bayesian framework is a fish model, e.g., P(fish model), where the fish model is a model representation of a fish (e.g., which may include biomass and other features such as eye size, 2D or 3D truss length ratios, etc.) The fish model additionally includes a set of variables that account for the movement of the fish swimming. The set of variables includes (i) the truss, (ii) the angular and linear velocity of the fish, (iii) the angular and linear acceleration of the fish, (iv) the initial location of the fish over a set of frames T.

[0088] The Bayesian framework can be used to estimate a set of joint variables, e.g., variables occurring in a dependent joint probability distribution of P(Hypothesis = {3D truss, fish pose, camera pose} | data), where solving for the joint probability yields the truss including a n x 3 matrix of key points in 3D space, where n is a number of key points. Rotation of the truss at time T is a 3x3 matrix with 3 degrees of freedom. Translation includes 3 degrees of freedom.

[0089] In some embodiments, a Rodrigues rotation formulation is 3D rotation group is used to transform a set of basis vectors to compute a rotation matrix in a SO(3) Lie algebra representation (6 degrees of freedom). In other words, to convert the samples, e.g., the 3D coordinates representing rotation, to the matrices, e.g., the 3x3 rotation matrices, for both Fish models and camera models.

[0090] The fish model can be a linear model or deep neural network used to estimate fish weights from truss lengths. The estimation of truss lengths and fish weight are separated such that other physically-based methods (e.g., described above with reference to the modules) can be applied to the fish model, e.g., structure from motion, truss ratios, and eye size, as constraints on the modeled truss lengths. The fish model can be written as:

[0091] Fish at time T = Truss * rotation at time T + translation at time T, where rotation and translation are governed by Newton’s laws of motion, and time T corresponds to an image frame captured by the camera. Joint prior values of the Bayesian framework can be applied on the joint distribution between frames of the camera to encourage smoothness, e.g., in the motion of the fish and / or motion of the camera.

[0092] The Bayesian framework can integrate an output, e.g., the camera motion, of the SFM module as:

[0093] P(Fish model | 2D key points, camera motion) = P(2D key points, camera motion | Fish Model) * P(Fish model)

[0094] Where the P(Fish model | 2D key points, camera motion) is searched for the maximum likelihood estimate (MLE) of the fish model by initially using a midpoint algorithm. For example, a MLE of the fish model can be found by simulating rays, e.g., raycasting, from the first frame and intersecting through each frame over time. In another example, only an initial frame is used, assuming a standard scale for a fish, e.g., 20 cm width. Monte Carlo, Gibbs sampling, or Hamiltonian Monte Carlo is used to propose sample variables that maximize the posterior probability, as described above.

[0095] In some embodiments, one or more of the outputs of the multiple modules can be incorporated into the Bayesian framework to extend the evidence term. For example,an output of the eye size-based calibration module can be incorporated as: P(Eye size | Fish model) as the evidence that supports P(Fish model | eye size). In another example, an output of the truss length ratio module can be incorporated as: P(Truss ratios | Fish model) as the evidence that supports P(Fish model | truss ratios). In another example, an output of the observed fiducial marker calibration module, e.g., in situ camera calibration, can be used to recover the camera intrinsic values.

[0096] In some embodiments, an output of the direct regression module can also be integrated by adding the final fish weight to the jointly distributed degrees of freedom in the fish model, i.e. P(truss, weight). Methods such as normalizing flow, or variational autoencoders can also be used to estimate a probability distribution of the regressed fish weight rather than a point sample. Alternatively, the output of the direct regression module can be used as a residual correction applied to the final weight estimate produced by combining the other methods. In this case, the direct regression needs only estimate unmodelled effects or bias in a prior estimate of the model.

[0097] Assuming independence the individual evidence terms, the evidence terms can be multiplied together as: P(Fish model | 2D keypoints, camera, truss ratios, eye size, direct regression) = P(Fish Model) * P(data | fish model) where P(data | fish model) = P(eye size | fish model) * P(truss ratios | fish model) * P(fiducial comparisons | fish model).

[0098] An output of the Bayesian framework for the above example, includes a fish model representative of a fish including at least a biomass of the fish, e.g., for a given input of 2D key points for the fish extracted from captured images by a monocular camera.

[0099] FIG. 5 is a flow diagram showing an example of a process 500 for monocular underwater camera biomass estimation. The process 500 may be performed by one or more systems, for example, the system 100 of FIG. 1 . One or more aspects of the process 500 can be performed by, for example, control unit 116. The system obtains multiple frames captured by a monocular camera over a period of time, each of the multiple frames including an image of a fish (502). The image data 110 can include multiple frames of 2D image data captured of an aquatic environment by a monocularcamera 102. The image data 110 can additionally include metadata, e.g., IMU data collected by an IMU 105 that is coupled to the camera 102.

[0100] The system identifies a first set of key points in a 2D space for the fish in the multiple frames (504). In some embodiments, an object detector 118 receives image data 110 as input and identifies fish within images of the image data, e.g., fish 113, 115 in image 112. The object detector can identify 2D key points for the identified fish, e.g., using trained models.

[0101] The system determines motion of the monocular camera of the period of time (506). Control unit 116 can receive IMU data from IMU 105 for the period of time corresponding to the image data 110 and extract, from the IMU data, information about the motion of the camera 102 during the period of time. For example, linear and angular velocity, acceleration, pose, etc.

[0102] The system generates, from the first set of key points in 2D space and the motion of the camera, a fish model including a second set of key points in a three- dimensional space for the fish (508). A Bayesian Framework 122 can incorporate one or more modules 120 and can receive the set of key points in 2D space and motion of the camera and provide as output a fish model including a set of 3D key points for the fish. For example, the Bayesian framework 122 can include P(Fish model | 2D key points, camera motion, {dataN of model(s)}) = P(Fish Model) * P(data | fish model) where P(data | fish model) = P(data1 | fish model) * P(data2 | fish model) * P(data3 | fish model)* .. *P(dataN | fish model).

[0103] The system generates a biomass estimate for the fish based on the second set of key points in 3D space (510). A biomass engine 124 can receive output from the Bayesian Framework 122, e.g., truss lengths or 3D key points, and generate a biomass estimate for the fish. The control unit 116 can generate a visualization of a biomass distribution 126 for the estimated biomasses for multiple (e.g., a population) of fish.

[0104] The system determines an action based on one or more biomass estimates including the biomass estimate of the fish (512). The decision engine 128 can, in response to the biomass distribution, determine an action. For example, the decisionengine can generate a recommendation related to a population profile (e.g., health, species, sizes, etc.) for a sampled population in the aquatic environment.

[0105] FIG. 6 is a diagram showing an example of a truss network. The truss lengths between key points are used to extract information about the fish including a weight of the fish. Various trusses, or lengths between key points, of the fish can be used. FIG. 6 shows a number of possible truss lengths including upper lip 602 to eye 604, upper lip 602 to leading edge dorsal fin 606, upper lip 602 to leading edge pectoral fin 608, leading edge dorsal fin 606 to leading edge anal fin 610, leading edge anal fin 610 to trailing low caudal peduncle 612, trailing lower caudal peduncle 612 to trailing upper caudal peduncle 614. Other key points and other separations, including permutations of key points mentioned, can be used. For different fish, or different species of fish, different key points may be generated. For any set of key points, a truss network may be generated as a model.

[0106] FIG. 7 is a diagram illustrating an example of a computing system used for monocular underwater camera biomass estimation. The computing system includes computing device 700 and a mobile computing device 750 that can be used to implement the techniques described herein. For example, one or more components of the system 100 could be an example of the computing device 700 or the mobile computing device 750, such as a computer system implementing the control unit 116, devices that access information from the control unit 116, or a server that accesses or stores information regarding the operations performed by the control unit 116.

[0107] The computing device 700 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The mobile computing device 750 is intended to represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smart-phones, mobile embedded radio systems, radio diagnostic computing devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only and are not meant to be limiting.

[0108] The computing device 700 includes a processor 702, a memory 704, a storage device 706, a high-speed interface 708 connecting to the memory 704 and multiple high-speed expansion ports 710, and a low-speed interface 712 connecting to a low- speed expansion port 714 and the storage device 706. Each of the processor 702, the memory 704, the storage device 706, the high-speed interface 708, the high-speed expansion ports 710, and the low-speed interface 712, are interconnected using various busses, and may be mounted on a common motherboard or in other manners as appropriate. The processor 702 can process instructions for execution within the computing device 700, including instructions stored in the memory 704 or on the storage device 706 to display graphical information for a GUI on an external input / output device, such as a display 716 coupled to the high-speed interface 708. In other implementations, multiple processors and / or multiple buses may be used, as appropriate, along with multiple memories and types of memory. In addition, multiple computing devices may be connected, with each device providing portions of the operations (e.g., as a server bank, a group of blade servers, or a multi-processor system). In some implementations, the processor 702 is a single threaded processor. In some implementations, the processor 702 is a multi-threaded processor. In some implementations, the processor 702 is a quantum computer.

[0109] The memory 704 stores information within the computing device 700. In some implementations, the memory 704 is a volatile memory unit or units. In some implementations, the memory 704 is a non-volatile memory unit or units. The memory 704 may also be another form of computer-readable medium, such as a magnetic or optical disk.

[0110] The storage device 706 is capable of providing mass storage for the computing device 700. In some implementations, the storage device 706 may be or include a computer-readable medium, such as a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid-state memory device, or an array of devices, including devices in a storage area network or other configurations. Instructions can be stored in an information carrier. The instructions, when executed by one or more processing devices (for example, processor 702), perform one or more methods, such as those described above. The instructions canalso be stored by one or more storage devices such as computer- or machine readable mediums (for example, the memory 704, the storage device 706, or memory on the processor 702). The high-speed interface 708 manages bandwidth-intensive operations for the computing device 700, while the low-speed interface 712 manages lower bandwidth-intensive operations. Such allocation of functions is an example only. In some implementations, the high speed interface 708 is coupled to the memory 704, the display 716 (e.g., through a graphics processor or accelerator), and to the high-speed expansion ports 710, which may accept various expansion cards (not shown). In the implementation, the low-speed interface 712 is coupled to the storage device 706 and the low-speed expansion port 714. The low-speed expansion port 714, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet) may be coupled to one or more input / output devices, such as a keyboard, a pointing device, a scanner, or a networking device such as a switch or router, e.g., through a network adapter.

[0111] The computing device 700 may be implemented in a number of different forms, as shown in the figure. For example, it may be implemented as a standard server 720, or multiple times in a group of such servers. In addition, it may be implemented in a personal computer such as a laptop computer 722. It may also be implemented as part of a rack server system 724. Alternatively, components from the computing device 700 may be combined with other components in a mobile device, such as a mobile computing device 750. Each of such devices may include one or more of the computing device 700 and the mobile computing device 750, and an entire system may be made up of multiple computing devices communicating with each other.

[0112] The mobile computing device 750 includes a processor 752, a memory 764, an input / output device such as a display 754, a communication interface 766, and a transceiver 768, among other components. The mobile computing device 750 may also be provided with a storage device, such as a micro-drive or other device, to provide additional storage. Each of the processor 752, the memory 764, the display 754, the communication interface 766, and the transceiver 768, are interconnected using various buses, and several of the components may be mounted on a common motherboard or in other manners as appropriate.

[0113] The processor 752 can execute instructions within the mobile computing device 750, including instructions stored in the memory 764. The processor 752 may be implemented as a chipset of chips that include separate and multiple analog and digital processors. The processor 752 may provide, for example, for coordination of the other components of the mobile computing device 750, such as control of user interfaces, applications run by the mobile computing device 750, and wireless communication by the mobile computing device 750.

[0114] The processor 752 may communicate with a user through a control interface 758 and a display interface 756 coupled to the display 754. The display 754 may be, for example, a TFT (Thin-Film-Transistor Liquid Crystal Display) display or an OLED (Organic Light Emitting Diode) display, or other appropriate display technology. The display interface 756 may include appropriate circuitry for driving the display 754 to present graphical and other information to a user. The control interface 758 may receive commands from a user and convert them for submission to the processor 752. In addition, an external interface 762 may provide communication with the processor 752, so as to enable near area communication of the mobile computing device 750 with other devices. The external interface 762 may provide, for example, for wired communication in some implementations, or for wireless communication in other implementations, and multiple interfaces may also be used.

[0115] The memory 764 stores information within the mobile computing device 750. The memory 764 can be implemented as one or more of a computer-readable medium or media, a volatile memory unit or units, or a non-volatile memory unit or units. An expansion memory 774 may also be provided and connected to the mobile computing device 750 through an expansion interface 772, which may include, for example, a SIMM (Single In Line Memory Module) card interface. The expansion memory 774 may provide extra storage space for the mobile computing device 750, or may also store applications or other information for the mobile computing device 750. Specifically, the expansion memory 774 may include instructions to carry out or supplement the processes described above, and may include secure information also. Thus, for example, the expansion memory 774 may be provided as a security module for the mobile computing device 750, and may be programmed with instructions that permitsecure use of the mobile computing device 750. In addition, secure applications may be provided via the SIMM cards, along with additional information, such as placing identifying information on the SIMM card in a non-hackable manner.

[0116] The memory may include, for example, flash memory and / or NVRAM memory (nonvolatile random access memory), as discussed below. In some implementations, instructions are stored in an information carrier such that the instructions, when executed by one or more processing devices (for example, processor 752), perform one or more methods, such as those described above. The instructions can also be stored by one or more storage devices, such as one or more computer- or machine-readable mediums (for example, the memory 764, the expansion memory 774, or memory on the processor 752). In some implementations, the instructions can be received in a propagated signal, for example, over the transceiver 768 or the external interface 762.

[0117] The mobile computing device 750 may communicate wirelessly through the communication interface 766, which may include digital signal processing circuitry in some cases. The communication interface 766 may provide for communications under various modes or protocols, such as GSM voice calls (Global System for Mobile communications), SMS (Short Message Service), EMS (Enhanced Messaging Service), or MMS messaging (Multimedia Messaging Service), CDMA (code division multiple access), TDMA (time division multiple access), PDC (Personal Digital Cellular), WCDMA (Wideband Code Division Multiple Access), CDMA2000, or GPRS (General Packet Radio Service), LTE, 7G / 6G cellular, among others. Such communication may occur, for example, through the transceiver 768 using a radio frequency. In addition, short-range communication may occur, such as using a Bluetooth, Wi-Fi, or other such transceiver (not shown). In addition, a GPS (Global Positioning System) receiver module 770 may provide additional navigation- and location-related wireless data to the mobile computing device 750, which may be used as appropriate by applications running on the mobile computing device 750.

[0118] The mobile computing device 750 may also communicate audibly using an audio codec 760, which may receive spoken information from a user and convert it to usable digital information. The audio codec 760 may likewise generate audible soundfor a user, such as through a speaker, e.g., in a handset of the mobile computing device 750. Such sound may include sound from voice telephone calls, may include recorded sound (e.g., voice messages, music files, among others) and may also include sound generated by applications operating on the mobile computing device 750.

[0119] The mobile computing device 750 may be implemented in a number of different forms, as shown in the figure. For example, it may be implemented as a cellular telephone 780. It may also be implemented as part of a smart-phone 782, personal digital assistant, or other similar mobile device.

[0120] A number of implementations have been described. Nevertheless, it will be understood that various modifications may be made without departing from the spirit and scope of the disclosure. For example, various forms of the flows shown above may be used, with steps re-ordered, added, or removed.

[0121] Embodiments of the invention and all of the functional operations described in this specification can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the invention can be implemented as one or more computer program products, e.g., one or more modules of computer program instructions encoded on a computer readable medium for execution by, or to control the operation of, data processing apparatus. The computer readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter effecting a machine-readable propagated signal, or a combination of one or more of them. The term “data processing apparatus” encompasses all apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagated signal is an artificially generatedsignal, e g., a machine-generated electrical, optical, or electromagnetic signal that is generated to encode information for transmission to suitable receiver apparatus.

[0122] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub programs, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communication network.

[0123] The processes and logic flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can also be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit).

[0124] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a processor for performing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a tablet computer, amobile telephone, a personal digital assistant (PDA), a mobile audio player, a Global Positioning System (GPS) receiver, to name just a few. Computer readable media suitable for storing computer program instructions and data include all forms of non volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto optical disks; and CD ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0125] To provide for interaction with a user, embodiments of the invention can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0126] This specification uses the term “configured to” in connection with systems, apparatus, and computer program components. That a system of one or more computers is configured to perform particular operations or actions means that the system has installed on it software, firmware, hardware, or a combination of them that in operation cause the system to perform the operations or actions. That one or more computer programs is configured to perform particular operations or actions means that the one or more programs include instructions that, when executed by data processing apparatus, cause the apparatus to perform the operations or actions. That specialpurpose logic circuitry is configured to perform particular operations or actions means that the circuitry has electronic logic that performs the operations or actions.

[0127] Embodiments of the invention can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., aclient computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the invention, or any combination of one or more such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (“LAN”) and a wide area network (“WAN”), e.g., the Internet.

[0128] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0129] In addition to the embodiments of the attached claims and the embodiments described above, the following numbered embodiments are also innovative.

[0130] Embodiment 1 is a method for estimating fish biomass using a monocular camera, the method comprising: obtaining a plurality of frames captured by the monocular camera over a period of time, wherein at least one frame of the plurality of frames includes a fish; identifying a first set of key points in a two-dimensional space for the fish in the plurality of frames; receiving motion data for the monocular camera over the period of time; generating, from the first set of key points in the two-dimensional space and the motion data for the monocular camera, a fish model comprising a second set of key points in a three-dimensional space for the fish; generating a biomass estimate of the fish based on the second set of key points; and determining an action based on one or more biomass estimates including the biomass estimate of the fish.

[0131] Embodiment 2 is the method of embodiment 1 , further comprising: generating, from the motion data for the monocular camera over the period of time, extrinsic camera parameters for the monocular camera; and applying the extrinsic camera parameters to the plurality of frames captured by the monocular camera for the period of time.

[0132] Embodiment 3 is the method of embodiments 1 or 2, wherein the extrinsic camera parameters comprise three-dimensional translation and three-dimensional rotation of the monocular camera during the period of time.

[0133] Embodiment 4 is the method of any one of embodiments 1 to 3, further comprising: obtaining, for the monocular camera, intrinsic camera parameters for the monocular camera for the period of time.

[0134] Embodiment 5 is the method of any one of embodiments 1 to 4, further comprising obtaining motion data for the fish, wherein generating the fish model comprising a second set of key points in the three-dimensional space for the fish comprises: generating, from the motion data for the fish, a set of motion parameters for the fish including (i) a pose, (ii) translation, and (iii) rotation for the fish; and generating, from the set of motion parameters, the second set of key points in the three-dimensional space for the fish.

[0135] Embodiment 6 is the method of any one of embodiments 1 to 5, wherein identifying the first set of key points in the two-dimensional space for the fish in the plurality of frames comprises: selecting, for each frame of the plurality of frames, a region of interest including a fish and at least one visual cue.

[0136] Embodiment 7 is the method of embodiment 6, wherein generating the fish model comprises providing the region of interest including the fish and the at least one visual cue to a deep regression model as input; and receiving, as output from the deep regression model, a biomass estimate of the fish.

[0137] Embodiment 8 is the method of any one of embodiments 1 to 7, wherein identifying the first set of key points in the two-dimensional space for the fish in the plurality of frames comprises: identifying, in the at least one frame of the plurality of frames including the fish, a size of an eye of the fish in the at least one frame of the plurality of frames; measuring, the size of the eye of the fish in the at least one frame; and generating a calibration for the size of the fish based on the measured size of the eye of the fish.

[0138] Embodiment 9 is the method of any one of embodiments 1 to 8, wherein identifying the first set of key points in the two-dimensional space for the fish in the plurality of frames comprises: providing to a deep learning perception model, theplurality of frames; and receiving an output from the deep learning perception model comprising the first set of key points in the two-dimensional space for the fish in the plurality of frames.

[0139] Embodiment 10 is the method of embodiment 9, wherein generating the biomass estimate of the fish based on the second set of key points further comprises: generating, from output of the deep learning perception model, two-dimensional sizes of a plurality of trusses; providing as input to a trained neural network, the second set of two-dimensional key points and at least one supporting variable; and receiving, from the trained neural network, the biomass estimate of the fish.

[0140] Embodiment 11 is the method of any one of embodiments 1 to 10, wherein identifying the first set of key points in the two-dimensional space for the fish in the plurality of frames further comprises: identifying, in the plurality of frames, a reference object comprising a known dimension within a frame of the plurality of frames; obtaining the known dimension of the reference object; and generating a calibration for a size of the fish based on the known dimension of the reference object.

[0141] Embodiment 12 is a system comprising: one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform the method of any one of embodiments 1 to 11 .

[0142] Embodiment 13 is a computer program carrier encoded with a computer program, the program comprising instructions that are operable, when executed by one or more computers, to cause the one or more computers to perform the method of any one of embodiments 1 to 11 .

[0143] Embodiment 14 is the computer program carrier of embodiment 13, wherein the computer program carrier is a propagated signal.

[0144] While this specification contains many specifics, these should not be construed as limitations on the scope of the invention or of what may be claimed, but rather as descriptions of features specific to particular embodiments of the invention. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also beimplemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.

[0145] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0146] Particular embodiments of the invention have been described. Other embodiments are within the scope of the following claims. For example, the steps recited in the claims can be performed in a different order and still achieve desirable results.

Claims

CLAIMSWhat is claimed is:1 . A method for estimating fish biomass using a monocular camera, the method comprising: obtaining a plurality of frames captured by the monocular camera over a period of time, wherein at least one frame of the plurality of frames includes a fish; identifying a first set of key points in a two-dimensional space for the fish in the plurality of frames; receiving motion data for the monocular camera over the period of time; generating, from the first set of key points in the two-dimensional space and the motion data for the monocular camera, a fish model comprising a second set of key points in a three-dimensional space for the fish; generating a biomass estimate of the fish based on the second set of key points; and determining an action based on one or more biomass estimates including the biomass estimate of the fish.

2. The method of claim 1 , further comprising: generating, from the motion data for the monocular camera over the period of time, extrinsic camera parameters for the monocular camera; and applying the extrinsic camera parameters to the plurality of frames captured by the monocular camera for the period of time.

3. The method of claim 2, wherein the extrinsic camera parameters comprise three- dimensional translation and three-dimensional rotation of the monocular camera during the period of time.

4. The method of any of the preceding claims, further comprising: obtaining, for the monocular camera, intrinsic camera parameters for the monocular camera for the period of time.

5. The method of any of the preceding claims, further comprising obtaining motion data for the fish, wherein generating the fish model comprising a second set of key points in the three-dimensional space for the fish comprises: generating, from the motion data for the fish, a set of motion parameters for the fish including (i) a pose, (ii) translation, and (iii) rotation for the fish; and generating, from the set of motion parameters, the second set of key points in the three-dimensional space for the fish.

6. The method of any of the preceding claims, wherein identifying the first set of key points in the two-dimensional space for the fish in the plurality of frames comprises: selecting, for each frame of the plurality of frames, a region of interest including a fish and at least one visual cue.

7. The method of claim 6, wherein generating the fish model comprises providing the region of interest including the fish and the at least one visual cue to a deep regression model as input; and receiving, as output from the deep regression model, a biomass estimate of the fish.

8. The method of any of the preceding claims, wherein identifying the first set of key points in the two-dimensional space for the fish in the plurality of frames comprises: identifying, in the at least one frame of the plurality of frames including the fish, a size of an eye of the fish in the at least one frame of the plurality of frames; measuring, the size of the eye of the fish in the at least one frame; and generating a calibration for the size of the fish based on the measured size of the eye of the fish.

9. The method of any of the preceding claims, wherein identifying the first set of key points in the two-dimensional space for the fish in the plurality of frames comprises: providing to a deep learning perception model, the plurality of frames; and receiving an output from the deep learning perception model comprising the first set of key points in the two-dimensional space for the fish in the plurality of frames.

10. The method of claim 9, wherein generating the biomass estimate of the fish based on the second set of key points further comprises: generating, from output of the deep learning perception model, two-dimensional sizes of a plurality of trusses; providing as input to a trained neural network, the second set of two-dimensional key points and at least one supporting variable; and receiving, from the trained neural network, the biomass estimate of the fish.11 . The method of any of the preceding claims, wherein identifying the first set of key points in the two-dimensional space for the fish in the plurality of frames further comprises: identifying, in the plurality of frames, a reference object comprising a known dimension within a frame of the plurality of frames; obtaining the known dimension of the reference object; and generating a calibration for a size of the fish based on the known dimension of the reference object.

12. A non-transitory, computer-readable medium storing one or more instructions executable by a computer system to perform the methods of any of claims 1 to 11 .

13. A non-transitory, computer-readable medium storing one or more instructions executable by a computer system to perform operations comprising: obtaining a plurality of frames captured by a monocular camera over a period of time, wherein at least one frame of the plurality of frames includes a fish;identifying a first set of key points in a two-dimensional space for the fish in the plurality of frames; receiving motion data for the monocular camera over the period of time; generating, from the first set of key points in the two-dimensional space and the motion data for the monocular camera, a fish model comprising a second set of key points in a three-dimensional space for the fish; generating a biomass estimate of the fish based on the second set of key points; and determining an action based on one or more biomass estimates including the biomass estimate of the fish.

14. The non-transitory, computer-readable medium of claim 13, further comprising: generating, from the motion data for the monocular camera over the period of time, extrinsic camera parameters for the monocular camera; and applying the extrinsic camera parameters to the plurality of frames captured by the monocular camera for the period of time.

15. The non-transitory, computer-readable medium of claim 14, wherein the extrinsic camera parameters comprise three-dimensional translation and three-dimensional rotation of the monocular camera during the period of time.

16. The non-transitory, computer-readable medium of any of claims 13 to 15, further comprising: obtaining, for the monocular camera, intrinsic camera parameters for the monocular camera for the period of time.

17. The non-transitory, computer-readable medium of any of claims 13 to 16, further comprising obtaining motion data for the fish, wherein generating the fish model comprising a second set of key points in the three-dimensional space for the fish comprises: generating, from the motion data for the fish, a set of motion parameters for thefish including (i) a pose, (ii) translation, and (iii) rotation for the fish; and generating, from the set of motion parameters, the second set of key points in the three-dimensional space for the fish.

18. The non-transitory, computer-readable medium of any of claims 13 to 17, wherein identifying the first set of key points in the two-dimensional space for the fish in the plurality of frames comprises: selecting, for each frame of the plurality of frames, a region of interest including a fish and at least one visual cue.

19. The non-transitory, computer-readable medium of claim 18, wherein generating the fish model comprises providing the region of interest including the fish and the at least one visual cue to a deep regression model as input; and receiving, as output from the deep regression model, a biomass estimate of the fish.

20. The non-transitory, computer-readable medium of any of claims 13 to 19, wherein identifying the first set of key points in the two-dimensional space for the fish in the plurality of frames comprises: identifying, in the at least one frame of the plurality of frames including the fish, a size of an eye of the fish in the at least one frame of the plurality of frames; measuring, the size of the eye of the fish in the at least one frame; and generating a calibration for the size of the fish based on the measured size of the eye of the fish.21 . The non-transitory, computer-readable medium of any of claims 13 to 20, wherein identifying the first set of key points in the two-dimensional space for the fish in the plurality of frames comprises: providing to a deep learning perception model, the plurality of frames; and receiving an output from the deep learning perception model comprising the first set of key points in the two-dimensional space for the fish in the plurality of frames.

22. The non-transitory, computer-readable medium of claim 21 , wherein generating the biomass estimate of the fish based on the second set of key points further comprises: generating, from output of the deep learning perception model, two-dimensional sizes of a plurality of trusses; providing as input to a trained neural network, the second set of two-dimensional key points and at least one supporting variable; and receiving, from the trained neural network, the biomass estimate of the fish.

23. The non-transitory, computer-readable medium of any of claims 13 to 22, wherein identifying the first set of key points in the two-dimensional space for the fish in the plurality of frames further comprises: identifying, in the plurality of frames, a reference object comprising a known dimension within a frame of the plurality of frames; obtaining the known dimension of the reference object; and generating a calibration for a size of the fish based on the known dimension of the reference object.

24. A computer-implemented system, comprising: one or more computers; and one or more computer memory devices interoperably coupled with the one or more computers and having tangible, non-transitory, machine-readable media storing one or more instructions that, when executed by the one or more computers, perform the methods of any of claims 1 to 11 .

25. A computer-implemented system, comprising: one or more computers; and one or more computer memory devices interoperably coupled with the one or more computers and having tangible, non-transitory, machine-readable media storingone or more instructions that, when executed by the one or more computers, perform one or more operations comprising: obtaining a plurality of frames captured by a monocular camera over a period of time, wherein at least one frame of the plurality of frames includes a fish; identifying a first set of key points in a two-dimensional space for the fish in the plurality of frames; receiving motion data for the monocular camera over the period of time; generating, from the first set of key points in the two-dimensional space and the motion data for the monocular camera, a fish model comprising a second set of key points in a three-dimensional space for the fish; generating a biomass estimate of the fish based on the second set of key points; and determining an action based on one or more biomass estimates including the biomass estimate of the fish.

26. The system of claim 25, further comprising: generating, from the motion data for the monocular camera over the period of time, extrinsic camera parameters for the monocular camera; and applying the extrinsic camera parameters to the plurality of frames captured by the monocular camera for the period of time.

27. The system of claim 26, wherein the extrinsic camera parameters comprise three-dimensional translation and three-dimensional rotation of the monocular camera during the period of time.

28. The system of any of claims 25 to 27, further comprising: obtaining, for the monocular camera, intrinsic camera parameters for the monocular camera for the period of time.

29. The system of any of claims 25 to 28, further comprising obtaining motion data for the fish, wherein generating the fish model comprising a second set of key points inthe three-dimensional space for the fish comprises: generating, from the motion data for the fish, a set of motion parameters for the fish including (i) a pose, (ii) translation, and (iii) rotation for the fish; and generating, from the set of motion parameters, the second set of key points in the three-dimensional space for the fish.

30. The system of any of claims 25 to 29, wherein identifying the first set of key points in the two-dimensional space for the fish in the plurality of frames comprises: selecting, for each frame of the plurality of frames, a region of interest including a fish and at least one visual cue.31 . The system of claim 30, wherein generating the fish model comprises providing the region of interest including the fish and the at least one visual cue to a deep regression model as input; and receiving, as output from the deep regression model, a biomass estimate of the fish.

32. The system of any of claims 25 to 31 , wherein identifying the first set of key points in the two-dimensional space for the fish in the plurality of frames comprises: identifying, in the at least one frame of the plurality of frames including the fish, a size of an eye of the fish in the at least one frame of the plurality of frames; measuring, the size of the eye of the fish in the at least one frame; and generating a calibration for the size of the fish based on the measured size of the eye of the fish.

33. The system of any of claims 25 to 32, wherein identifying the first set of key points in the two-dimensional space for the fish in the plurality of frames comprises: providing to a deep learning perception model, the plurality of frames; and receiving an output from the deep learning perception model comprising the first set of key points in the two-dimensional space for the fish in the plurality of frames.

34. The system of claim 33, wherein generating the biomass estimate of the fish based on the second set of key points further comprises: generating, from output of the deep learning perception model, two-dimensional sizes of a plurality of trusses; providing as input to a trained neural network, the second set of two-dimensional key points and at least one supporting variable; and receiving, from the trained neural network, the biomass estimate of the fish.

35. The system of any of claims 25 to 34, wherein identifying the first set of key points in the two-dimensional space for the fish in the plurality of frames further comprises: identifying, in the plurality of frames, a reference object comprising a known dimension within a frame of the plurality of frames; obtaining the known dimension of the reference object; and generating a calibration for a size of the fish based on the known dimension of the reference object.