Computer program, estimation method, and estimation device

The computer program uses machine learning to analyze fish images, identifying key features and projecting body height measurements onto segmented images, addressing inaccuracies in existing size estimation methods for aquatic organisms.

JP7811205B2Active Publication Date: 2026-02-04FURUNO ELECTRIC CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2023510653
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-03-31
Filing Date
2022-02-22
Publication Date
2026-02-04
Estimated Expiration
2042-02-22

AI Technical Summary

Technical Problem

Existing methods for accurately measuring the size of aquatic organisms, such as fish, are hindered by the lack of distinctive features for body height measurement, leading to inaccuracies and instability in size estimation.

Method used

A computer program that utilizes machine learning models to analyze underwater organism images, identifying specific parts like the caudal fin and snout tip, calculates body height based on perpendicular lines, and projects these measurements onto segmented images for precise size estimation.

Benefits of technology

Enables stable and accurate estimation of fish size by leveraging machine learning models to enhance measurement accuracy and biological stability, even when optimal features are absent.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007811205000001
    Figure 0007811205000001
  • Figure 0007811205000002
    Figure 0007811205000002
  • Figure 0007811205000003
    Figure 0007811205000003
Patent Text Reader

Abstract

[Problem] To provide a computer program, a model generation method, an estimation method and an estimation device with which the size of an aquatic organism can be estimated. [Solution] This computer program causes a computer to execute processing for: acquiring image data in which an aquatic organism is captured; inputting the acquired image data into a first learning model that outputs location data of a predetermined part of an aquatic organism if image data has been input, and acquiring location data of a predetermined part of the captured aquatic organism; inputting the acquired image data into a generation model that generates a segmentation image of an aquatic organism if image data has been input, and generating a segmentation image of the captured aquatic organism; and estimating the size of the captured aquatic organism on the basis of the location data of the predetermined part and the segmentation image.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a computer program , recommendation The present invention relates to a method and an apparatus for estimating a temperature. [Background technology]

[0002] The size of farmed fish in fish pens is important information for determining feeding amounts and harvesting times. Several analysis services have been offered to automate underwater fish length measurement. These services measure fish length by extracting multiple characteristic points on the fish's body and measuring the distance between these points (for example, between the tip of the mouth and the fork of the tail).

[0003] Patent Document 1 discloses an information processing device that detects two characteristic parts of a fish (the mouth and the caudal fin) from a photographed image of the fish and calculates the size of the fish based on the length between the detected characteristic parts. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] International Publication No. 2018 / 061925 Summary of the Invention [Problem to be solved by the invention]

[0005] However, the best part of a fish to measure its body height is not necessarily distinctive, making it difficult to measure its body height accurately in terms of measurement accuracy and biological stability, and making it difficult to estimate the size of aquatic organisms such as fish.

[0006] The present invention has been made in view of the above circumstances, and provides a computer program capable of estimating the size of aquatic organisms. , recommendation The present invention aims to provide a method and an apparatus for determining the [Means for solving the problem]

[0007] The computer program of the present invention causes a computer to execute the following processes: acquire image data of an underwater organism; input the acquired image data into a first learning model that outputs position data of a specific part of the underwater organism when image data is input, thereby acquiring position data of the specific part of the imaged underwater organism; input the acquired image data into a generation model that generates a segmented image of the underwater organism when image data is input, thereby generating a segmented image of the imaged underwater organism; and estimate the size of the imaged underwater organism based on the position data of the specific part and the segmentation image.

[0008] The computer program of the present invention causes a computer to execute a process of inputting acquired image data into a second learning model that outputs first position data of an aquatic organism and second position data of a specified part of the aquatic organism when image data is input, outputting the first position data and the second position data, inputting image data of a second region including the second position data into the first learning model, and inputting image data of a first region including the first position data into the generation model.

[0009] In the computer program according to the present invention, the predetermined part includes the caudal fin and the snout tip, and the computer is caused to execute a process of calculating body height based on a line perpendicular to a line connecting the caudal fin and the snout tip.

[0010] The computer program of the present invention causes a computer to execute the following process: input image data captured by a stereo camera into the first learning model; calculate the three-dimensional positions of the snout tip and the fork based on the acquired positional data of specified parts; generate a body height extension line based on the calculated three-dimensional positions; and two-dimensionally project the generated body height extension line onto the segmentation image to calculate body height.

[0011] The computer program according to the present invention causes a computer to execute a process of calculating the withers height using the withers height extension line located at a position that is a predetermined percentage of the distance between the snout tip and the fork from the position of the snout tip or the fork.

[0012] The computer program of the present invention causes a computer to execute the following process: input image data for each of multiple frames captured by a stereo camera into the first learning model, calculate the three-dimensional position of the fork for each frame based on the acquired position data of a specified part, identify the displacement of the fork for each frame based on the calculated three-dimensional position of the fork, and select a frame for estimating the size of the aquatic creature based on the identified displacement.

[0013] The computer program of the present invention causes a computer to execute a process of displaying an estimated position of the caudal fin or snout tip based on position data output by the first learning model, accepting a correction to the displayed estimated position of the caudal fin or snout tip, and re-learning the first learning model based on the accepted corrected position and image data at the time the corrected position was accepted.

[0014] The computer program of the present invention causes a computer to execute a process of displaying a segmentation image generated by the generative model, accepting modifications to the displayed segmentation image, and re-learning the generative model based on the modified segmentation image and the image data at the time the modifications were accepted.

[0015] The computer program according to the present invention causes a computer to execute a process of displaying an image of the aquatic organism being cultivated in a fish pen with its fork length or body height added thereto.

[0016] The computer program of the present invention causes a computer to execute a process of classifying the accuracy of the estimated sizes of multiple aquatic organisms cultivated in a fish tank into multiple ranks and displaying the number of aquatic organisms for each rank.

[0017] The computer program according to the present invention causes a computer to execute a process of displaying a size distribution including at least one of fork length and body height of a plurality of aquatic organisms cultivated in a fish pen.

[0018] The model generation method of the present invention acquires first training data including image data of an underwater organism and positional data of a specified part of the underwater organism, acquires second training data including the image data and a segmentation image of the underwater organism, generates a first learning model based on the first training data so as to output positional data of a specified part of the underwater organism when image data of the underwater organism is input, and generates a generative model based on the second training data so as to generate a segmentation image of the underwater organism when image data of the underwater organism is input.

[0019] The model generation method according to the present invention includes acquiring third training data including image data of an underwater organism, first position data of the underwater organism, and second position data of a predetermined part of the underwater organism, and based on the third training data, outputting the first position data of the underwater organism and the second position data of the predetermined part of the underwater organism when image data of the underwater organism is input. A second learning model is generated as follows:

[0020] The estimation method of the present invention acquires image data of an underwater organism, inputs the acquired image data into a first learning model that outputs position data of a specific part of the underwater organism when image data is input, acquires position data of a specific part of the imaged underwater organism, inputs the acquired image data into a generation model that generates a segmented image of the underwater organism when image data is input, generates a segmented image of the imaged underwater organism, and estimates the size of the imaged underwater organism based on the position data of the specific part and the segmentation image.

[0021] The estimation device of the present invention comprises a first acquisition unit that acquires image data of an underwater organism, a second acquisition unit that inputs the image data acquired by the first acquisition unit into a first learning model that outputs position data of a specific part of the underwater organism when image data is input, and acquires position data of the specific part of the imaged underwater organism, a generation unit that inputs the image data acquired by the first acquisition unit into a generation model that generates a segmentation image of the underwater organism when image data is input, and generates a segmentation image of the imaged underwater organism, and an estimation unit that estimates the size of the imaged underwater organism based on the position data of the specific part and the segmentation image. [Effects of the Invention]

[0022] According to the present invention, the size of aquatic organisms can be estimated. [Brief explanation of the drawings]

[0023] [Figure 1] 1 is a block diagram showing an example of the configuration of an estimation device according to an embodiment of the present invention. [Figure 2] FIG. 2 is a schematic diagram showing an example of the configuration of an AI unit. [Figure 3] FIG. 10 is a schematic diagram showing the process of estimating the size of a fish. [Figure 4] FIG. 2 is a schematic diagram showing an example of the configuration of a second learning model. [Figure 5] FIG. 2 is a schematic diagram showing an example of the configuration of a first learning model. [Figure 6] FIG. 1 is a schematic diagram illustrating an example of the configuration of a generative model. [Figure 7] FIG. 10 is a schematic diagram illustrating an example of a method for generating each model by machine learning. [Figure 8] FIG. 10 is a schematic diagram showing an example of fish frame pairing between cameras. [Figure 9] FIG. 10 is a diagram showing an example of a pair list of fish frames and body part frames. [Figure 10] FIG. 1 is a schematic diagram showing an example of pairing of fish and body parts. [Figure 11] FIG. 10 is a diagram illustrating an example of a fish list. [Figure 12] FIG. 2 is a block diagram showing an example of the configuration of a tracking unit. [Figure 13] FIG. 2 is a block diagram showing an example of the configuration of a tail beat removal unit. [Figure 14] FIG. 10 is a schematic diagram showing an example of an approximate straight line of a forked track. [Figure 15] FIG. 2 is a block diagram showing an example of the configuration of a measurement unit. [Figure 16] FIG. 1 is a schematic diagram showing an example of body height measurement. [Figure 17] FIG. 1 is a schematic diagram showing an example of automatic measurement of fish size. [Figure 18] FIG. 10 is a schematic diagram showing an example of the configuration of an estimation result DB. [Figure 19] FIG. 10 is a block diagram showing another example of the configuration of the estimation device according to the present embodiment. [Figure 20] FIG. 10 is a schematic diagram showing a first display example of the estimation result. [Figure 21] FIG. 10 is a schematic diagram illustrating an example of a cause of error. [Figure 22] FIG. 10 is a schematic diagram showing a second display example of the estimation result. [Figure 23] FIG. 10 is a schematic diagram showing a third display example of the estimation result. [Figure 24] FIG. 10 is a schematic diagram showing a fourth display example of the estimation result. [Figure 25] 10 is a flowchart illustrating an example of a processing procedure of the estimation device. [Figure 26] 10 is a flowchart illustrating an example of a processing procedure of the estimation device. DETAILED DESCRIPTION OF THE INVENTION

[0024] An embodiment of the present invention will now be described. Fig. 1 is a block diagram showing an example of the configuration of an estimation device 100 according to this embodiment. The estimation device 100 includes an input unit 10, an AI unit 20, a pairing unit 30, a tracking unit 40, a tail beat removal unit 50, a measurement unit 60, an estimation unit 70 as an estimation unit, and an output unit 80.

[0025] The camera unit 200 is, for example, a waterproof stereo camera installed at a predetermined position underwater in the fish tank. The camera unit 200 can capture images of aquatic creatures such as fish swimming in the fish tank. The camera unit 200 may be a monocular camera. Other imaging devices may be used instead of a camera. In the following, fish will be used as an example of aquatic creatures. The camera unit 200 outputs captured image data (video). The captured image data is made up of a plurality of frame-by-frame images (frame images).

[0026] The estimation device 100 may be a personal computer or the like, or may be a server (cloud) on the Internet. When the estimation device 100 is a personal computer or the like, the estimation device 100 may acquire captured image data directly from the camera unit 200. When the estimation device 100 is a server on the Internet, the estimation device 100 can acquire captured image data captured by the camera unit 200 via a client terminal device (not shown) or the like.

[0027] The input unit 10 acquires captured image data captured by the camera unit 200. The input unit 10 outputs the acquired captured image data to the AI ​​unit 20.

[0028] The AI ​​unit 20 is a model generated by machine learning, and outputs data necessary for estimating the size of a fish based on captured image data. The size of the fish is, for example, fork length and body height. The AI ​​unit 20 will be described in detail below. Fork length is the body length, and more specifically, it is the length from the tip of the fish's upper jaw (snout) to the most concave part (fork) in the center of the forked caudal fin. Body height is the vertical distance from the dorsal edge to the ventral edge of the fish, and more specifically, it is the distance from the base of the pelvic fin to the dorsal edge.

[0029] 2 is a schematic diagram showing an example of the configuration of the AI ​​unit 20. The AI ​​unit 20 includes a first learning model 21 as a second acquisition unit, a generative model 22 as a generation unit, a second learning model 23, and an image cropping unit 24 as a first acquisition unit. The estimation unit 70 can estimate the size of the fish based on the position data of the body part output by the first learning model 21 and the fish segmentation image data output by the generative model 22. In this case, the camera unit 200 may be a monocular camera.

[0030] The second learning model 23 is generated by performing machine learning so that, when captured image data is input, it outputs position data of a single fish (first position data of an aquatic organism) and position data of a part (second position data of a predetermined part of the aquatic organism). The position data output by the second learning model 23 is coordinates (x, y) or area (x, y, width, height) on the captured image. do.

[0031] The image cropping unit 24 crops out an image of the surrounding area including the coordinates or area based on the position data (coordinates or area) of the individual fish output by the second learning model 23, and acquires individual fish image data. The individual fish image includes an image of the entire fish. The image cropping unit 24 crops out an image of the surrounding area including the coordinates or area based on the position data (coordinates or area) of the part output by the second learning model 23, and acquires part image data. The part image includes an image of a specific part of the fish. The image cropping unit 24 outputs the individual fish image data to the generation model 22 and outputs the part image data to the first learning model 21. The image cropping unit 24 may be provided outside the AI ​​unit 20.

[0032] The first learning model 21 is generated by performing machine learning so as to output position data of a part when inputting the part image data output by the image cropping unit 24. The generation model 22 is generated by performing machine learning so as to output segmentation image data of a fish when inputting the individual fish image data output by the image cropping unit 24.

[0033] FIG. 3 is a schematic diagram showing the flow of estimating fish size. For convenience, in FIG. 3, it is assumed that a single fish is captured in the captured image. As described above, the second learning model 23, in cooperation with the image cropping unit 24, outputs individual fish image data (image data of a first region including first position data) and partial image data (image data of a second region including second position data) when captured image data is input. As shown in FIG. 3, the individual fish image is an image of a rectangular region surrounding the entire individual fish. The individual fish image data is data including rectangular region data and a class (e.g., fish) within the rectangular region. The partial image is a snout tip image including a snout tip P1 as a predetermined region, and a caudal fin image including a caudal fin (or a fork) P2 as a predetermined region. The partial image data is data including rectangular region data and a class (e.g., snout tip and caudal fin) within the rectangular region.

[0034] When the first learning model 21 receives image data of a part, it outputs position data of the part. The position data of the part is position data of the snout tip P1 and the caudal fin P2, and is expressed as coordinates on the image. If the coordinate system of the image is XY, the position data can be expressed by coordinates (x, y) or area (x, y, width, height). In this case, the position data in the direction of the optical axis of the camera The coordinate can be expressed as Z.

[0035] When image data of a single fish is input, the generative model 22 outputs segmentation image data. The segmentation image is a classification of the class of each pixel of the image of a single fish. In the example of FIG. 3, pixels classified as "fish body excluding fins" are shown as "no pattern," and pixels classified as "non-fish body" are shown as "patterned." The segmentation image makes it possible to determine the boundary between the fish body and parts other than the fish body (the outline of the fish body).

[0036] The estimation unit 70 can estimate the size of the fish based on the position data of the body parts and the segmentation image data. As shown in Figure 3, the length of a straight line L connecting the snout tip P1 and the caudal fin P2 can be estimated as the fork length. In addition, the length between the intersection of a straight line H perpendicular to the line L and the outline of the fish body can be estimated as the body height. Here, the position of the intersection of the line H and the line L can be set to a point on the line L that is a predetermined percentage of the position of the snout tip P1 (for example, 40% of the fork length).

[0037] As described above, by combining the position data of specific parts (snout tip and caudal fin) with the segmentation image, body height measurement can be performed based on the body height line H that is perpendicular to the line representing the fork length and the segmentation image.Therefore, even if there is no characteristic part that is optimal for body height measurement, the size of the fish can be estimated stably and with high accuracy.

[0038] By carrying out the processing from the pairing unit 30 onwards, which will be described later, it is possible to measure the body height with even higher accuracy.

[0039] 4 is a schematic diagram showing an example of the configuration of the second learning model 23. The second learning model 23 can be configured, for example, by a Faster R-CNN (Convolutional Neural Network). The second learning model 23 includes a CNN layer 231, a region proposal network (RPN) 232, a ROI (Region-of-Interest) POOL 233, and a classification network 234. The CNN layer 231 generates image features (feature maps) from input captured image data and outputs them to the region proposal network (RPN) 232 and the ROI POOL 233.

[0040] The region proposal network (RPN) 232 calculates candidate regions from the input image features and outputs them to the ROI POOL 233. The region proposal network (RPN) 232 can detect where an object is captured in the captured image, that is, the region where the object is captured and its rectangular shape.

[0041] The ROI POOL 233 connects the image features output from the CNN layer 231 with the candidate regions output from the region proposal network (RPN) 232, and outputs fixed-length ROI region features to the classification network 234.

[0042] The classification network 234 recalculates the correct region and class from the input ROI region features, classifies the region, and outputs the class. Here, the classes are single fish, snout tip, and caudal fin. The second learning model 23 is not limited to Faster R-CNN, but may be, for example, R-CNN, Mask R-CNN, YOLO (You Only Look Once), SSD (Single Shot Multibox Detector), etc.

[0043] FIG. 5 is a schematic diagram showing an example of the configuration of the first learning model 21. The first learning model 21 can be configured, for example, using RetinaNet. The first learning model 21 includes a Feature Pyramid Network 211, a classifier 212, and a region regressor 213. The regional image data output by the second learning model 23 is input to the first learning model 21. The regional image data includes image data of the caudal fin and image data of the snout tip. The Feature Pyramid Network 211 calculates a feature hierarchy constituting feature maps at various scales in a bottom-up direction, and generates high-resolution features in a top-down direction by upsampling feature maps from higher layers that are spatially coarse but semantically strong. Each feature is associated with a feature calculated in the bottom-up direction.

[0044] In this way, the Feature Pyramid Network211 can extract features with both high-order and low-order characteristics, achieving a well-balanced accuracy in both semantics and location.

[0045] The area regressor 213 outputs the position of the snout tip and the position of the caudal fin. The positions of the snout tip and the caudal fin are represented by coordinate values ​​on the image. The classifier 212 identifies the class (snout tip and caudal fin) of the position (coordinate values) output by the area regressor 213. The first learning model 21 can accurately determine the position of the snout tip on the image of the snout tip, and can accurately determine the position of the caudal fin on the image of the caudal fin. The first learning model 21 is not limited to RetinaNet, and may be, for example, SegNet, Mask R-CNN, SVM (Support Vector Machine), etc.

[0046] FIG. 6 is a schematic diagram showing an example of the configuration of the generative model 22. The generative model 22 can be configured, for example, with a U-Net. The generative model 22 includes encoders 221 to 225 and decoders 226 to 229. The generative model 22 repeatedly performs convolution processing on input image data of individual fish using the encoders 221 to 225. The decoders 226 to 229 repeatedly perform upsampling (deconvolution) processing on the image convolved by the encoder 225. When decoding the convolved image, the feature maps generated by the encoders 224 to 221 are added to the image to be deconvolved. The generative model 22 outputs a segmentation image of the fish body in the input image of the individual fish. The segmentation image makes it possible to determine the boundary between the fish body and parts other than the fish body (the outline of the fish body).

[0047] This allows us to retain the positional information that would otherwise be lost during the convolution process, resulting in more accurate segmentation (which pixels belong to which class). The generative model 22 is not limited to U-Net, and may be, for example, a Generative Adversarial Network (GAN), SegNet, or the like.

[0048] FIG. 7 is a schematic diagram showing an example of a method for generating each model by machine learning. FIG. 7A shows a method for generating the second learning model 23. The second learning model 23 can be generated by performing machine learning so that, when captured image data is input as learning input data, it outputs position data of individual fish (first position data) and position data of body parts (second position data). In this case, the training data uses the position data of individual fish and position data of body parts created by performing annotation based on the captured image data, which is the learning input data. A large amount of training data and training data is prepared, and the second learning model 23 is trained by machine learning. The second learning model 23 can be generated by updating the internal parameters of the second learning model 23 so that the output data approaches the training data.

[0049] FIG. 7B shows a method for generating the first learning model 21. The first learning model 21 can be generated by machine learning so that when part image data is input as learning input data, part position data is output. In this case, the training data uses part position data created by annotating the part image data that is the learning input data. A large amount of training data and training data is prepared, and the first learning model 21 is trained by machine learning. The first learning model 21 can be generated by updating the internal parameters of the first learning model 21 so that the output data approaches the training data.

[0050] FIG. 7C shows a method for generating the generative model 22. The generative model 22 can be generated by performing machine learning to output segmentation image data when image data of individual fishes is input as learning input data. In this case, the training data is segmentation image data created by performing annotation based on the image data of individual fishes, which is the learning input data. A large amount of training data and training data is prepared, and the generative model 22 is trained by machine learning. The internal parameters of the generative model 22 can be updated to generate the generative model 22 so that the output data approaches the training data.

[0051] Each of the above-mentioned models can also be retrained. For example, as shown in FIG. 3, the estimated positions of the caudal fin and snout tip are displayed on the screen based on the position data output by the first learning model 21, and an operation to correct the displayed estimated positions of the caudal fin and snout tip is accepted. The estimated positions can be corrected, for example, by operating a mouse or the like to move a pointer on the screen to the estimated position and performing a predetermined operation (touch, click, drag-and-drop, etc.), thereby moving the estimated position to the correct position on the screen. The accepted corrected position and image data when the corrected position was accepted are prepared as training data, and the first learning model 21 can be retrained based on the accepted corrected position and the image data when the corrected position was accepted. This improves the accuracy of the position data output by the first learning model 21.

[0052] In addition, as shown in FIG. 3, the segmentation image generated by the generative model 22 is displayed on the screen and an operation to modify the displayed segmentation image is accepted. The segmentation image can be modified, for example, by launching an application such as paint software, displaying the segmentation image on the screen, operating a mouse or the like to move a pointer on the screen to a desired position, and performing a predetermined operation (touch, click, drag-and-drop, etc.) to convert pixels other than those of the fish body into pixels of the fish body, or converting pixels of the fish body into pixels other than those of the fish body. The modified segmentation image and the image data when the modification was accepted are prepared as training data, and the generative model 22 can be retrained based on the modified segmentation image and the image data when the modification was accepted. This improves the accuracy of the segmentation image generated by the generative model 22.

[0053] In addition, the position data of individual fish, snout tip position data, and caudal fin position data output by the second learning model 23 are displayed on the screen, and an operation to correct the displayed position data is accepted. In this case, from the viewpoint of improving workability and visibility, cropped images (individual fish image, snout tip image, caudal fin image) containing the position of the individual fish, the position of the snout tip, and the position of the caudal fin, respectively, may be displayed on the screen along with the positions. When correcting the position data of an individual fish, the position of the individual fish output by the second learning model 23 is displayed on the screen, and an operation to correct the displayed position of the individual fish is accepted. The position of the individual fish can be corrected, for example, by operating a mouse or the like to move the position of the individual fish on the screen and correcting the position. When correcting the position of the snout tip, the position of the snout tip output by the second learning model 23 is displayed on the screen, and an operation to correct the displayed position of the snout tip is accepted. The position of the snout tip can be corrected, for example, by operating a mouse or the like to move the position of the snout tip on the screen and correcting the position. When correcting the position of the caudal fin, the position of the caudal fin output by the second learning model 23 is displayed on the screen, and an operation to correct the displayed position of the caudal fin is accepted. The position of the caudal fin can be corrected, for example, by operating a mouse or the like to move the position of the caudal fin on the screen. In addition, the corrected position data for re-training the first learning model 21 can also be used as the position data of the snout tip and the caudal fin.

[0054] The corrected position of the individual fish and the image data when the correction was accepted are prepared as training data, and the second learning model 23 can be retrained based on the corrected position of the individual fish and the image data when the correction was accepted. The corrected position of the snout tip and the image data when the correction was accepted are prepared as training data, and the second learning model 23 can be retrained based on the corrected position of the snout tip and the image data when the correction was accepted. Similarly, the corrected position of the caudal fin and the image data when the correction was accepted are prepared as training data, and the second learning model 23 can be retrained based on the corrected position of the caudal fin and the image data when the correction was accepted. This improves the accuracy of the position data of the individual fish and the position data of the body parts output by the second learning model 23.

[0055] Next, a method for further improving the accuracy of body height measurement will be described. Below, the pairing process, tracking process, tail beat removal process, and body height measurement process will be described in order. It is assumed that the camera unit 200 is a stereo camera.

[0056] The pairing unit 30 performs pairing processing between cameras and pairing processing between fish and body parts. First, the pairing processing between cameras will be described. The pairing processing between cameras includes fish frame pairing and body part frame pairing.

[0057] FIG. 8 is a schematic diagram showing an example of fish frame pairing between cameras. In FIG. 8, the left-side identification image is an image showing an individual fish image estimated based on an image captured by the left camera of the stereo camera, and the right-side identification image is an image showing an individual fish image (also referred to as a fish frame) estimated based on an image captured by the right camera of the stereo camera. In the example of FIG. 8, three fish frames are shown, but the number of fish frames in an actual image is not limited to three. Instead of arranging the stereo cameras in the left-right direction (horizontal direction), they may be arranged in the up-down direction or in an oblique direction (for example, at an angle greater than 0 degrees and less than 90 degrees with respect to the horizontal direction).

[0058] The pairing unit 30 performs a round-robin pairing for all fish frames in the left and right identification images. In the example of FIG. 8, pairing is performed between the fish frame G1 in the left identification image and all fish frames G4, G5, and G6 in the right identification image. When pairing, the closest pair within the allowable error range is selected. In the example of FIG. 8, the fish frame G6 is selected as the closest pair for the fish frame G1. Similarly, for the fish frames G2 and G3 in the left identification image, corresponding pairs are selected from the fish frames in the right identification image.

[0059] Although not shown, body part frame pairing can be performed in the same way as fish frame pairing. The pairing unit 30 performs a round-robin pairing for all body part frames in the left and right identification images. Pairing is performed between the body part frame in the left identification image and all body part frames in the right identification image. When pairing, the closest pair within the allowable error range is selected.

[0060] Figure 9 shows an example of a pair list of fish frames and body part frames. Figure 9A shows the fish frame pair list, and Figure 9B shows the body part frame pair list. As shown in Figure 9A, fish frame G1 in the left identification image and fish frame G6 in the right identification image are paired. Similarly, fish frame G2 in the left identification image and fish frame G4 in the right identification image are paired, and fish frame G3 in the left identification image and fish frame G5 in the right identification image are paired.

[0061] 9B, body part frame g1 in the left identification image is paired with body part frame g12 in the right identification image, body part frame g2 in the left identification image is paired with body part frame g10 in the right identification image, and body part frame g3 in the left identification image is paired with body part frame g8 in the right identification image, and the same applies to the other body parts.

[0062] FIG. 10 is a schematic diagram showing an example of pairing of fish and body parts. Assume that fish frames G1 to G3 and body part frames g1 to g6 are paired through the pairing between cameras described above. Note that fish frames G1 to G3 and body part frames g1 to g6 each have a paired fish frame and body part frame. FIG. 10 illustrates one side (for example, the left identification image) of the paired fish frame and body part frame.

[0063] The pairing unit 30 performs a round-robin pairing for all fish frames and body part frames in the identification image. In this case, pairing can be performed using only the left identification image, only the right identification image, or both the left and right identification images. In the example of FIG. 10, pairing is performed between the fish frame G1 and all body part frames g1 to g6 in the identification image. When pairing, the closest pair within the allowable error range is selected. The same process is performed for the other fish frames G2 and G3.

[0064] Figure 11 is a diagram showing an example of a fish list. As shown in Figure 11, fish frame G1 is paired with body part frames g3 and g4. Similarly, fish frame G2 is paired with body part frames g1 and g2, and fish frame G3 is paired with body part frames g5 and g6. That is, for a fish with an ID of 1, fish frame G1 is paired with body part frames g3 and g4, for a fish with an ID of 2, fish frame G2 is paired with body part frames g1 and g2, and for a fish with an ID of 3, fish frame G3 is paired with body part frames g5 and g6.

[0065] With the above-described configuration, for example, it is possible to associate the fish frame and body part frame of each of a plurality of fish obtained by capturing images of the inside of a fish pen.

[0066] 12 is a block diagram showing an example of the configuration of the tracking unit 40. The tracking unit 40 tracks fish swimming in the fish tank between frames. By tracking, double counting of fish can be prevented. The tracking unit 40 includes a 2D tracking unit 41, a 3D conversion unit 42, and a 3D tracking unit 43.

[0067] The 2D tracking unit 41 tracks each fish using the fish frame pair list, body part frame pair list, and fish list generated by the pairing unit 30. The 2D tracking unit 41 tracks the fish based on a 2D image of the fish.

[0068] The 3D conversion unit 42 converts the 2D image of the fish into a 3D image based on the fish frame pair list, body part frame pair list, and fish list generated by the pairing unit 30, using the single fish image captured by the left camera, the single fish image captured by the right camera, and the distance measurement principle of the stereo camera. Alternatively, the 3D positions of the snout and fork can be measured based on the principle of triangulation. The distance Z to each pixel in the 2D image of the fish can be calculated as Z = (B × F) / D, where B is the distance between the cameras, F is the focal length, and D is the parallax.

[0069] The 3D tracking unit 43 tracks each fish using the fish frame pair list, body part frame pair list, and fish list generated by the pairing unit 30. The 3D tracking unit 43 can track each fish based on the position of the snout tip on the 3D image of the fish. This can reduce the influence of the fish's tail beat when it swims in the water.

[0070] 13 is a block diagram showing an example of the configuration of the tail beat removal unit 50. The tail beat removal unit 50 includes a fork 3D measurement unit 51, a plane projection unit 52, an approximate line calculation unit 53, and a frame removal unit .

[0071] The tail fork 3D measurement unit 51 measures the 3D position (X, Y, Z) from the (X, Y) coordinates on the image of the tail fork (caudal fin) that has been paired by the pairing unit 30. Here, the horizontal direction of the image can be defined as the X axis, the vertical direction as the Y axis, and the direction of the camera's optical axis as the Z axis.

[0072] The plane projection unit 52 projects the 3D position of the fork onto the XZ plane. This allows the position of the fork to be as seen from the back of the fish. The plane projection unit 52 records the position of the fork as seen from the back of the fish for each frame. This allows the position of the fork to be plotted on the XZ plane.

[0073] The approximate line calculation unit 53 calculates an approximate line of the position of the tail fork (tail fork trajectory) plotted on the XZ plane.

[0074] If the deviation of the tail fork from the approximate line exceeds the allowable range, the frame removal unit 54 removes the frame corresponding to the tail fork. Instead of calculating the approximate line after projecting the 3D position of the tail fork onto a plane, an approximate line in 3D space may be calculated from the 3D position of the tail fork. In this case, if the deviation of the 3D position of the tail fork from the calculated 3D approximate line exceeds the allowable range, the frame corresponding to the tail fork can be removed.

[0075] FIG. 14 is a schematic diagram showing an example of an approximate line of the fork trajectory. In FIG. 14, the horizontal axis represents the fork position X, and the vertical axis represents the fork position Z. In FIG. 14, O represents the measured fork position, symbols t1 to t13 represent frame time points, and the dashed line represents the approximate line of the fork trajectory. The fork position changes from time t1 to t13. In FIG. 14, the fork positions at times t3 and t7 to t9 deviate from the approximate line. Therefore, frames t3 and t7 to t9 are removed, and the remaining frames t1 to t2, t4 to t6, and t10 to t13 are used to measure the body height, as described below. The fork position deviates from the approximate line because, for example, the fish changes course, causing the fork to fluctuate significantly. This can cause errors when estimating the fish's size. Therefore, if the deviation of the fork from the approximate line exceeds the allowable range, the frame corresponding to that fork is removed. Note that frame removal is not limited to the method using an approximated line, and other machine learning methods such as SVM, neural network, clustering, etc. Also, if the amount of change in the position of the fork over time exceeds a predetermined threshold, the frame may be removed.

[0076] The tail beat removal unit 50 outputs the frames other than the frames from which the tail beats have been removed from the captured frames to the measurement unit 60 as tail beat removed frames (frames to be measured).

[0077] That is, the estimation device 100 calculates a plurality of frames captured by the stereo camera. Each image data is input into the first learning model 21, and the three-dimensional position of the fork for each frame is calculated based on the acquired position data of the specified part.The calculated three-dimensional position is projected onto a specified two-dimensional plane to identify the displacement of the fork for each frame, and a frame for estimating the size of the fish is selected based on the identified displacement.

[0078] 15 is a block diagram showing an example of the configuration of the measurement unit 60. The measurement unit 60 includes an optimal frame selection unit 61, a snout tip / tail fork 3D measurement unit 62, a body height auxiliary line generation unit 63, a body height auxiliary line projection unit 64, and a body height measurement unit 65.

[0079] The optimal frame selection unit 61 selects the frame at the time corresponding to the tail fork that is closest to the approximate straight line of the tail fork trajectory. Note that instead of selecting the optimal frame, it may be selected randomly from among the frames remaining after the frame is removed, or it may be selected a frame several frames before or after the removed frame.

[0080] The snout / fork 3D measurement unit 62 measures the 3D position of the snout and the 3D position of the fork in the selected frame. Note that the 3D measurement may use the 3D position of the snout tracked by the 3D tracking unit 43 and the 3D position of the fork measured by the fork 3D measurement unit 51.

[0081] The body height auxiliary line generating unit 63 generates nine body height auxiliary lines that are perpendicular to the 3D line connecting the 3D measured 3D position of the snout tip and the 3D position of the fork, so as to divide the 3D line into, for example, ten parts.

[0082] The body height extension line projection unit 64 projects the body height extension line generated by the body height extension line generating unit 63 onto a 2D plane (for example, an XY plane).

[0083] The body height measurement unit 65 measures the body height based on the intersection position between the body height auxiliary line projected onto a 2D plane and the fish body contour (for example, the contour on the segmentation image). If the body height cannot be measured due to a defect in the fish body contour, etc., the optimal frame selection unit 61 selects the frame at the time corresponding to the fork that is next closest to the approximate line of the fork trajectory, and repeats the processes of the snout tip / fork 3D measurement unit 62, body height auxiliary line generation unit 63, body height auxiliary line projection unit 64, and body height measurement unit 65. A defect in the fish body contour, etc., can be determined, for example, by comparing the distance between the intersection points of the body height auxiliary line and the fish body contour with a predetermined threshold. Alternatively, it can be determined by comparing the amount of change in the distance between the intersection points of each adjacent body height auxiliary line and the fish body contour with a predetermined threshold. The missing portion may be interpolated.

[0084] Fig. 16 is a schematic diagram showing an example of body height measurement. As shown in Fig. 16A, nine body height auxiliary lines A1 to A9 perpendicular to the 3D line L are generated so as to divide the 3D line L connecting the 3D position of the snout tip and the 3D position of the fork, which have been 3D measured, into, for example, 10 parts.

[0085] As shown in FIG. 16B, nine body height auxiliary lines A1 to A9 are projected onto a 2D plane (for example, an XY plane), and the body height is measured based on the positions of intersections between the body height auxiliary lines A1 to A9 projected onto the 2D plane and the outline of the fish body (for example, the outline on a segmentation image). Specifically, the body height can be measured as the distance H between two positions where the body height auxiliary line A4, which is at a distance equivalent to 40% of the fork length L from the tip of the snout, intersects with the outline of the fish body. Note that the number of body height auxiliary lines is just an example and is not limited to nine. Also, the ratio of 40% is just an example and is not limited to 40%.

[0086] FIG. 17 is a schematic diagram showing an example of automatic measurement of fish size. In FIG. 17, the ID of the fish shown in each image of each frame number is shown. For example, a fish with a fish ID of 0001 is shown in frames 1 to 10, and the size of fish 0001 is automatically measured in frame 7 indicated by O. Frame 7, which is measured by automatic measurement, is the largest of frames 1 to 10. For example, this is the frame with the smallest tail beat.

[0087] Similarly, a fish with fish ID 0002 is captured in frames 3 to 12, and the size of fish 0002 is automatically measured in frame 9 indicated by O. A fish with fish ID 0003 is captured in frames 5 to 9, but the size of fish 0003 is not automatically measured. Furthermore, a fish with fish ID 0004 is captured in frames 7 to 14, and the size of fish 0004 is automatically measured in frame 13 indicated by O. A fish with fish ID 0005 is captured in frames 8 to 16, and the size of fish 0005 is automatically measured in frame 15 indicated by O. The tracking unit 40 recognizes the same individual fish between frames.

[0088] The measurement unit 60 outputs the measurement results to the estimation unit 70. The estimation unit 70 collects the automatically measured fork lengths and body heights of the fish in the fish pen, and can estimate the size of the fish in the fish pen.

[0089] The output unit 80 can convert the result estimated by the estimation unit 70 into displayable data and output it to an external terminal device, display device, or the like.

[0090] FIG. 18 is a schematic diagram showing an example of the configuration of the estimation result DB 90. The estimation result DB 90 may be provided inside the estimation device 100, or may be provided in an external data server or the like as long as it is accessible from the estimation device 100. The estimation result DB 90 can store the estimation results obtained by the estimation device 100. The estimation result DB 90 stores, for example, a fish ID, fork length, body height, rank, and image in association with the fish ID. Although not shown, position information (coordinate values) of the snout tip and position information (coordinate values) of the caudal fin (fork) may also be stored in association with the fish ID. The fish ID is, for example, an identifier for identifying a fish in a fish pen. The rank specifies the measurement accuracy of the fish size. Details of the rank will be described later. The image is an image of a fish, and a line indicating the fork length and a line indicating the body height of the fish may be displayed superimposed on the image of the fish. In addition, a frame number, a value of the fork length, a value of the body height, and the rank to which the fish belongs may be displayed on the image of the fish.

[0091] FIG. 19 is a block diagram showing another example of the configuration of the estimation device 100 according to this embodiment. As shown in FIG. 19, the estimation device 100 may be, for example, a personal computer. The estimation device 100 may be configured with a CPU 101, a ROM 102, a RAM 103, a GPU 104, a video memory 105, a recording medium reading unit 106, and the like. A computer program (computer program product) recorded on a recording medium 1 (e.g., an optically readable disk storage medium such as a CD-ROM) can be read by the recording medium reading unit 106 (e.g., an optical disk drive) and stored in the RAM 103. Here, the computer program (computer program product) includes the processing procedures described in FIGS. 25 and 26 (described later). The computer program may also be stored on a hard disk (not shown) and then stored in the RAM 103 when the computer program is executed.

[0092] By causing the CPU 101 to execute a computer program (computer program product) stored in the RAM 103, it is possible to execute each process in the input unit 10, the AI ​​unit 20, the pairing unit 30, the tracking unit 40, the tail beat removal unit 50, the measurement unit 60, the estimation unit 70, and the output unit 80. The video memory 105 can temporarily store data for various image processes, processing results, and the like. Furthermore, instead of being configured to be read by the recording medium reading unit 106, the computer program (computer program product) can also be downloaded from another computer or network device via a network such as the Internet.

[0093] Next, we will explain the estimation results displayed on the display screen of a terminal device connected to the estimation device 100 via a communication network. The following display process is performed in response to a request from the terminal device. This can be done by the estimation device 100 accessing the estimation result DB 90 and, if necessary, calculating the values ​​to be displayed and outputting them to the terminal device. Also, if the terminal device has already downloaded the data from the estimation result DB, the CPU of the terminal device can display the downloaded data on the display unit.

[0094] 20 is a schematic diagram showing a first display example of the estimation results. The estimation result screen 300 displays columns such as rank, number of measurements, average fish weight, average fork length, average body height, average body fatness, average shooting distance, and average angle. In addition, a box 301 is displayed for each rank, allowing users to select whether to download detailed data. By operating the "save" icon 302, detailed data can be downloaded from the estimation device 100. The data to be downloaded includes videos captured by the camera unit 200.

[0095] The rank specifies the accuracy of measuring fish size, and for example, rank A may indicate an error of less than 5%, rank B an error of 5% or more but less than 10%, rank C an error of 10% or more but less than 20%, and rank F an error of 20% or more. Note that the number of rank categories and the definition of the ranks are not limited to these.

[0096] Figure 21 is a schematic diagram showing an example of an error factor. Figure 21A shows the distribution of errors in the position of body parts. As the errors in the position of the snout tip and the position of the tail fin increase, the errors in the fork length and body height also increase. The errors in the position of the snout tip and the tail fin are calculated based on the reprojection error when performing three-dimensional measurement using the position data of the snout tip and the tail fin (fork) output by the first learning model 21. Figure 21B shows the error in the tilt (angle) of the fish body. The greater the tilt of the fish body, the greater the error in the body height. The tilt error can be reflected in the body height error as a simple measurement position error, resulting from the angular deviation. For ranking, the worse of the error in the position of the body parts or the tilt error can be used.

[0097] The number of measurements indicates the number of fish included in each rank. The average fork length and average withers height represent the average values ​​within each rank. The average fish weight and average body fatness can be calculated from the average fork length and average withers height using a predetermined formula. In aquaculture, it is also important to understand the weight of reared fish. Measuring the weight of reared fish using an actual weighing scale requires a lot of effort. However, according to this embodiment, the weight of reared fish can be converted from the fork length and withers height estimated by the estimation device 100 without using an actual weighing scale. In this case, even if there is no characteristic part that is optimal for measuring the withers height, the withers height can be estimated stably and with high accuracy, and the weight of reared fish can also be estimated accurately.

[0098] As described above, the estimation device 100 classifies the accuracy of the estimated sizes of multiple fish (aquatic organisms) cultivated in the fish pen into multiple ranks and displays the fish for each rank. This allows the user to know the size of the fish in the pen and the reliability of the data, providing information for determining the amount of feeding and the timing of landing.

[0099] Figure 22 is a schematic diagram showing a second display example of the estimation results. A distribution selection list 311 and a distribution display area 312 are displayed on the estimation result screen 310. In the example of Figure 22, the fork length distribution is selected from items such as fork length, body height, and fish weight, and the fork length distribution of the fish in the fish pen is displayed in the distribution display area 312. Ranks may also be displayed as a breakdown of the fork length distribution. As with fork length, the distribution of body height, fish weight, etc. can be displayed.

[0100] As described above, the estimation device 100 can display the distribution of at least one of the fork length and body height of multiple fish being cultured in a fish pen. This allows the size distribution of the fish in the pen to be known, providing information for determining the amount of feeding and the timing of landing.

[0101] FIG. 23 is a schematic diagram showing a third display example of the estimation results. The estimation result screen 320 displays a rank selection list 321 and a fish measurement data display area 322. In the example of FIG. 23, rank A is selected from ranks A, B, C, and F. The fish measurement data display area 322 displays the measurement data of each fish belonging to rank A. In the example of FIG. 23, the measurement data includes, but is not limited to, fish weight, fork length, body height, and an image. By operating the box 323 in the image column, an image of the selected fish can be displayed, as shown in FIG. 24, which will be described later.

[0102] This allows the measurement data of each individual fish (individuals) belonging to each rank to be checked.

[0103] 24 is a schematic diagram showing a fourth display example of the estimation results. An image of the selected fish is displayed on the estimation result screen 330. A line indicating the fork length of the fish and a line indicating its height may also be displayed superimposed on the image of the fish. The frame number, fork length value, height value, fish weight value, and rank may also be displayed on the image of the fish.

[0104] As described above, the estimation device 100 can display an image of a fish with its fork length and body height attached, allowing the measurement data of each individual fish (individual) belonging to a rank to be confirmed.

[0105] 25 and 26 are flowcharts showing an example of a processing procedure of the estimation device 100. The estimation device 100 acquires captured image data (S11), inputs the acquired captured image data to the second learning model 23, and acquires individual fish image data and partial image data from the image cropping unit 24 (S12). Here, the partial image data includes snout tip image data and caudal fin image data. The estimation device 100 inputs the acquired partial image data to the first learning model 21 to acquire partial position data (S13). Here, the snout tip image data is input to the first learning model 21 to acquire snout tip position information (coordinates or area values) output by the first learning model 21, and the caudal fin image data is input to the first learning model 21 to acquire caudal fin position information (coordinates or area values) output by the first learning model 21.

[0106] The estimation device 100 inputs the acquired image data of a single fish into the generation model 22 to acquire segmentation image data (S14), and estimates (generates) measurement points for the fork length based on the acquired position data of the body part and the segmentation image data (S15). Here, the measurement points for the fork length (position information (coordinates or area values) of the snout tip and caudal fin) are stored in the estimation result DB 90. Note that the size of the fish (fork length and body height) may be estimated by determining the distance Z from the camera to the fish or a predetermined part of the fish using the processing of steps S16 and S17 described below.

[0107] The estimation device 100 performs pairing between cameras (S16), and pairs fish with body parts to generate a fish list (S17). The estimation device 100 performs a fish tracking process (S18) and a tail beat removal process (S19). The estimation device 100 selects the optimal frame for calculating body height (S20).

[0108] The estimation device 100 determines whether or not it is possible to calculate the body height (S21), and if it is not possible to calculate the body height (NO in S21), it performs the process of step S20. If it is possible to calculate the body height (YES in S21), the estimation device 100 calculates the body height (S22). The body height is calculated using the method shown in FIG. 16B. The term "calculating" the body height is synonymous with "measuring," but it also means measuring indirectly. The estimation device 100 determines whether or not there are other fish (fish for which measurements have not been performed) (S23).

[0109] If there are other fish (YES in S23), the estimation device 100 repeats the processes from step S11 onwards. If there are no other fish (NO in S23), the estimation device 100 collects the calculation data of the fish (S24) and classifies the calculation data into ranks according to the error (S25). The estimation device 100 calculates statistical values ​​(e.g., average values) of the calculation data for each rank (S26), creates a distribution for each estimation item (e.g., fork length, withers height, etc.) (S27), outputs the estimation results (S28), and ends the process.

[0110] The estimation device 100 can be configured with a CPU, a GPU, a ROM, a RAM, a recording medium reading unit, and the like. A computer program recorded on a recording medium can be read by the recording medium reading unit and stored in the RAM. The computer program stored in the RAM can be executed by the CPU and the GPU to perform the processing performed by the estimation device 100. Furthermore, instead of being read by the recording medium reading unit, the computer program can also be downloaded via a network such as the Internet.

[0111] According to this embodiment, body height measurement is performed by combining a segmentation image and a body height line, so that even if there is no optimal part for body height measurement, the size of the fish can be estimated stably and with high accuracy.

[0112] Furthermore, fish fins expand and retract while swimming, which can reduce measurement accuracy. However, according to this embodiment, segmentation is performed on the fish body excluding the fins, so body height can be measured with high accuracy without being affected by the fins.

[0113] Furthermore, extracting feature points from high-resolution images with high accuracy increases the computational load for image processing. However, according to this embodiment, by combining the first learning model and the generative model, it is possible to estimate fish size with high accuracy by combining the segmentation image and the body height line, and also reduce the computational load. [Explanation of symbols]

[0114] 100 Estimator 10 Input section 20 AI Department 21 First Learning Model 211 Feature Pyramid Network 212-class classifier 213 Area Regressor 22 Generative Model 221, 222, 223, 224, 225 Encoders 226, 227, 228, 229 decoders 23 Second Learning Model 24 Image extraction section 231 CNN layer 232 Domain Proposal Network 233 ROI POOL 234 Identification Network 30 Pairing Section 40 Tracking part 41 2D tracking section 42 3D conversion unit 43 3D tracking section 50 Tail beat removal section 51 Tail fork 3D measurement unit 52 Plane projection section 53 Approximate straight line calculation section 54 Frame removal section 60 Measurement section 61 Optimal frame selection section 62 Snout and tail 3D measurement unit 63 Body height auxiliary line generation section 64 Body height auxiliary line projection 65 Height Measurement Section 70 Estimation part 80 Output section 90 Estimation result DB 101 CPU 102 ROM 103 RAM 104 GPU 105 Video Memory 106 Recording medium reading unit 200 Camera Department

Claims

1. On the computer, Acquire image data of underwater creatures, inputting the acquired image data into a first learning model that outputs position data of a predetermined part of an aquatic organism when image data is input, and acquiring position data of a predetermined part of the captured aquatic organism; inputting the acquired image data into a generative model that generates a segmented image of an aquatic organism when image data is input, and generating a segmented image of the captured aquatic organism; estimating the size of the captured underwater creature based on the position data of the predetermined portion and the segmentation image; inputting the acquired image data into a second learning model that outputs first position data of an aquatic organism and second position data of a predetermined part of the aquatic organism when image data is input, and outputting the first position data and the second position data; image data of a second region including the second position data is input to the first learning model; image data of a first region including the first position data is input to the generative model; A computer program that executes a process.

2. A computer, Acquire image data of underwater creatures, inputting the acquired image data into a first learning model that outputs position data of a predetermined part of an aquatic organism when image data is input, and acquiring position data of a predetermined part of the captured aquatic organism; inputting the acquired image data into a generative model that generates a segmented image of an aquatic organism when image data is input, and generating a segmented image of the captured aquatic organism; estimating the size of the captured underwater creature based on the position data of the predetermined portion and the segmentation image; the predetermined site includes a caudal fin and a snout tip, In addition, the computer generating a body height extension line based on the three-dimensional position data of the caudal fin and the three-dimensional position data of the snout tip; two-dimensionally projecting the generated body height auxiliary line onto the segmentation image to calculate the body height; A computer program that executes a process.

3. On the computer, Acquire image data of underwater creatures, inputting the acquired image data into a first learning model that outputs position data of a predetermined part of an aquatic organism when image data is input, and acquiring position data of a predetermined part of the captured aquatic organism; inputting the acquired image data into a generative model that generates a segmented image of an aquatic organism when image data is input, and generating a segmented image of the captured aquatic organism; estimating the size of the captured underwater creature based on the position data of the predetermined portion and the segmentation image; inputting image data captured by the stereo camera into the first learning model, and calculating the three-dimensional positions of the snout tip and the tail fork based on the acquired position data of the predetermined part; A height auxiliary line is generated based on the calculated three-dimensional position. two-dimensionally projecting the generated body height auxiliary line onto the segmentation image to calculate the body height; A computer program that executes a process.

4. On the computer, Calculating the withers height using the withers height extension line located at a position that is a predetermined percentage of the distance between the snout tip and the tail fork from the position of the snout tip or the tail fork.

4. A computer program according to claim 3, which causes a process to be executed.

5. On the computer, inputting image data of each of a plurality of frames captured by the stereo camera into the first learning model, and calculating a three-dimensional position of the tail fork for each frame based on the acquired position data of the predetermined part; The displacement of the tail fork is determined for each frame based on the calculated three-dimensional position of the tail fork. selecting a frame for estimating the size of the aquatic creature based on the determined displacement; 5. A computer program product according to claim 1, which causes a process to be executed.

6. On the computer, displaying an estimated position of the caudal fin or the snout tip based on the position data output by the first learning model; Accepts corrections to the estimated position of the displayed caudal fin or snout tip, re-learning the first learning model based on the received correction position and the image data at the time when the correction position was received; 6. A computer program product according to claim 1, which causes a process to be executed.

7. On the computer, Displaying the segmentation image generated by the generative model; Accept corrections to the displayed segmentation image, re-training the generative model based on the corrected segmentation image and the image data at the time the correction was received; 7. A computer program product according to claim 1, which causes a process to be executed.

8. On the computer, Displaying the image of the aquatic organism being cultivated in the fish tank with its fork length or body height. A computer program product according to any one of claims 1 to 7, which causes a process to be executed.

9. On the computer, classifying the accuracy of the estimated sizes of the plurality of aquatic organisms cultivated in the fish tank into a plurality of ranks, and displaying the number of aquatic organisms for each rank; A computer program product according to any one of claims 1 to 8, which causes a process to be executed.

10. On the computer, displaying a size distribution including at least one of fork length and body height of a plurality of aquatic organisms cultured in the fish pen; 10. A computer program product according to claim 1, which causes a process to be executed.

11. Acquire image data of underwater creatures, inputting the acquired image data into a first learning model that outputs position data of a predetermined part of an aquatic organism when image data is input, and acquiring position data of a predetermined part of the captured aquatic organism; inputting the acquired image data into a generative model that generates a segmented image of an aquatic organism when image data is input, and generating a segmented image of the captured aquatic organism; estimating the size of the captured underwater creature based on the position data of the predetermined portion and the segmentation image; inputting the acquired image data into a second learning model that outputs first position data of an aquatic organism and second position data of a predetermined part of the aquatic organism when image data is input, and outputting the first position data and the second position data; image data of a second region including the second position data is input to the first learning model; image data of a first region including the first position data is input to the generative model; An estimation method that causes a computer to execute processing.

12. Acquiring image data of underwater organisms, inputting the acquired image data into a first learning model that outputs position data of a predetermined part of an aquatic organism when image data is input, and acquiring position data of a predetermined part of the captured aquatic organism; inputting the acquired image data into a generative model that generates a segmented image of an aquatic organism when image data is input, and generating a segmented image of the captured aquatic organism; estimating the size of the captured underwater creature based on the position data of the predetermined portion and the segmentation image; the predetermined site includes a caudal fin and a snout tip, moreover, generating a body height extension line based on the three-dimensional position data of the caudal fin and the three-dimensional position data of the snout tip; two-dimensionally projecting the generated body height auxiliary line onto the segmentation image to calculate the body height; An estimation method that causes a computer to execute processing.

13. Acquiring image data of underwater organisms, inputting the acquired image data into a first learning model that outputs position data of a predetermined part of an aquatic organism when image data is input, and acquiring position data of a predetermined part of the captured aquatic organism; inputting the acquired image data into a generative model that generates a segmented image of an aquatic organism when image data is input, and generating a segmented image of the captured aquatic organism; estimating the size of the captured underwater creature based on the position data of the predetermined portion and the segmentation image; inputting image data captured by the stereo camera into the first learning model, and calculating the three-dimensional positions of the snout tip and the tail fork based on the acquired position data of the predetermined part; A height auxiliary line is generated based on the calculated three-dimensional position. two-dimensionally projecting the generated body height auxiliary line onto the segmentation image to calculate the body height; An estimation method that causes a computer to execute processing.

14. a first acquisition unit that acquires image data of an underwater organism; a second acquisition unit that inputs the image data acquired by the first acquisition unit into a first learning model that outputs position data of a predetermined part of an aquatic organism when image data is input, and acquires position data of a predetermined part of the captured aquatic organism; a generation unit that inputs the image data acquired by the first acquisition unit into a generation model that generates a segmentation image of an aquatic organism when image data is input, and generates a segmentation image of the captured aquatic organism; an estimation unit that estimates the size of the captured aquatic organism based on the position data of the predetermined portion and the segmentation image; an output unit that inputs the acquired image data into a second learning model that outputs first position data of an aquatic organism and second position data of a predetermined part of the aquatic organism when the image data is input, and outputs the first position data and the second position data; a first input unit that inputs image data of a second region including the second position data into the first learning model; a second input unit for inputting image data of a first region including the first position data into the generative model; An estimation device comprising:

15. A first acquisition unit that acquires image data of an underwater organism; a second acquisition unit that inputs the image data acquired by the first acquisition unit into a first learning model that outputs position data of a predetermined part of an aquatic organism when image data is input, and acquires position data of a predetermined part of the captured aquatic organism; a generation unit that inputs the image data acquired by the first acquisition unit into a generation model that generates a segmentation image of an aquatic organism when image data is input, and generates a segmentation image of the captured aquatic organism; an estimation unit that estimates the size of the captured underwater creature based on the position data of the predetermined portion and the segmentation image; Equipped with the predetermined site includes a caudal fin and a snout tip, moreover, a body height extension line generating unit that generates a body height extension line based on the three-dimensional position data of the caudal fin and the three-dimensional position data of the snout tip; a calculation unit that calculates the body height by two-dimensionally projecting the generated body height auxiliary line onto the segmentation image; An estimation device comprising:

16. A first acquisition unit that acquires image data of an underwater organism; a second acquisition unit that inputs the image data acquired by the first acquisition unit into a first learning model that outputs position data of a predetermined part of an aquatic organism when image data is input, and acquires position data of a predetermined part of the captured aquatic organism; a generation unit that inputs the image data acquired by the first acquisition unit into a generation model that generates a segmentation image of an aquatic organism when image data is input, and generates a segmentation image of the captured aquatic organism; an estimation unit that estimates the size of the captured aquatic organism based on the position data of the predetermined portion and the segmentation image; a calculation unit that inputs image data captured by a stereo camera into the first learning model and calculates the three-dimensional positions of the snout tip and the fork based on the acquired position data of the predetermined part; a body height extension line generating unit that generates a body height extension line based on the calculated three-dimensional position; a body height calculation unit that calculates the body height by two-dimensionally projecting the generated body height auxiliary line onto the segmentation image; An estimation device comprising:

Citation Information

Patent Citations

  • Monitor device for ecology of fishes

    JP1989195364A

  • Image recognition device, image recognition method, and image recognition device-purpose program

    JP2019003554A

  • JPP6842100B

  • Fish biomass, shape, and size determination

    US20200184206A1

  • Information processing device, length measurement system, length measurement method, and program storage medium

    WO2018061925A1