Observation device, observation method, and program

The observation device uses passive acoustic observation and additional units to dynamically gather underwater sound data, addressing inefficiencies in existing methods by improving the acquisition and verification of sound source training data, enhancing classification accuracy.

WO2026105649A1PCT designated stage Publication Date: 2026-05-21BIOLOGGING SOLUTIONS INC
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BIOLOGGING SOLUTIONS INC
Filing Date
2025-11-06
Publication Date
2026-05-21

AI Technical Summary

Technical Problem

Existing methods for acquiring training data and verifying the accuracy of sound source identification in underwater environments are laborious and inefficient, as they require extensive manual effort and are limited by the unpredictability of sound occurrence and environmental differences.

Method used

An observation device equipped with passive acoustic observation, discrimination AI, and additional observation units that dynamically determine the need for further data collection based on sound source characteristics and environmental conditions, using hydrophones, active sonars, and optical cameras to efficiently gather true value data.

Benefits of technology

Facilitates the efficient acquisition of training data and verification of sound source identification by reducing manual labor and environmental limitations, enhancing the accuracy and efficiency of underwater sound source classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025038894_21052026_PF_FP_ABST
    Figure JP2025038894_21052026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention efficiently acquires information (generated sound, shape and position of an object, etc.) related to discrimination of underwater living things, artificial objects, and natural objects. An observation device 10 comprises a passive acoustic observation unit 11 that acquires candidate information of sound emitted by an observation object existing in water on the basis of a method for passively observing sound generated in water, a determination unit 14 that determines whether to perform additional observation on the basis of the acquired candidate information, and an additional observation unit 15 that performs additional observation related to a sound source on the basis of the determination result of the determination unit 14.
Need to check novelty before this filing date? Find Prior Art

Description

Observation device, observation method, and program

[0001] The present invention relates to an observation device, an observation method, and a program.

[0002] In water, sound travels much farther than light. In air, the speed of sound is about 330 m per second, but in water, it is approximately 4.5 times that, traveling about 1,500 m per second. Although much slower than light, the transmission distance is long and the influence is immediate. Many marine organisms make good use of sound and are sensitive to it. For example, they generate sound for underwater exploration, communication, intimidation, reproduction, etc.

[0003] In recent years, many reports have been made on the impact of sound exposure on organisms. There are various artificial sound sources in the sea. Among them, the most numerous are ships. It has been pointed out that the sound emitted from ships along with maritime transportation reduces the distance at which whales and fish can communicate, and in some cases, it may also affect reproduction.

[0004] On the other hand, there are also naturally occurring noises in the sea. The main sound sources are pulse sounds such as those of snapping shrimp constantly emitted from coastal areas. Underwater sound is also generated by weather and sea conditions. The broadband sound caused by rainfall is radiated from the sea surface into the sea. Broadband sound is also generated from the bubbles formed when waves break in rough weather.

[0005] As described above, underwater sound is composed of artificial sound, natural sound, and biological sound. By using passive acoustic observation (hereinafter referred to as passive acoustic observation) to passively observe underwater sound, various information caused by various sound sources can be collected. Biological sounds include not only the sounds emitted by organisms (fish, whales, etc.), but also all sounds associated with the behavior and movement of organisms, such as chewing sounds and the swimming sounds of fish schools. Natural sounds include wind, rain, waves, currents, volcanic sounds, etc. Artificial sounds include sounds emitted from ships, piling, air guns, offshore wind power, sonar, and divers.

[0006] "Bio-logging" [online] May 2019: First edition published, National Institute for Environmental Studies [Accessed May 13, 2024], Internet<URL:https: / / tenbou.nies.go.jp / science / description / detail.php?id=109>

[0007] By acquiring acoustic information from passive acoustic observations and utilizing the fact that the characteristics of sound differ depending on the sound source, it is possible to create an AI (Artificial Intelligence; hereinafter referred to as "discrimination AI") that can identify the sound source. Examples of AI include classification of time-series waveforms of sound using signal processing and deep learning models, as well as image classification and image recognition of spectrogram images, etc.

[0008] By associating generated sounds with their sources, it is possible to train a discrimination AI. However, acquiring training data sets (data used to train a learning model; hereinafter referred to as training data) that include input data and corresponding output data used for training artificial intelligence in the field (outdoor environment) requires a great deal of effort.

[0009] For example, fish are known to produce sounds for purposes such as courtship, intimidation, and communication, but the timing for capturing these sounds is limited. Therefore, it is not always possible to capture sounds when acquiring underwater sounds or when installing underwater cameras to identify the sound source. There are also limits to the amount of time humans can spend underwater, and the presence of humans may cause organisms to behave in an abnormal way. Attempts to record target sounds in captivity tanks are also unsuccessful, as the environmental conditions in tanks differ from those in natural environments, and the target sounds are often not emitted. Natural phenomena such as wind, rain, and waves also require timely recording of the sounds as they occur, but in rough seas, even the task of installing underwater observation equipment can be difficult. Even if a correspondence between generated sounds and sound sources is established at one site, the characteristics of the sounds and the types of sounds produced will differ depending on the time and location. As a result, untrained data exists, and acquiring new training data that identifies generated sounds and sound sources requires considerable effort.

[0010] Furthermore, verifying the accuracy of sounds identified by the AI ​​also requires considerable effort. To verify the accuracy of sounds identified by the AI, it is necessary to check the sound source in a natural environment. However, since it is impossible to know when a sound will occur, obtaining sound source information is, for the same reasons mentioned above, extremely laborious.

[0011] Furthermore, when identifying sound sources through passive acoustic observation, one method involves installing acoustic recorders and cameras (other methods described later) in natural bodies of water for extended periods, retrieving them from their locations after a certain time, downloading the data from the devices to a computer, and analyzing the obtained data using discrimination AI. Cameras and other devices can acquire data about the sound source from the video, and are used to obtain true value data (training data) that accurately identifies what the sound source actually is, for example, whether a particular sound originates from a ship, a fish, or another marine organism. However, with this method, not only is it impossible to know information about the sound source in real time while it is being generated, but the amount of data becomes enormous, including the true value data of sound sources that are not the target, requiring a large amount of memory and power from the device. In one aspect, the present invention aims to efficiently acquire training data for underwater sounds.

[0012] To achieve the above objective, the disclosed observation device is provided. This observation device includes a first observation unit that acquires candidate sound information emitted by an object present in the water based on a method of passively observing sound generated in water, a determination unit that determines whether or not to perform additional observations based on the acquired candidate information, and a second observation unit that performs additional observations regarding the sound source based on the determination result of the determination unit.

[0013] In one embodiment, training data of underwater sounds can be efficiently acquired. The above and other objects, features and advantages of the present invention will become apparent from the following description in conjunction with the accompanying drawings illustrating preferred embodiments as examples of the present invention.

[0014] This is a diagram showing the observation device of the embodiment. This is a diagram showing the hardware configuration of the observation device of the embodiment. This is a diagram showing the hardware configuration of the observation device of the embodiment. This is a block diagram explaining the functions of the observation device of the embodiment. This is a flowchart explaining the first process of the observation device. This is a flowchart explaining the first process of the server device. This is a flowchart explaining the second process of the observation device. This is a flowchart explaining the third process of the observation device. This is a flowchart explaining the fourth process of the observation device.

[0015] The observation system of this embodiment will be described in detail below with reference to the drawings.

[0016] The positions, sizes, shapes, and ranges of each component shown in the following drawings may not represent the actual positions, sizes, shapes, and ranges in order to facilitate understanding of the invention. Therefore, the present invention is not necessarily limited to the positions, sizes, shapes, and ranges disclosed in the drawings. Elements expressed in the singular form in the embodiments shall include plural forms unless explicitly stated in the text. <Embodiment> Figure 1 is a diagram showing an observation system of an embodiment.

[0017] The observation system 100 of this embodiment includes an observation device (computer) 10 and a server device 20. The observation device 10 and the server device 20 are connected by a network 30. When the communication unit of the observation device 10 is exposed in the air, the means of communication between the observation device 10 and the network 30 may include, for example, wireless communication using a mobile phone network or satellite communication network, infrared communication, or short-range wireless communication using Bluetooth®. Furthermore, when the observation device 10 is completely submerged in water, it can also communicate with the network 30 via a surface relay device such as an ASV (Autonomous Surface Vehicle) or buoy using underwater communication. The means of communication between the server device 20 and the network 30 may include, for example, a VPN (Virtual Private Network) or the Internet.

[0018] One of the objectives of the observation device 10 of this embodiment is to efficiently acquire information (such as generated sound, shape, and position of objects) that allows for the identification of observation targets such as organisms, artificial objects, and natural objects present in the water. This observation device 10 is installed, for example, in rivers or the sea. The observation device 10 may be attached to living organisms (for example, turtles or sharks). The observation device 10 may also be attached to an autonomous underwater vehicle (AUV) or a remotely operated vehicle (ROV). Furthermore, the observation device 10 may be placed in the mid-water layer (the area of ​​water located between the surface and the bottom) by attaching it to a buoy or the like (not shown). It may also be placed on the seabed or ocean floor. Figures 2 and 3 show the hardware configuration of the observation device of this embodiment.

[0019] The observation device 10 is controlled as a whole by a CPU (Central Processing Unit) 101. However, control may be performed using other components besides the CPU 101, such as an FPGA (Field Programmable Gate Array), ASIC (Application-Specific Integrated Circuit), microcontroller, other logic circuits, or a PLD (Programmable Logic Device). The CPU 101 is connected to a RAM (Random Access Memory) 102 and several peripheral devices via a bus 111.

[0020] The RAM 102 is used as the main memory of the observation device 10. At least a portion of the OS (Operating System) program and application programs to be executed by the CPU 101 are temporarily stored in the RAM 102. The RAM 102 also stores various data used for processing by the CPU 101.

[0021] The bus 111 is connected to an internal memory 103, an optical camera (with light source) 104 (104a, 104b), an active sonar 105 (105a, 105b), a hydrophone 106 (106a, 106b, 106c, 106d), an environmental sensor 107, an acoustic modem 108, an acceleration / geomagnetic / gyro sensor 109, and a GPS 110.

[0022] The internal memory 103 performs data writing and reading. The internal memory 103 is used as a secondary storage device for the observation device 10. The internal memory 103 stores the OS program, application programs, and various data. Examples of internal memory include semiconductor storage devices such as flash memory. The optical cameras 104a and 104b capture images of the underwater environment.

[0023] The active sonars 105a and 105b determine the distance and position of an object by generating sound waves and recording the sound waves reflected back from the object using a microphone. The active sonars 105a and 105b calculate the distance by measuring the round-trip time of the sound waves, thereby determining the precise position of the object. Furthermore, the size and shape of the object can be estimated from the characteristics of the reflected echoes. In addition, the material from which the object is made can be determined by the difference in the reflectivity of the sound waves. The active sonars 105a and 105b also provide important information regarding the movement of objects. Because the frequency of reflected waves from a moving object changes due to the Doppler effect, the speed and direction of movement of the object can be determined by analyzing this. In particular, this technology is used in fish finders and other devices for identifying marine life, and is useful in tracking the position and movement of large marine creatures such as whales and dolphins. The environmental sensor 107 detects depth, temperature, electrical conductivity, dissolved oxygen, current direction and velocity, turbidity, etc.

[0024] The acoustic modem 108 can connect to the network 30 via a floating relay device. The acoustic modem 108 sends and receives data to and from other computers or communication devices via the network 30. It is also possible that the system does not include an acoustic modem. The acceleration, geomagnetic, and gyro sensors 109 determine the motion and attitude of the observation device 10.

[0025] The processing functions of this embodiment can be realized with the hardware configuration described above. Note that multiple optical cameras 104, active sonar 105, and environmental sensors 107 can be mounted as variations of the device. Furthermore, regarding the acceleration, geomagnetic, and gyro sensors 109, which detect the movement and attitude of the device, it is possible to use only a portion of them, or to use other sensors (such as tilt sensors) to achieve the objective, as variations. Alternatively, the acceleration, geomagnetic, and gyro sensors 109 do not necessarily need to be mounted. Regarding the GPS 110, which detects position, it is possible to omit it, or to use other sensors (such as Wi-Fi® or other radio wave or ultrasonic position detection methods) to achieve the objective, as variations. The environmental DNA analysis unit for additional observation, the active sonar 105, and the optical camera 104 can also be omitted, as variations. The optimal configuration will vary depending on the overall size and shape of the device, power design, and observation target. Figure 4 is a block diagram illustrating the functions of the observation device of this embodiment.

[0026] The observation device 10 of this embodiment includes a passive acoustic observation unit 11, a chemical / physical environment measurement unit 12, a storage unit 13, a judgment unit 14, an additional observation unit 15, a control unit 16, and a data transmission / reception unit 17. Note that, depending on the overall size, shape, and power design of the device, a variation is also conceivable in which the data transmission / reception unit is omitted, and the results are stored in the storage unit, downloaded to a PC after the observation device is retrieved, and the results are then retrieved.

[0027] The passive acoustic observation unit 11 is an example of the first observation unit. The passive acoustic observation unit 11 continuously measures sounds (including noise) generated from living organisms, ships, etc., using a hydrophone 106. That is, the passive acoustic observation unit 11 does not generate sound itself, but passively collects external sounds. By using multiple hydrophones 106 (three or more), and positioning the hydrophones at an appropriate distance from the frequency characteristics of the target sound, and acquiring sound, the position and direction of the target sound can be estimated by acoustic positioning. Even when using two hydrophones, if the hydrophones are positioned at an appropriate distance from the frequency characteristics of the target sound and sound is acquired, the direction of the target sound can be estimated. When estimating direction and position by measuring the TDOA (Time Difference of Arrival), which is the time difference in the arrival of sound waves, or the phase difference, it is preferable that the placement distance of the hydrophones be less than half the wavelength of the sound. Assuming a sound speed of 1500 m / s, if the sound frequency is 1 kHz, it is preferable to place the hydrophones at a distance of 0.75 meters or less. On the other hand, if the sound frequency is 100 kHz, it is preferable to place the hydrophones at a distance of 7.5 millimeters or less. However, placing the hydrophones at too short a distance can also cause problems. Specifically, if the distance is too short, the time difference and phase difference of the sound waves become very small, requiring an accuracy that exceeds the time resolution of the equipment, which can increase measurement errors and reduce the accuracy of estimating direction and position. Therefore, although it is preferable to place the hydrophones at a distance of less than half the wavelength, it is important to maintain an appropriate distance considering the performance of the equipment and the measurement environment. Whether only direction can be estimated or position can also be estimated depends on whether two hydrophones or three or more hydrophones are used, but for simplicity, we will refer to this collectively as acoustic positioning from now on.

[0028] The chemical and physical environment measurement unit 12 uses environmental sensors 107 to detect, for example, dissolved oxygen in water (the amount of oxygen dissolved in water), electrical conductivity, pH, chlorophyll turbidity, temperature, etc.

[0029] Here, chlorophyll turbidity refers to the chlorophyll concentration in water and the resulting turbidity. Chlorophyll is a pigment involved in the photosynthesis of phytoplankton, and its concentration indicates the amount of phytoplankton in the water.

[0030] Underwater acoustics differ in characteristics and sound propagation depending on physical environmental factors such as dissolved oxygen, electrical conductivity, and temperature. Therefore, considering physical environmental information can improve the estimation of the location of the target sound. In the following explanation, the data detected by the chemical and physical environment measurement unit 12 will be collectively referred to as "chemical and physical environment data."

[0031] The memory unit 13 may also store pre-prepared training data for sound sources. Generally, due to power and size limitations of the observation device, the discrimination AI is often an AI that has been trained and built on a PC or the like beforehand and then introduced into the observation device. However, as power efficiency improves in the future, it is conceivable that an AI trained within the observation device will also be used. In that case, the training data for sound sources will be stored within the observation device. The judgment unit 14 has a discrimination AI 14a and an OOD (Out Of Distribution) detection unit 14b.

[0032] The discrimination AI 14a learns from training data and distinguishes between artificial objects and living organisms from acoustic data (acoustic data) obtained from observations by the passive acoustic observation unit 11. Discrimination can be performed not only on broad categories such as the type of ship among artificial objects (tankers and fishing boats) or fish, crustaceans, and marine mammals among living organisms, but also on specific species (moray eels, bottlenose dolphins, dugongs), and even on multiple types of sounds within the same species (for dolphins, whistles, clicks, burst pulses, etc.). The degree of detail in the discrimination depends on the design of the discrimination AI. Discrimination AI 14a can perform signal processing, classification of time-series waveforms of sound using deep learning models, and image classification and image recognition of spectrogram images, etc.

[0033] Here, "discrimination" refers to judging and identifying the likelihood of a given object belonging to a specific class. Here, a class refers to a category or group of data. For example, in the classification of acoustic data, a class refers to a specific type of sound source, such as "ship sounds," "fish sounds," "natural sounds," or "moray eel sounds." Machine learning models learn to assign new data to these classes. Here, "likelihood" refers to a probability value or score (a broad indicator including, for example, Softmax probability, Logit value, energy score, feature distance, reconstruction error, etc.).

[0034] This process involves preparing and classifying training data for each class in advance. The training data includes input data and the corresponding correct output (label). For example, in the classification of acoustic data, the training data consists of various acoustic samples (input data) and labels (outputs) that indicate which class each sample belongs to. In other words, the label is an indicator that shows the class to which each input data in the training data belongs. For example, if a particular acoustic sample belongs to "ship noise," that sample will be labeled "ship noise." The label functions as a class name or ID and serves as a guideline for the model to learn which class to assign the data to.

[0035] The model uses this training data to learn the characteristics of each class. If the classes are not comprehensive, the target may not be included in the pre-prepared classes. However, if the classes are sufficiently comprehensive, the target may be included in the pre-prepared classes, and the classification result can be considered as discriminative. Since the AI ​​ultimately aims to accurately identify the target, it is referred to as "discrimination AI" rather than "classification AI."

[0036] The decision unit 14 determines whether or not to perform additional observations with the additional observation unit 15, based on the results of the discrimination accuracy by the discrimination AI 14a, the detection results of abnormal data and unknown data by OOD detection, and the results of acoustic characteristics and sound source distance and direction calculated from the acoustic positioning of the underwater chemical / physical environment measurement unit 12 and the passive acoustic observation unit 11 (hereinafter referred to as candidate information). It is also conceivable that the chemical / physical environment measurement unit and acoustic positioning may not necessarily be installed, and the decision on whether or not to perform additional observations may be based only on the results of the discrimination accuracy by the discrimination AI 14a and the detection results of abnormal data and unknown data by OOD detection. Whether or not to perform chemical / physical environment measurement and acoustic positioning depends on the size and shape of the equipment and the power design.

[0037] Here, abnormal data refers to data that is not included in the training data and is significantly different from the known training data. For example, in marine environmental monitoring, equipment malfunction sounds or sudden abnormal sounds mixed in with normal acoustic data would be considered abnormal data.

[0038] Unknown data refers to a new type of data that is not included in the training data. For example, if an AI whose training data only includes acoustic data of ordinary marine organisms and ships encounters a completely new type of sound source (the sound of a new machine or the sound of an unknown organism), that data becomes unknown data. Specifically, for example, when discriminative AI 14a discriminates the class of an input sound source, it calculates the probability or score of belonging to all the classes it has learned in advance. For example, discriminative AI 14a based on training data acquired in the sea near Okinawa only includes organisms that inhabit the sea near Okinawa (such as tropical fish) in its classes. When this discriminative AI is used to detect a specific sound in the sea near Hokkaido, it is highly likely that almost all of the sounds of organisms inhabiting the sea near Hokkaido will be sounds of an unknown class that the AI ​​has never heard before. In the case of sounds of an unknown class, the probability or score of belonging to that class may be low, or it may be an abnormally high probability or score, for all of the classes it has learned in advance. In this way, an indicator of the certainty of the discriminative AI can be used for OOD event detection (other OOD detection algorithms will be described later). When OOD is detected, it is determined to be unknown data, and additional observation is performed. In this way, the determination unit 14 calculates the probability or score that the candidate information belongs to any of the multiple classes that have been learned in advance, using the probability calculation function or score calculation function set during learning, and when it is determined that the probability or score is less than a predetermined threshold, it decides to perform additional observation with the additional observation unit 15.

[0039] Additional observations can also be performed when a pre-specified class is identified. For example, the additional observation unit 15 can be instructed to perform additional observations only when a dugong sound is produced in order to perform additional learning. If the sound of a dugong is included in the training data, the discrimination AI 14a is highly likely to identify it as a dugong when the sound is actually produced. To confirm the reliability of the information and whether it was actually a dugong, additional observations can be performed to acquire information using active acoustics or optical cameras to investigate whether it was truly a dugong.

[0040] Incidentally, if the target sound is too far away, there is a high probability that the target cannot be detected even if the additional observation unit 15 performs observations. For this reason, the judgment unit 14 can decide to perform additional observations only when the target sound is within a certain distance range. The distance to the target varies depending on water quality (transparency, etc.) and time of day (nighttime), and may be changed depending on chemical and physical environmental conditions, time of day, and season. Also, the range targeted by the optical camera and the range targeted by the active sonar are different, and generally the active sonar can target a greater range, so the sensing configuration used for additional observations can also be changed according to the distance and environmental conditions. The first distance and the second distance are preset thresholds and can be changed according to the aforementioned environmental conditions and operating conditions. Furthermore, since image acquisition by active sonar and optical cameras is directional, estimating the direction of the sound source in advance before performing additional observations increases the likelihood of detecting the target sound. By estimating the direction of the target sound in advance, the orientation of the active sonar and optical cameras in the additional observation unit can be determined. This allows for efficient additional observation of the target sound by activating only the active sonar or optical camera that matches the direction and distance of the target sound from among multiple active sonars and optical cameras installed in different directions, or by mechanically changing the orientation of the active sonar or optical camera to match the direction of the target sound before additional observation.

[0041] The OOD detection unit 14b identifies inputs (events) that have a data distribution different from the training data, which is determined by the machine learning model (AI). This can be used to enable the model to react appropriately when it encounters unknown or anomalous data. Methods for detecting events in the OOD detection unit 14b include using an autoencoder or the probability or score of the discrimination AI 14a to detect changes in the distribution. Other statistical methods (such as measuring the distance in the feature space or statistical values ​​of the distribution) and ensemble methods (judgment based on predictions from multiple models) are also possible. An autoencoder is a type of unsupervised learning using a neural network, aiming to learn an efficient representation (encoding) of data and then reconstruct (decode) the original data. It is mainly used for data compression, noise reduction, and feature extraction. A neural network is a computational model that mimics the neural circuits of the human brain and consists of multiple layers (input layer, hidden layer, output layer). Each layer contains numerous "neurons," which are connected to each other to process information. Neural networks are used for pattern recognition, data classification, and prediction, and are particularly widely used in the fields of machine learning and deep learning.

[0042] A concrete example of using the probability or score of the discrimination AI 14a is to analyze the distribution of the probability or score of the class that the discrimination AI 14a predicts for a new input, and if it has a low probability or score or an unusually high probability or score, it is determined that the input may be coming from a different distribution than the training data. In the following description, the detection of events by the OOD detection unit 14b is referred to as "detection of OOD events". When an OOD event is detected by the OOD detection unit 14b, or when a pre-specified class is determined, it is possible to decide whether to perform additional observations.

[0043] The additional observation unit 15 is an example of the second observation unit. The additional observation unit 15 includes an environmental DNA analysis unit 15a, an active sonar 105, and an optical camera 104. The additional observation unit may be composed of a part of the environmental DNA analysis unit 15a, the active sonar 105, and the optical camera 104, or other sensors for achieving the purpose of additional observation may also be considered as variations (the details of the additional observation will be described later). Additional observations using remote sensing, underwater drones or aerial drones, LiDAR (Light Detection And Ranging) may also be possible. Which configuration to use for additional observation depends on the device size, shape, and power design.

[0044] The environmental DNA analysis unit 15a analyzes environmental DNA to examine the organisms present in that environment. The environmental DNA analysis unit 15a is not only for in-situ analysis, but also for sampling by water collection, and it is assumed that the actual analysis may be carried out in a laboratory after the observation device 10 is recovered.

[0045] In the following description, the operations and detection operations for identifying the sounds of organisms performed by the active sonar 105, the optical camera 104, and the environmental DNA analysis unit 15a are collectively referred to as "additional observations". Also, the data obtained from the additional observations is referred to as "true value candidate data". The true value candidate data is data that serves as a candidate for identifying accurate information indicating what the sound source actually is. This true value candidate data is very important in the learning and evaluation of the AI model as teacher data and serves as a criterion for enabling the model to accurately discriminate the sound source.

[0046] Through the observation by the additional observation unit 15, true value candidate data regarding various sound sources can be obtained, and based on various multimodal information combined with chemical and physical environmental information, the reliability verification and re-learning of the discrimination AI 14a can be carried out.

[0047] The control unit 16 has a primary processing function for the data observed by the observation device 10. Also, the control unit 16 has a function for calculating the acoustic characteristics and acoustic positioning in the observation device 10.

[0048] The data transmission / reception unit 17 can transmit the data obtained by the observation device 10 to, for example, a server device 20 on the ground. Examples of the data to be transmitted to the server device 20 include acoustic data obtained by the observation of the passive acoustic observation unit 11 and true value candidate data obtained by the additional observation of the additional observation unit 15. Since the work of installing the observation device 10 underwater may also involve the labor of launching a ship and installing it by a diver, data transmission is also useful in order to reduce the frequency of recovering and reinstalling the observation device 10. However, since the data transmission / reception unit 17 increases the size and required power of the entire device, a variation is also conceivable in which the result is saved in the storage unit without being mounted, the observation device is recovered, and then downloaded to a PC to recover the result.

[0049] By checking the data transmitted to the server device 20, it is possible to check the discrimination result of the real-time discrimination AI 14a and the true value candidate data. As a data transmission method from underwater, for example, underwater communication using sound, light, etc. (and wireless communication via an above-water repeater), or communication using radio waves such as a satellite communication network, a mobile phone communication network, or short-range wireless by surfacing the device to the water surface by automatically ascending and descending from underwater is conceivable. Further, the data transmission / reception unit 17 can receive update data of the discrimination AI 14a provided in the determination unit 14 from the server device 20.

[0050] The server device 20 determines whether the determination of the determination unit 14 was correct based on the data received from the observation device 10. Then, the server device 20 transmits data for updating the discrimination AI 14a model itself and the OOD detection unit 14b model itself based on the determination result, data for updating the teacher data of the discrimination AI 14a and the OOD detection unit 14b, data for updating or changing the monitoring methods (sampling period, date and time, measurement parameters for calibration, etc.) of the passive acoustic observation unit 11, the chemical / physical environment measurement unit 12, and the additional observation unit 15, and the program of the control unit 16 to the data transmission / reception unit 17.

[0051] The observation device 10 receives judgment results and update data and stores them in the storage unit 13, thereby learning the judgment criteria of the discrimination AI 14a and the OOD detection unit 14b, improving the accuracy of the discrimination and changing the target of OOD detection. The observation device 10 also receives data related to the monitoring method and control program and stores it in the storage unit 13, thereby changing the measurement timing of the passive acoustic observation unit 11, the chemical / physical environment measurement unit 12, and the additional observation unit 15, as well as the measurement parameters for calibration and the frequency of data transmission and reception of the data transmission and reception unit 17. The server device 20 can also evaluate the accuracy of the discrimination AI 14a based on the judgment results. Furthermore, the server device 20 can determine the reliability of the observation by the passive acoustic observation unit 11 based on the acoustic data received from the observation device 10.

[0052] Furthermore, the observation device 10 may also perform a variation in which it retrains the models of the discrimination AI 14a and OOD detection unit 14b using the candidate true value data obtained from additional observations by the observation device 10 as training data for model retraining, thereby updating the models within the observation device. In addition, the judgment unit 14 may be configured to select whether or not to use the model if a judgment result or update data is received.

[0053] Next, an example of the processing of the observation system 100 will be explained. The observation system 100 can perform the following four processes (the first to the fourth processes). <First Process> The first process is to obtain training data associated with new sound sources for retraining the discrimination AI. Figure 5 is a flowchart illustrating the first process of the observation device.

[0054] [Step S1] The observation device 10 installed underwater has a passive acoustic observation unit 11 that performs monitoring. In addition, the chemical and physical environment measurement unit 12 measures chemical and physical environment data periodically (for example, once every 60 seconds).

[0055] [Step S2] The judgment unit 14 performs OOD detection by the OOD detection unit 14b based on the candidate information obtained through monitoring. If an OOD event is detected, the process proceeds to step S3. Examples of OOD event detection algorithms include the following: (Case 1): Input data that differs from the probability distribution of data learned so far is detected as an OOD event. (Case 2): The discrimination AI 14a is allowed to recognize the data, and if the discrimination AI 14a determines that the probability or score of belonging to any of the classes learned so far does not meet the threshold condition, that determination is detected as an OOD event.

[0056] [Step S3] The control unit 16 calculates the direction and distance of the sound source using the acoustic data observed by the passive acoustic observation unit 11. As mentioned above, the observation device 10 is equipped with multiple optical cameras 104a, 104b and active sonars 105a, 105b to cover as many directions as possible. Therefore, when performing optical camera observation or active acoustic observation, the control unit 16 decides which camera or sonar to use. After that, the process proceeds to step S4.

[0057] [Step S4] Based on the calculation result of step S3, the additional observation unit 15 determines whether the sound source is within a first distance. The first distance is a predetermined distance, which is the distance at which it is determined that the detection target can be imaged by underwater photography. If the sound source is within the first distance (No. of step S4), the process proceeds to step S5. If the sound source is within a second distance (No. of step S4), the process proceeds to step S6. The second distance is a predetermined distance, which is greater than the first distance, making it difficult to identify the target by underwater photography, but at which it is determined that the detection target can be identified by active sonar. If the sound source is farther than the second distance (No. of step S4), the process ends.

[0058] [Step S5] The additional observation unit 15 performs additional observations related to the sound source. Specifically, the additional observation unit 15 acquires the position, movement, shape, etc. of objects and organisms in the water using the active sonars 105a and 105b (true value candidate data). The additional observation unit 15 also activates the optical camera 104 in the direction determined in step S3 and images the water in the direction of the sound source (true value candidate data). After that, the process proceeds to step S7. Information may also be acquired by activating either the active sonars 105, 105b or the optical camera 104.

[0059] [Step S6] The additional observation unit 15 acquires the position, movement, shape, etc. of objects and organisms in the water using the active sonars 105a and 105b (true value candidate data). Alternatively, true value candidate data may be acquired by operating only the LiDAR or other observation means. After that, the process proceeds to step S7.

[0060] [Step S7] The data transmission / reception unit 17 transmits to the server device 20 the raw passive acoustic data (spectrogram image, etc.) obtained by the control unit 16 through primary processing of the raw passive acoustic data in which OOD was detected, together with the true value candidate data acquired in step S5 or step S6 and the chemical and physical environmental data measured in step S2. Environmental DNA data acquired by the environmental DNA analysis unit 15a may also be transmitted to the server device 20. Next, the processing on the server device 20 side in the first processing will be explained. Figure 6 is a flowchart illustrating the first processing of the server device.

[0061] [Step S11] The server device 20 compares the data sent from the data transmission / reception unit 17 of the observation device 10. Methods of comparison include, for example, automatic labeling by AI of the server device 20, or manual verification and correction by a human labeler. In this process, training data is accumulated using the compared data.

[0062] [Step S12] The server device 20 retrains the AI ​​model using the updated training data. This process may be performed after a certain amount of updated data has been accumulated.

[0063] [Step S13] The server device 20 transmits the retrained AI model to the observation device 10. After that, it terminates the process shown in Figure 6. The observation device 10, having received the AI ​​model, updates the AI ​​14a using this AI model.

[0064] Note that the operations shown in Figures 5 and 6 are examples, and some steps may be omitted or other steps may be added. For example, the optical camera 15b may be activated to take images regardless of the distance to the sound source. <Second Processing> The second processing is the process of verifying the results identified by the discrimination AI 14a using passive acoustic data as input data (verifying the accuracy of the discrimination AI 14a). In the second processing, only the program of the discrimination AI 14a is run, and the processing is advanced by triggering a specific sound emitted by a dugong or the like in advance. Figure 7 is a flowchart illustrating the second processing of the observation device.

[0065] [Step S21] The observation device 10 installed underwater has a passive acoustic observation unit 11 that performs monitoring. In addition, the chemical and physical environment measurement unit 12 measures chemical and physical environment data periodically (for example, once every 60 seconds).

[0066] [Step S22] The determination unit 14 operates the discrimination AI 14a based on the candidate information obtained through monitoring. If the discrimination AI 14a detects the occurrence of a specific sound that has been specified in advance, the process proceeds to step S23.

[0067] [Step S23] The control unit 16 calculates the direction and distance of the sound source using the acoustic data observed by the passive acoustic observation unit 11. The additional observation unit 15 also determines which camera or sonar to use when performing additional observations. The process then proceeds to step S24.

[0068] [Step S24] Based on the calculation result of step S23, the additional observation unit 15 determines whether the sound source is within the first distance. If the sound source is within the first distance (first step of step S24), the process proceeds to step S25. If the sound source is within the second distance (second step of step S24), the process proceeds to step S26. If the sound source is farther than the second distance (No. of step S24), the process ends.

[0069] [Step S25] The additional observation unit 15 performs additional observations related to the sound source. Specifically, the additional observation unit 15 acquires the position, movement, shape, etc. of objects and organisms in the water using active sonar (true value candidate data). The additional observation unit 15 also activates the optical camera 104 in the direction determined in step S23 and images the water in the direction of the sound source (true value candidate data). After that, the process proceeds to step S27.

[0070] [Step S26] The additional observation unit 15 acquires the position, movement, shape, etc. of objects and organisms in the water using the active sonars 105a and 105b (true value candidate data). After that, the process proceeds to step S27.

[0071] [Step S27] The data transmission / reception unit 17 transmits to the server device 20 the raw passive sound data identified by the discrimination AI 14a or the data (spectrogram image, etc.) obtained by the control unit 16 from the raw passive sound data, together with the true value candidate data acquired in step S25 or step S26 and the chemical / physical environment data measured in step S22. Environmental DNA data acquired by the environmental DNA analysis unit 15a may also be transmitted to the server device 20. The processing of the server device 20 in the second process is the same as that of the server device 20 in the first process, so the explanation is omitted. <Third process> The third process is the process of obtaining training data associated with a new sound source for retraining the discrimination AI 14a.

[0072] In the third process, the OOD detection program is operated using the active acoustic data as input data. The following explanation will omit details that are the same as those in the first process. Figure 8 is a flowchart illustrating the third process of the observation device.

[0073] [Step S31] The observation device 10 installed underwater has an additional observation unit 15 that activates the active sonar 105 to perform monitoring. The chemical and physical environment measurement unit 12 also periodically measures chemical and physical environment data (for example, once every 60 seconds).

[0074] [Step S32] The determination unit 14 performs OOD detection by the OOD detection unit 14b based on the candidate information obtained through monitoring. If an OOD event is detected, the process proceeds to step S33.

[0075] [Step S33] The additional observation unit 15 calculates the direction and distance of the target object using the active sonars 105a and 105b. As mentioned above, the observation device 10 is equipped with optical cameras 104a and 104b to cover as many sides as possible. Therefore, when performing active acoustic observation, it is decided which camera to use. After that, the process proceeds to step S34.

[0076] [Step S34] Based on the calculation result of step S33, the additional observation unit 15 determines whether the sound source is within the second distance. If the sound source is within the first distance (No. of step S34), the process proceeds to step S35. If the sound source is farther than the first distance (No. of step S34), the process ends.

[0077] [Step S35] The additional observation unit 15 activates the optical camera 104 corresponding to the active sonar 105 that detected the OOD event in step S32 and captures an image of the underwater area in the direction of the sound source (true value candidate data). After that, the process proceeds to step S36.

[0078] [Step S36] The data transmission / reception unit 17 transmits to the server device 20 the sonar image of the active acoustics in which OOD was detected, or the raw data of the active acoustics that has been processed by the control unit 16, together with the true value candidate data acquired in step S35 and the chemical and physical environmental data measured in step S32. Environmental DNA data acquired by the environmental DNA analysis unit 15a may also be transmitted to the server device 20. The processing of the server device 20 in the third process is the same as that of the server device 20 in the first process, so the explanation is omitted. <Fourth process> The fourth process is a process to verify the results identified by the discrimination AI 14a using the active acoustics as input data (verifying the accuracy of the discrimination AI 14a). In the fourth process, only the program of the discrimination AI 14a is run, and the process proceeds using a specific sound emitted by a specific identification target (for example, a school of fish) as a trigger. Figure 9 is a flowchart illustrating the fourth process of the observation device.

[0079] [Step S41] The observation device 10 installed underwater has an additional observation unit 15 that activates the active sonar 105 to perform monitoring. The chemical and physical environment measurement unit 12 also periodically measures chemical and physical environment data (for example, once every 60 seconds).

[0080] [Step S42] The determination unit 14 operates the discrimination AI 14a based on the candidate information obtained through monitoring. If a pre-specified identification target (such as a school of fish) is found, the process proceeds to step S43.

[0081] [Step S43] The additional observation unit 15 calculates the direction and distance of the target object using the active sonars 105a and 105b. The additional observation unit 15 also determines which camera to use when performing active acoustic observation. After that, the process proceeds to step S44.

[0082] [Step S44] Based on the calculation result in step S43, the additional observation unit 15 determines whether the sound source is within the second distance. If the sound source is within the first distance (Yes in step S44), the process proceeds to step S45. If the sound source is farther than the first distance (No in step S44), the process ends.

[0083] [Step S45] The additional observation unit 15 activates the optical camera 104 facing the direction of the target and captures an image of the underwater area in the direction of the sound source (true value candidate data). Then, the process proceeds to step S46.

[0084] [Step S46] The data transmission / reception unit 17 transmits to the server device 20 the data obtained by the control unit 16 through primary processing of the sonar image of the active acoustics identified by the discrimination AI 14a or the raw data of the active acoustics, together with the true value candidate data acquired in step S45 and the chemical and physical environmental data measured in step S42. Environmental DNA data acquired by the environmental DNA analysis unit 15a may also be transmitted to the server device 20. The processing of the server device 20 in the fourth process is the same as that of the server device 20 in the first process, so the explanation is omitted.

[0085] As described above, the observation device 10 of the embodiment includes a passive acoustic observation unit 11 that acquires candidate sound information emitted by organisms, artificial objects, and natural objects present in the water based on a method of passively observing sounds generated in water, and an additional observation unit 15 that performs additional observations related to the sound source based on the judgment result of the judgment unit 14.

[0086] Therefore, observations by the additional observation unit 15 are performed only when there is a high probability of obtaining useful data, thereby reducing the amount of data required to obtain training data and saving power consumption. Since the work of installing the observation device 10 underwater may involve labor such as launching a boat and having divers install it, data transmission is useful in reducing the frequency of retrieving and reinstalling the device.

[0087] By performing real-time processing within the observation device 10 using the discrimination AI 14a (edge ​​processing), the amount of data can be reduced, and a decision can be made to acquire the true value based on the processing according to the discrimination result. Furthermore, although sound propagates from various directions in 360 degrees, cameras and the like have limited directivity and distance ranges for acquiring information. Therefore, even when acquiring true value information with an optical camera 104, if information on the direction and distance of the sound source is available in advance, a decision can be made to acquire the true value based on that direction and distance information. In addition, by obtaining true value candidate data through the data transmission / reception unit 17, the server device 20 can verify the accuracy of the discrimination AI 14a and obtain new training data associated with sound sources for retraining the discrimination AI 14a. The server device 20 can also retrain the discrimination AI 14a by sending the training data generated based on the true value candidate data to the judgment unit 14 through the data transmission / reception unit 17.

[0088] Furthermore, additional observations may be performed again by the passive acoustic observation unit 11. Additional observations may also be performed by other sensors such as environmental DNA or remote sensing. Additional observations may also be performed by underwater drones or aerial drones. In addition, as mentioned above, the judgment unit 14 may classify the acoustic information obtained from active acoustic observations. In this embodiment, the judgment unit 14 has been described as having both the discrimination AI 14a and the OOD detection unit 14b. However, it is not limited to this, and may have at least one of the discrimination AI 14a and the OOD detection unit 14b.

[0089] Although the observation apparatus, observation method, and program of the present invention have been described above based on the illustrated embodiments, the present invention is not limited thereto, and the configuration of each part can be replaced with any configuration having a similar function. Furthermore, other arbitrary components or processes may be added to the present invention. Moreover, the present invention may be a combination of any two or more configurations (features) from the embodiments described above. The above merely illustrates the principle of the present invention. Furthermore, numerous modifications and changes are possible for those skilled in the art, and the present invention is not limited to the exact configurations and applications shown and described above, and all corresponding modifications and equivalents are considered to be within the scope of the present invention as defined by the appended claims and equivalents.

[0090] The above processing functions can be implemented by a computer. In this case, a program describing the processing details of the functions of the observation device 10 is provided. By executing this program on a computer, the above processing functions are implemented on the computer. The program describing the processing details can be recorded on a computer-readable recording medium. Examples of computer-readable recording media include magnetic storage devices, optical discs, magneto-optical recording media, and semiconductor memory. Examples of magnetic storage devices include hard disk drives, flexible disks (FDs), and magnetic tapes. Examples of optical discs include DVDs, DVD-RAMs, and CD-ROMs / RWs. Examples of magneto-optical recording media include MOs (Magneto-Optical disks).

[0091] When distributing a program, portable recording media such as DVDs and CD-ROMs containing the program are sold. Alternatively, the program can be stored in the storage device of a server computer and transferred from the server computer to other computers via a network.

[0092] A computer executing a program stores programs, for example, those recorded on a portable storage medium or transferred from a server computer, in its own memory. The computer then reads the program from its memory and executes the processing according to the program. Alternatively, the computer can directly read the program from the portable storage medium and execute the processing according to that program. Furthermore, the computer can sequentially execute the processing according to the programs received from a server computer connected via a network, each time a program is transferred.

[0093] Furthermore, at least some of the above processing functions can be implemented using electronic circuits such as DSPs (Digital Signal Processors), ASICs (Application Specific Integrated Circuits), and PLDs (Programmable Logic Devices).

[0094] 10 Observation device 11 Passive acoustic observation unit 12 Chemical / physical environment measurement unit 13 Memory unit 14 Judgment unit 14a Discrimination AI 14b OOD detection unit 15 Additional observation unit 15a Environmental DNA analysis unit 16 Control unit 17 Data transmission / reception unit 20 Server device 30 Network 100 Observation system 101 CPU 102 RAM 103 Internal memory 104, 104a, 104b Optical camera 105, 105a, 105b Active sonar 106, 106a, 106b, 106c, 106d Hydrophone 107 Environmental sensor 108 Acoustic modem 109 Acceleration / geomagnetic / gyro sensor 110 GPS

Claims

1. An observation device characterized by comprising: a first observation unit that acquires candidate sound information of an object to be observed present in water based on a method of passively observing sound generated in water; a determination unit that determines whether or not to perform additional observations based on the acquired candidate information; and a second observation unit that performs additional observations of the sound source based on the determination result of the determination unit.

2. The observation device according to claim 1, characterized in that the determination unit determines to perform the additional observation when it determines that the candidate information has a distribution different from a known distribution based on a set of training data including input data and corresponding output data used for training artificial intelligence.

3. The observation device according to claim 1, wherein the determination unit comprises a determination AI that performs a determination process, and the determination AI calculates the probability or score that the candidate information belongs to any of a plurality of classes that have been learned in advance, using a probability or score calculation function set during learning, and determines to perform the additional observation when it determines that the probability or score is less than a predetermined threshold.

4. The observation device according to claim 3, characterized in that the determination unit determines to perform the additional observation when the discrimination AI detects the occurrence of a specific sound designated in advance.

5. The observation device according to claim 1, further comprising a data transmission / reception unit that transmits additional observation results acquired by the second observation unit to an external server device, wherein the data transmission / reception unit is capable of receiving an updated model retrained using the additional observation results in the server device, and the determination unit is updated using the received updated model if one exists.

6. The observation device according to claim 3, characterized in that the determination unit retrains itself within the observation device using additional observation results acquired by the second observation unit, and updates the OOD detection model in which the discrimination AI or machine learning model identifies inputs having a data distribution different from the training data.

7. The observation device according to claim 5 or 6, characterized in that the determination unit can select whether or not to use the received updated model if one exists.

8. The observation device according to claim 6, characterized in that the determination unit comprises at least one of the discrimination AI and the OOD detection model.

9. The observation apparatus according to claim 1, characterized in that the second observation unit comprises an optical camera, an active sonar, a LiDAR, an environmental DNA analyzer, other observation means capable of acquiring environmental information, or one or more thereof.

10. The observation device according to claim 9, characterized in that the determination unit calculates the distance to the sound source, the second observation unit activates an optical camera and / or active sonar if the distance is less than or equal to a first threshold, and activates only the active sonar, LiDAR or other observation means if the distance is greater than the first threshold and less than or equal to a second threshold.

11. An observation method comprising the steps of passively observing sounds generated in water, determining whether additional observations are necessary based on candidate information, and performing additional observations based on the determination result.

12. A program characterized by causing a computer to perform processes for acquiring underwater sounds, determining whether additional observations are necessary, and executing additional observations.