Sound source location estimation method, learning model generation method, sound source location estimation device, and sound source location estimation system

The method and system for estimating raptor call sound sources using sound pressure level attenuation and learning models address visibility and efficiency issues in raptor surveys, enabling accurate and cost-effective identification of nesting sites and breeding success.

JP7811346B2Active Publication Date: 2026-02-05ORIENTAL CONSULTANTS +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2021098573
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-06-14
Publication Date
2026-02-05
Estimated Expiration
2041-06-14

AI Technical Summary

Technical Problem

Raptor surveys face challenges in identifying nesting sites and breeding success due to reduced visibility in obstructed forests and low encounter rates of nocturnal species, especially during limited time frames, requiring efficient and cost-effective methods to determine nesting sites and breeding success.

Method used

A method and system for estimating the sound source position of raptor calls using sound pressure level attenuation, involving multiple sound acquisition positions, analysis of audio data, and a learning model to identify call types, and solving simultaneous equations to calculate sound source positions.

Benefits of technology

Enables accurate and efficient identification of raptor nesting sites and breeding success, reducing survey costs and time, particularly effective in obstructed environments and for nocturnal species.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007811346000009
    Figure 0007811346000009
  • Figure 0007811346000010
    Figure 0007811346000010
  • Figure 0007811346000011
    Figure 0007811346000011
Patent Text Reader

Abstract

To provide a sound source position estimate method capable of estimating the sound source position of a call of a prey bird for identifying the nesting point of the prey bird.SOLUTION: The sound source position estimation method is a method for estimating the sound source position of a call of a prey bird using a computer. The sound source position estimate method includes at least an estimate step (S2) of estimating the sound source position by using the sound pressure level attenuation corresponding to the distance between the sound acquisition position and the sound source position based on the sound pressure level of the call included in the audio data acquired at each of the multiple audio acquisition points.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a sound source localization method, a learning model generation method, a sound source localization device, and a sound source localization system. [Background technology]

[0002] Nature is an essential component for a rich human life, and therefore there is a need to conserve the natural environment with the aim of achieving coexistence between humans and nature.

[0003] In the natural environment, birds of prey such as eagles, hawks, and owls are at the top of the food chain and indicate the health of the ecosystem. Protecting birds of prey leads to protecting the ecosystem in which they live. For this reason, birds of prey are selected as target species in environmental assessments conducted for public works projects, etc.

[0004] Raptor surveys are planned and conducted mainly during the breeding season. This is based on the recognition that the presence or absence of nesting and breeding success during the breeding season in and around the project site is important for predicting and evaluating the impact of the project and for considering and implementing conservation measures. In particular, identifying nesting trees in raptor nesting areas is essential for considering conservation measures during construction.

[0005] Additionally, identifying nesting core areas is important when predicting and evaluating the impact of a project in advance in environmental assessments, etc. A nesting core area is defined as "an area where mating and courtship behavior (such as vocalizations and courtship feeding) take place around the nesting trees at the nesting site, and where the birds spend time incubating, nurturing, and the young birds until they leave the nest." The nesting core area is the most important area for the breeding of birds of prey. Any alteration of the nesting core area, human intrusion during the breeding season, or construction work must be handled with caution, as they have a significant impact on breeding. In particular, in the case of the Northern Goshawk, for example, the nesting core area is estimated from the activity range of the young birds, so it is essential to understand the behavior of the young birds at the nesting site.

[0006] For example, Patent Document 1 discloses an automatic raptor abnormal behavior analysis system that directly monitors raptor nesting areas around construction sites, monitors the behavior of raptors, detects abnormal behavior, automatically determines whether the behavior is due to the effects of construction or other factors, notifies the construction manager, and takes measures to avoid crises such as nest abandonment and breeding failure. [Prior art documents] [Patent documents]

[0007] [Patent Document 1] Japanese Patent Application Laid-Open No. 2008-158745 Summary of the Invention [Problem to be solved by the invention]

[0008] In raptor surveys, to identify nesting sites, the flights of raptors are observed visually with a telescope, and possible nesting areas are narrowed down based on breeding behavior, after which the forests in those areas are surveyed to directly confirm the nesting trees. Alternatively, the behavior of young birds such as goshawks is tracked to identify central nesting areas.

[0009] However, when conducting surveys in forests where there are many obstacles to the view, visibility is significantly reduced. Furthermore, in the case of nocturnal owls, visual observation itself is difficult. Furthermore, many raptor species are designated as endangered species, and their populations are small. Therefore, the probability of encountering a raptor is low when conducting surveys within a limited time frame, and there is a high risk of missing one. When conducting raptor surveys, it is necessary to take these issues into account, ensure survey accuracy, and conduct surveys effectively.

[0010] Furthermore, large-scale public works projects such as roads and dams take a long time to complete. Raptor surveys are conducted over a 1.5-year period, covering two nesting periods, but surveys are required at each stage of the project, from planning and implementation to operation, which requires enormous costs and time. As the financial environment surrounding public works becomes increasingly severe, raptor surveys must also be conducted efficiently in terms of cost.

[0011] The main object of the present invention is to provide an apparatus, system, and method for estimating the location of the sound source of raptor calls in order to efficiently and effectively determine whether or not raptors have nested, whether or not breeding is successful, and identify nesting sites in raptor surveys. [Means for solving the problem]

[0012] The present invention provides a method for estimating the sound source position of a bird of prey's call using a computer, which includes at least an estimation step of estimating the sound source position based on the sound pressure level of the call contained in sound data acquired at each of a plurality of sound acquisition positions, using the attenuation amount of the sound pressure level corresponding to the distance between the sound acquisition position and the sound source position. The number of the sound acquisition positions may be at least four, and the estimation step may estimate the sound source position by solving a three-dimensional simultaneous equation to calculate the attenuation amount of the sound pressure level contained in each of the four sound data. The method may further include an analysis step before the estimation step, in which the analysis step determines the type of the call contained in the audio data. In the analysis stage, a learning model trained using image data containing information about the cry and training data containing information about the type of the cry may obtain the image data and determine the type. The present invention also provides a learning model generation method including at least an acquisition step of acquiring image data containing information about the calls of raptors and training data containing species information about the calls of the raptors, and a generation step of using the training data to generate a learning model that uses the image data as input and the species information as output. The present invention also provides a sound source position estimation device that includes at least an estimation unit that estimates the sound source position based on the sound pressure level of the raptor's call contained in sound data acquired at each of a plurality of sound acquisition positions, using the attenuation amount of the sound pressure level corresponding to the distance between the sound acquisition position and the sound source position. The present invention also provides a sound source location estimation system that is realized via an information and communication network and that estimates the sound source location of the calls of birds of prey, the system comprising at least a sound acquisition device that acquires sound data at each of a plurality of sound acquisition positions, and a sound source location estimation device, wherein the sound source location estimation device estimates the sound source location based on the sound pressure level of the call contained in the sound data acquired by the sound acquisition device, using an attenuation amount of the sound pressure level that corresponds to the distance between the sound acquisition position and the sound source position. [Effects of the Invention]

[0013] According to the present invention, a method, device, and system for estimating the source location of the calls of birds of prey can be provided. [Brief explanation of the drawings]

[0014] [Figure 1] 1 is an explanatory diagram illustrating the arrangement of a sound capture device used in a sound source localization method according to an embodiment of the present invention. [Figure 2] 1 is a flowchart of a sound source localization method according to an embodiment of the present invention. [Figure 3] 1 is an explanatory diagram showing planar coordinates of a sound source position and a sound acquisition position, sound pressure, etc. according to an embodiment of the present invention. [Figure 4] 1 is a flowchart of a sound source localization method according to an embodiment of the present invention. [Figure 5] 1 is a flow chart of an analysis stage according to one embodiment of the present invention. [Figure 6] 3 is an example of a spectrogram used in the sound source localization method according to the embodiment of the present invention. [Figure 7] FIG. 1 is an explanatory diagram showing the types of goshawk cries. [Figure 8] FIG. 1 is a diagram illustrating a learning model according to an embodiment of the present invention. [Figure 9] 1 is a flowchart of a learning model generation method according to an embodiment of the present invention. [Figure 10]FIG. 1 is a hardware configuration diagram of an embodiment of a computer used in the present invention. [Figure 11] 1 is a configuration diagram showing a sound source position estimation device according to an embodiment of the present invention; [Figure 12] 1 is a configuration diagram showing a sound source localization system according to an embodiment of the present invention; [Figure 13] 1 is a hardware configuration diagram of an embodiment of a voice acquisition device used in the present invention. [Figure 14] FIG. 10 is a diagram illustrating verification results of a trained model according to one embodiment of the present invention. [Figure 15] 1 is a diagram for explaining an estimation result obtained by a sound source position estimation method according to an embodiment of the present invention. [Figure 16] 1 is a diagram for explaining an estimation result obtained by a sound source position estimation method according to an embodiment of the present invention. [Figure 17] 1 is a diagram showing an estimation result obtained by a sound source localization method according to an embodiment of the present invention; [Figure 18] 1 is a diagram showing an estimation result obtained by a sound source localization method according to an embodiment of the present invention; [Figure 19] 1 is a diagram showing an estimation result obtained by a sound source localization method according to an embodiment of the present invention; DETAILED DESCRIPTION OF THE INVENTION

[0015] Preferred embodiments for implementing the present technology will be described below. The embodiments described below are examples of typical embodiments of the present technology, and the scope of the present technology will not be narrowed by them. Unless otherwise specified, in the drawings, "upper" means the upper direction or upper side in the drawing, "lower" means the lower direction or lower side in the drawing, "left" means the left direction or left side in the drawing, and "right" means the right direction or right side in the drawing. Furthermore, in the drawings, the same or equivalent elements or components are denoted by the same reference numerals, and redundant explanations will be omitted.

[0016] The present invention will be described in the following order. 1. First embodiment of the present invention (sound source position estimation method) (1) Overview (2) Estimation stage (3) Analysis stage (4) Learning model generation method (5) Hardware configuration 2. Second embodiment of the present invention (sound source position estimation device) 3. Third embodiment of the present invention (sound source localization system) 4. Working Example (1) Identifying the type of call (2) Estimating location by bird calls

[0017] <1. First embodiment of the present invention (sound source position estimation method)>

[0018] <(1) Overview> The sound source location estimation method according to the present invention is a method for estimating the sound source location of the cry of, for example, a bird of prey, using a computer. The type of bird of prey is not particularly limited, but for example, a goshawk may be selected. The goshawk is designated as a near-threatened species in the Ministry of the Environment's Red List. The goshawk is a top species in the ecosystem, and is also a species that is targeted for conservation in the "Progression of Raptor Conservation" (Ministry of the Environment, 2012), which is used for environmental impact assessments, etc.

[0019] Alternatively, for example, an owl may be selected as a type of bird of prey. Since owls are nocturnal, human survey work must also be carried out at night, which is extremely difficult.

[0020] In a sound source localization method according to the present invention, sound data including the calls of birds of prey is acquired. This will be described with reference to Fig. 1. Fig. 1 is an explanatory diagram of the arrangement of a sound acquisition device used in a sound source localization method according to one embodiment of the present invention.

[0021] As shown in Figure 1, multiple sound capture devices are placed randomly or regularly around locations where raptor nesting sites are suspected to exist. Numbers indicate the locations where the sound capture devices are placed. Stars indicate locations where raptor nesting sites are suspected to exist. There is no particular limit to the number of sound capture devices to be placed.

[0022] The voice capture device may be, for example, an IC recorder or a PCM recorder, or a highly sensitive microphone, which allows the number of voice capture devices to be reduced.

[0023] Each of the plurality of audio capture devices captures audio data containing the calls of birds of prey. The audio data may have a format with, for example, a sampling frequency of 44.1 kHz, a quantization bit rate of 16 bits, and two channels.

[0024] <(2) Estimation stage> The sound source localization method according to the present invention will be described with reference to Fig. 2. Fig. 2 is a flowchart of the sound source localization method according to an embodiment of the present invention. As shown in Fig. 2, the sound source localization method according to the present invention includes at least an estimation step (S2).

[0025] In the estimation step (S2), the sound source position is estimated based on the sound pressure level of the call contained in the sound data acquired at the same time at each of a plurality of sound acquisition positions, using the attenuation amount of the sound pressure level corresponding to the distance between the sound acquisition position and the sound source position.

[0026] The sound pressure level P contained in the audio data can be calculated, for example, by the following formula (1): p is the sound pressure for each sampling frequency, T is the length of the audio, and m is the number of sampling data.

[0027]

number

[0028] The sound pressure level P0 at the sound source position can be calculated, for example, by the following equation (2): P i is the sound pressure level at the audio acquisition position i, where audio data is acquired. i is the distance from the sound source position to the sound acquisition position i.

[0029]

number

[0030] The sound pressure p0 at the sound source position, the plane coordinates (x0, y0) of the sound source position, and the sound pressure p at the sound acquisition position i i and the plane coordinates of the audio acquisition position i (x i , y i ), the above equation (2) is transformed into the following equation (3).

[0031]

number

[0032] These plane coordinates will be explained with reference to Fig. 3. Fig. 3 is an explanatory diagram showing the plane coordinates of sound source positions and sound acquisition positions, sound pressure, etc. according to one embodiment of the present invention. As shown in Fig. 3, sound source position 0 is located in the center of the diagram, and four sound acquisition positions (1 to 4) are located around the sound source position.

[0033] The plane coordinates of sound source position 0 are (x0, y0). The sound pressure level at sound source position 0 is P0. The plane coordinates of sound acquisition position 1, which is located to the upper right of sound source position 0, are (x1, y1). The sound pressure level at sound acquisition position 1 is P1. The sound pressure at sound acquisition position 1 is p1. The distance from sound source position 0 to sound acquisition position 1 is r1. Similarly, the plane coordinates of sound acquisition position 2, which is located to the upper left of sound source position 0, are (x2, y2). The sound pressure level at sound acquisition position 2 is P2. The sound pressure at sound acquisition position 2 is p2. The distance from sound source position 0 to sound acquisition position 2 is r2. Similarly, the plane coordinates of sound acquisition position 3, which is located to the lower left of sound source position 0, are (x3, y3). The sound pressure level at sound acquisition position 3 is P3. The sound pressure at sound acquisition position 3 is p3. The distance from sound source position 0 to sound acquisition position 3 is r3. Similarly, the plane coordinates of sound acquisition position 4, which is located to the lower right of sound source position 0, are (x4, y4). The sound pressure level at sound acquisition position 4 is P4. The sound pressure at sound acquisition position 4 is p4. The distance from sound source position 0 to sound acquisition position 4 is r4.

[0034] Using these variables, the above equation (3) is transformed into the three simultaneous equations shown in the following equations (4), (5), and (6).

[0035]

number

[0036]

number

[0037]

number

[0038] By solving the simultaneous equations consisting of the above equations (4), (5), and (6), the plane coordinates (x0, y0) of the sound source position 0 can be found.

[0039] Although it is theoretically possible to estimate the plane coordinates of the sound source position from the sound data acquired at each of the three sound acquisition positions, it is preferable to acquire sound data at at least four sound acquisition positions in order to reduce errors and achieve highly accurate estimation.

[0040] <(3) Analysis stage> In order to identify the nesting sites of birds of prey, it is preferable to first analyze the audio data to extract the calls of birds of prey (for example, goshawks, etc.). Furthermore, since the types of calls made at nesting sites differ from those made at non-nesting locations, it is preferable to determine the types of calls.

[0041] Therefore, the sound source localization method according to the present invention may further include an analysis step of analyzing audio data before the estimation step (S2). This will be described with reference to Fig. 4. Fig. 4 is a flowchart of the sound source localization method according to an embodiment of the present invention. As shown in Fig. 4, the sound source localization method according to the present invention further includes an analysis step (S1) before the estimation step (S2).

[0042] Specific processing in the analysis stage (S1) will be described with reference to Fig. 5. Fig. 5 is a flowchart of an example of the analysis stage according to one embodiment of the present invention.

[0043] 5, in the analysis step (S1), first, frequency analysis using fast Fourier transform (FFT analysis) is performed on the audio data (S101). A spectrogram is created by this frequency analysis.

[0044] Here, a spectrogram will be described with reference to Fig. 6. Fig. 6 is an example of a spectrogram used in a sound source localization method according to one embodiment of the present invention. As shown in Fig. 6, a spectrogram is image data that displays audio waveform information in three dimensions of time, frequency, and intensity. In this spectrogram, the horizontal axis represents time, the vertical axis represents frequency, and color represents sound pressure. Note that Figs. 6A, 6B, and 6C will be described later.

[0045] Returning to the explanation of Figure 5, next, in the analysis stage (S1), noise is removed from this spectrogram (S102). Specifically, for example, when the main frequency range of the call of a bird of prey (such as a goshawk) is 1.0 to 6.5 kHz, noise is removed by setting the power values ​​of frequency bands outside this main frequency range to zero.

[0046] Next, in the analysis step (S1), the feature quantities of this spectrogram are obtained. Specifically, for example, while scanning the spectrogram, the convolution integral shown in the following equation (7) is performed (S103). In the following equation (7), Q i,j is the i, j component of the output spectrogram (feature map). i,j is the i, j component of the input spectrogram. m,n are the m,n components of the kernel (filter).

[0047]

number

[0048] For example, four types of kernels are used for this convolution integral. Specifically, kernel V1 (-1, -1, -1, 2, 2, 2, -1, -1, -1) detects horizontal lines, kernel V2 (2, -1, -1, -1, 2, -1, -1, -1, 2) detects left diagonal lines, kernel V3 (-1, -1, 2, -1, 2, -1, 2, -1, -1) detects right diagonal lines, and kernel V4 (-1, 2, -1, -1, 2, -1, -1, 2, -1) detects vertical lines. In addition to these, kernel V5 for smoothing may also be used.

[0049] By this convolution, the analysis stage (S1) can obtain the feature quantities of the spectrogram.

[0050] Next, in the analysis stage (S1), the size of the image data, which is a spectrogram, is reduced by pooling (S104).

[0051] Next, in the analysis stage (S1), the type of call contained in the spectrogram is determined (S105). The types of calls will be explained with reference to FIG. 7. FIG. 7 is an explanatory diagram showing the types of calls made by goshawks. As shown in FIG. 7, for example, the sound pattern for the type "alert" is "kek kek kek." This type "alert" tends to be emitted mainly by adult males.

[0052] The species will be further explained with reference again to the spectrograms shown in Figure 6. Figure 6A is an example of a spectrogram in which the species is "adult bird (alert)". Figure 6B is an example of a spectrogram in which the species is "adult bird (begging)". Figure 6C is an example of a spectrogram in which the species is "young bird".

[0053] Returning to the explanation of Figure 5, the analysis stage (S1) uses a machine-learned learning model to determine the type of bird call contained in the spectrogram (S105). As the learning model, for example, a decision tree model can be used. This learning model is trained using image data, which is a spectrogram, and training data containing information on the type of bird call.

[0054] The learning model will be described with reference to Fig. 8. Fig. 8 is a diagram showing a learning model according to one embodiment of the present invention. As shown in Fig. 8, this example uses a decision tree model, which is an example of a learning model. The root node located at the top layer contains the conditional expression "V11107>=92". In this conditional expression, the first two characters indicate the kernel type (V1 to V5). The next four characters indicate the coordinates of the audio frame. The last two characters indicate the brightness (0 to 255). In other words, the conditional expression "V11107>=92" means that when kernel V1, which detects horizontal lines, is used, the brightness at the position where the horizontal x coordinate is 11 and the vertical y coordinate is 07 is 92 or greater. If this conditional expression is met, the process proceeds to the node on the bottom left; if not, the process proceeds to the node on the bottom right. The parameters included in the conditional expression can be changed using machine learning.

[0055] In this way, the computer can obtain image data, which is a spectrogram, and determine the type of bird call using a learning model. In this embodiment, a decision tree model is used as the learning model, but this is not limiting. For example, a neural network may also be used as the learning model.

[0056] <(4) Learning model generation method> A method for generating a learning model used in the present invention will be described with reference to Fig. 9. Fig. 9 is a flowchart of a method for generating a learning model according to one embodiment of the present invention.

[0057] As shown in FIG. 9, the learning model generation method according to the present invention includes at least an acquisition stage (S3) and a generation stage (S4).

[0058] First, in the acquisition step (S3), training data for the learning model is acquired. The training data includes image data (spectrograms) containing information about the calls of raptors and information about the types of calls of raptors. This image data is preferably subjected to the above-mentioned convolution and pooling processes.

[0059] Next, in the generation step (S4), a learning model is generated using the training data, with the image data as input and the type information as output. The learning model may be, for example, a decision tree model or a neural network.

[0060] The verification method of the learning model generated in the generation stage (S4) will be explained. The verification of the learning model can be performed by calculating the accuracy rate for each type of classification using verification data (data other than the training data). For example, let n be the number of data that the learning model has classified as A, and let n be the number of data that were actually A. * In this case, the precision (correct answer rate) q when the learning model discriminates A can be calculated, for example, by the following formula (8).

[0061]

number

[0062] <(5) Hardware Configuration> The hardware configuration of a computer used in the present invention will be described with reference to Fig. 10. Fig. 10 is a diagram showing the hardware configuration of one embodiment of a computer 50 used in the present invention.

[0063] 10, the computer 50 may include, as components, a CPU 101, a storage 102, a RAM (Random Access Memory) 103, and a display 104. The components are connected to each other by, for example, a bus serving as a data transmission path.

[0064] The CPU 101 is realized by, for example, a microcomputer, and controls each component of the computer 50. The CPU 101 can perform, for example, the estimation step (S2) and the like. The estimation step (S2) and the like can be realized by, for example, a program. The CPU 101 can function by reading this program.

[0065] The storage 102 stores control data such as programs and calculation parameters used by the CPU 101. The storage 102 can be realized by using, for example, a hard disk drive (HDD) or a solid state drive (SSD), etc. The storage 102 holds, for example, audio data, image data, etc.

[0066] The RAM 103 temporarily stores, for example, programs executed by the CPU 101.

[0067] The display 104 displays information. For example, the display 104 can display the location of a sound source. The display 104 can be realized by, for example, an LCD (Liquid Crystal Display) or an OLED (Organic Light-Emitting Diode).

[0068] Although not shown, the computer 50 may also include a communication interface. This communication interface has a function of communicating via an information and communication network using communication technologies such as Wi-Fi, Bluetooth (registered trademark), and LTE (Long Term Evolution). For example, the CPU 101 can estimate the position of a sound source based on audio data obtained via this communication interface.

[0069] The computer 50 may be, for example, a server, a smartphone terminal, a tablet terminal, a mobile phone terminal, a PDA (Personal Digital Assistant), a PC (Personal Computer), a portable music player, a portable game console, or a wearable terminal (HMD: Head Mounted Display, glasses-type HMD, watch-type terminal, band-type terminal, etc.).

[0070] The program for realizing the estimation step (S2) etc. may be stored in a computer device or computer system other than the computer 50. In this case, the computer 50 can use a cloud service that provides the functions of this program. Examples of such cloud services include SaaS (Software as a Service), IaaS (Infrastructure as a Service), PaaS (Platform as a Service), etc.

[0071] Furthermore, this program can be stored and supplied to a computer using various types of non-transitory computer-readable media. Non-transitory computer-readable media include various types of tangible storage media. Examples of non-transitory computer-readable media include magnetic storage media (e.g., flexible disks, magnetic tapes, hard disk drives), magneto-optical storage media (e.g., magneto-optical disks), Compact Disc Read Only Memory (CD-ROM), CD-R, CD-R / W, and semiconductor memory (e.g., mask ROM, programmable ROM (PROM), erasable PROM (EPROM), flash ROM, random access memory (RAM)). The program may also be supplied to a computer by various types of transitory computer-readable media. Examples of transitory computer-readable media include electrical signals, optical signals, and electromagnetic waves. The transitory computer-readable medium can supply the program to a computer via a wired communication path such as an electric wire or optical fiber, or via a wireless communication path.

[0072] In addition to this, the configurations given in the above embodiments can be selected or changed to other configurations as appropriate without departing from the spirit of the present technology.

[0073] 2. Second embodiment of the present invention (sound source position estimation device) A sound source localization device according to an embodiment of the present invention will be described with reference to Fig. 11. Fig. 11 is a configuration diagram showing a sound source localization device according to an embodiment of the present invention. As shown in Fig. 11, a sound source localization device 20 according to an embodiment of the present invention includes at least an estimation unit 21.

[0074] The estimation unit 21 estimates the sound source position based on the sound pressure level of the raptor's call contained in the sound data acquired at each of multiple sound acquisition positions, using the attenuation amount of the sound pressure level corresponding to the distance between the sound acquisition position and the sound source position.

[0075] The estimation unit 21 can perform the estimation step (S2) described in the first embodiment, and therefore, detailed description thereof will be omitted.

[0076] The hardware of the sound source localization device 20 may be the same as the hardware of the computer 50 described in the first embodiment, so a detailed description will not be given again.

[0077] 3. Third Embodiment of the Present Invention (Sound Source Localization System) A sound source localization system according to one embodiment of the present invention is realized via an information and communication network, and is a sound source localization system capable of estimating the sound source location of the calls of birds of prey.

[0078] A sound source localization system according to an embodiment of the present invention will be described with reference to Fig. 12. Fig. 12 is a configuration diagram showing a sound source localization system according to an embodiment of the present invention. As shown in Fig. 12, a sound source localization system 100 according to an embodiment of the present invention includes at least a sound acquisition device 10 that acquires sound data at each of a plurality of sound acquisition positions, and a sound source localization device 20.

[0079] A plurality of voice capturing devices 10 and a sound source localization device 20 are connected via an information and communication network 30. Note that there may be voice capturing devices 10 that are not connected to the sound source localization device 20. Furthermore, not all voice capturing devices 10 may be connected to the sound source localization device 20.

[0080] The sound source position estimation device 20 estimates the sound source position based on the sound pressure level of the cry contained in the sound data acquired by the sound acquisition device 10, using the attenuation amount of the sound pressure level corresponding to the distance between the sound acquisition position and the sound source position.

[0081] The sound source localization device 20 can perform the estimation step (S2) described in the first embodiment, and therefore a detailed description thereof will be omitted.

[0082] Each of the plurality of voice capturing devices 10 can have some or all of the functions of the sound source localization device 20. For example, each of the plurality of voice capturing devices 10 can have an estimation unit 21.

[0083] Similarly, the sound source localization device 20 can have some or all of the functions of each of the plurality of sound capturing devices 10. For example, the sound source localization device 20 can have a sound capturing function.

[0084] Here, the hardware configuration of the voice capturing device 10 will be described with reference to Fig. 13. Fig. 13 is a diagram showing the hardware configuration of one embodiment of the voice capturing device 10 used in the present invention.

[0085] 13, the speech capturing device 10 may include, as its components, a speech capturing unit 1001, a storage unit 1002, and a control unit 1003. The components are connected to each other by, for example, a bus serving as a data transmission path.

[0086] The voice acquisition unit 1001 acquires voice and can be realized by, for example, a microphone.

[0087] The storage unit 1002 stores, as audio data, the audio acquired by the audio acquisition unit 1001. The storage unit 1002 can be realized by using, for example, a hard disk drive (HDD) or a solid state drive (SSD).

[0088] The control unit 1003 controls each of the components of the voice capturing device 10. The control unit 1003 can be realized by, for example, a microcomputer.

[0089] Although not shown in the figures, the voice capturing device 10 may include a current location acquisition unit. This location information acquisition unit has a function of detecting the current location of the voice capturing device 10 based on an acquired signal from an external device. Specifically, the location information acquisition unit is realized by, for example, a GPS (Global Positioning System) positioning unit, and receives radio waves from a GPS satellite to detect the location where the location information acquisition unit is located. Alternatively, the location information acquisition unit may detect the location by, for example, Wi-Fi (registered trademark), transmission and reception with a mobile phone, PHS, smartphone, etc., or short-range communication, in addition to GPS.

[0090] Although not shown in the figure, the voice capturing device 10 may include a communication interface. Voice data, position information, and the like captured by the voice capturing device 10 can be transmitted to the sound source position estimation device 20 via this communication interface.

[0091] The hardware of the sound source localization device 20 may be the same as the hardware of the computer 50 described in the first embodiment, so a detailed description will not be given again.

[0092] The effects described in this specification are merely examples and are not limiting, and other effects may also be present.

[0093] <4. Example> <(1) Identifying the type of call> As shown in Figure 1, IC recorders (SONY, ICD-UX560F) were placed at 16 locations around suspected goshawk nesting sites. The recording format for each IC recorder was linear PCM with a sampling frequency of 44.1 kHz, quantization bit rate of 16, and two channels.

[0094] Among the multiple audio data recorded, the audio data recorded by the IC recorder placed at point 10 in Figure 1 contained particularly clear and frequent calls of the Northern Goshawk during a specific time period (6:00 AM to 8:00 AM on July 9, 2018). Therefore, audio frames such as those shown in Figure 7 were extracted from the audio data during this time period. Furthermore, when calls were continuous, such as for the species "adult bird (alert)," audio segments consisting of multiple audio frames were also extracted.

[0095] Based on the audio frames, the number of data extracted from the audio data was 119 for the species "adult bird (alert)", 179 for the species "adult bird (begging)", and 426 for the species "young bird", for a total of 724.

[0096] Additionally, frequency analysis (FFT analysis) was performed based on the audio data from this time period, and the spectrogram shown in Figure 6 was created. To obtain the characteristics of the created spectrogram, the convolution shown in equation (7) above was performed. Five kernel patterns were used for the convolution: kernel V1 (-1, -1, -1, 2, 2, 2, -1, -1, -1) for detecting horizontal lines, kernel V2 (2, -1, -1, -1, 2, -1, -1, -1, 2) for detecting left diagonal lines, and kernel V5 for smoothing.

[0097] Next, the size of the spectrogram image data was reduced to 11 × 33 pixels by pooling.

[0098] Next, the type of calls contained in the spectrogram was determined. A decision tree model, an example of a machine learning model, was used to make the determination. The objective variables of the decision tree model were three types of calls heard during the nest-rearing period of goshawks: "adult bird (alert)," "adult bird (begging)," and "juvenile." The explanatory variables of the decision tree model were the values ​​for each pixel in the image data after convolution and pooling processes.

[0099] Of the 724 pieces of data extracted from the speech data, half (362 pieces) were used as training data, and the other half (362 pieces) were used as validation data. The training data were used as training data to train a decision tree model.

[0100] The precision (accuracy rate) q of the judgment result by the decision tree model, which is a trained model, obtained by the above formula (8) will be described with reference to FIG. 14. FIG. 14 is a diagram showing the verification results of a trained model according to one embodiment of the present invention. FIG. 14A shows the verification results using training data. FIG. 14B shows the verification results using verification data.

[0101] 14A and 14B, the items arranged horizontally are the types of calls determined by the decision tree model, and the items arranged vertically are the correct (actual) types of calls.

[0102] As shown in Figure 14A, for example, for the species "adult bird (alert)," the decision tree model determined that 46 of the 51 pieces of data were "adult bird (alert)," and the correct species was found to be the "adult bird (alert)." In this case, the precision rate (correctness rate) of the determination, calculated using the above formula (8), was approximately 90.2%. Similarly, the precision rate for the species "adult bird (begging)" was 87.5%, and for the species "young bird," the precision rate was approximately 83.4%. The precision rate was high because the training data was training data used for learning.

[0103] Referring to Figure 14B, which shows the verification results using the verification data, the precision rate for the species "adult bird (alert)" was approximately 67.9%, the precision rate for the species "adult bird (begging)" was approximately 53.0%, and the precision rate for the species "young bird" was approximately 77.4%.

[0104] In this example, the matching rates for the types "adult bird (alert)" and "young bird" were particularly high.

[0105] (2) Estimating location by bird calls Frequency analysis was performed using the fast Fourier transform (FFT) on the sound data from all 16 locations over all time periods. Next, to remove noise, the power values ​​of frequency bands other than 1.0 to 6.5 kHz, the main frequency range of goshawk calls, were set to zero.

[0106] Next, to extract the goshawk's call, the created spectrogram was scanned, and the values ​​for each pixel after convolution and pooling were input into a decision tree model as explanatory variables. Note that the species "adult bird (alert)" was extracted for each segment, so consecutive segments were aggregated together as an audio frame.

[0107] Next, the data was restored to audio data by performing an inverse fast Fourier transform (inverse FFT), and the sound pressure level of the goshawk's call was calculated using the above equation (1).

[0108] Finally, the sound source position was estimated using the above equations (2) to (6). The estimation results will be explained with reference to Fig. 15 and Fig. 16. Fig. 15 and Fig. 16 are diagrams for explaining the estimation results obtained by a sound source position estimation method according to one embodiment of the present invention. Fig. 15 shows the results estimated by a computer, and Fig. 16 shows the results of a field survey conducted by a human. The positions indicated by ovals in Fig. 16 are positions where the calls of goshawks were confirmed.

[0109] As shown in Figure 15, the estimated sound source locations are plotted for each type of call. The estimated locations of nesting sites are plotted with stars. Comparing this with Figure 16, the locations of the nesting sites are roughly consistent.

[0110] According to the present invention, it is also possible to track the activity range of young birds by time during the out-of-nest nest breeding period. This will be described with reference to Figures 17, 18, and 19. Figures 17, 18, and 19 are diagrams showing the estimation results using a sound source location estimation method according to one embodiment of the present invention. Figure 17 shows the estimated sound source location from 6:00 AM to 6:30 AM on July 9, 2018, Figure 18 shows the estimated sound source location from 6:30 AM to 7:00 AM on the same day, and Figure 19 shows the estimated sound source location from 7:00 AM to 7:30 AM on the same day. As shown in Figures 17, 18, and 19, it can be seen that the estimated sound source location of young birds moves over time. This makes it possible to track the activity range of raptors.

[0111] The present invention can also be configured as follows. [1] A method for estimating the source location of a bird of prey call using a computer, comprising: A sound source position estimation method comprising at least an estimation step of estimating the sound source position based on the sound pressure level of the bird's cry contained in sound data acquired at each of a plurality of sound acquisition positions, using an attenuation amount of the sound pressure level corresponding to the distance between the sound acquisition position and the sound source position. [2] the number of said audio capture locations is at least four; the estimating step estimates the sound source position by solving simultaneous equations with three unknowns for calculating the attenuation amount of the sound pressure level included in each of the four pieces of audio data; The sound source localization method according to [1]. [3] further comprising an analysis step prior to said estimation step, the analyzing step determines the type of the cry contained in the audio data; The sound source localization method according to [1] or [2]. [4] In the analysis step, a learning model trained using image data including information about the bird's cry and training data including information about the bird's cry type obtains the image data and determines the type. The sound source localization method according to [3]. [5] an acquisition step of acquiring image data including information about the calls of birds of prey and teacher data including type information of the calls of the birds of prey; and a generation step of generating a learning model using the training data, the learning model having the image data as input and the type information as output. Learning model generation method. [6] A sound source position estimation device including at least an estimation unit that estimates the sound source position based on the sound pressure level of the calls of birds of prey contained in sound data acquired at each of a plurality of sound acquisition positions, using the attenuation amount of the sound pressure level corresponding to the distance between the sound acquisition position and the sound source position. [7] A sound source location estimation system that is realized via an information and communication network and estimates the sound source location of a bird of prey cry, a voice capture device that captures voice data at each of a plurality of voice capture positions; a sound source location estimation device; A sound source position estimation system in which the sound source position estimation device estimates the sound source position based on the sound pressure level of the cry contained in the sound data acquired by the sound acquisition device, using the attenuation amount of the sound pressure level corresponding to the distance between the sound acquisition position and the sound source position. [Explanation of symbols]

[0112] S1 Analysis stage S2 Estimation stage S3 Acquisition stage S4 Generation stage 10. Audio capture device 20 Sound source position estimation device 30 Information and Communications Networks 100 Sound Source Localization System

Claims

1. A method for estimating the source location of a bird of prey call using a computer, comprising: an analysis step of determining the type of the call contained in the audio data; and an estimation step of estimating the sound source position by solving a three-dimensional simultaneous equation for calculating an attenuation amount of the sound pressure level corresponding to a distance between the sound acquisition position and the sound source position, based on the sound pressure level of the sound included in the sound data acquired at each of at least four sound acquisition positions, for the specific type of sound determined in the analysis step. A sound source location estimation method in which, in the analysis stage, a learning model trained using image data containing information about the bird cry and training data containing information about the bird cry type obtains the image data and determines the type.

2. The bird of prey is a goshawk or an owl. The sound source position estimation method according to claim 1 .

3. the analyzing step further includes a process of removing noise from the image data and a process of performing a convolution integral. The sound source position estimation method according to claim 1 or 2.

4. The method further comprises identifying a central nesting area or a range of activity of the raptor based on the type of call determined in the analysis step and the estimated sound source location. The sound source position estimation method according to claim 1 or 2.

5. an estimation unit that estimates the sound source position based on the sound pressure level of the bird of prey call included in the sound data acquired at each of a plurality of sound acquisition positions, using an attenuation amount of the sound pressure level corresponding to the distance between the sound acquisition position and the sound source position; an analysis unit that determines the type of the bird cry included in the audio data, the analysis unit obtains the image data and determines the type using a learning model trained using image data including information about the cry and training data including information about the type of the cry; A sound source position estimation device, wherein the estimation unit estimates the sound source position by solving a ternary simultaneous equation to calculate the attenuation amount of the sound pressure level contained in each of the four sound data at at least four sound acquisition positions for the specific type of call determined by the analysis unit.

6. A sound source location estimation system that is realized via an information and communication network and estimates the sound source location of a bird of prey cry, a voice capture device that captures voice data at each of a plurality of voice capture positions; a sound source location estimation device; A sound source localization system, wherein the sound source localization device is the sound source localization device according to claim 5 .

Citation Information

Patent Citations

  • Apparatus and method for recognizing song of wild bird

    JP2003255984A

  • Automatic abnormal bird behavior analysis system and construction work management method using the same

    JP2008158745A

  • Determination / distribution system of position and kind of wild animal

    JP2018099114A

  • Specifying device

    JP2019124513A

  • Position estimation device and position estimation method

    JP2020159705A