Event recognition device, event recognition method, and program
By using a deep learning model trained with acoustic signals and performing SNR improvement on intermediate optical fiber signals, the low SNR issue is addressed, enabling effective event recognition without model training.
Patent Information
- Application Number
- JP2024028067
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-28
- Publication Date
- 2025-09-09
AI Technical Summary
Optical fiber sensing signals have a low signal-to-noise ratio (SNR) due to high optical noise, making it difficult to perform model training and generate effective deep learning models for event recognition.
Utilize a deep learning model trained with acoustic signals from acoustic sensors to recognize events, and perform SNR improvement processing on intermediate signals with a time-frequency structure to bridge the SNR gap between optical fiber and acoustic signals.
Enables event recognition without model learning, improving recognition performance by adapting the optical fiber signals to the higher SNR of acoustic signals.
Smart Images

Figure 2025130785000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an event recognition device, an event recognition method, and a program. [Background technology]
[0002] Optical fiber sensing, typified by Distributed Acoustic Sensing (DAS), is capable of detecting sound and vibrations generated at points along an optical fiber.
[0003] In recent years, a technology has been proposed that acquires observation signals indicating sound or vibration that occurs at a point along an optical fiber and is detected by optical fiber sensing, and recognizes an event that has occurred at a point along the optical fiber (for example, a car traveling) based on the acquired observation signals.
[0004] Furthermore, a technology for recognizing an event based on an observation signal acquired by optical fiber sensing includes a technology for performing model learning on the observation signal to generate a deep learning model and recognizing an event using the generated deep learning model. For example, Patent Document 1 discloses a technology for detecting an acoustic alarm pattern event using a deep learning system. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Special Publication No. 2023-538196 Summary of the Invention [Problem to be solved by the invention]
[0006] When using a deep learning model for optical fiber sensing, observed signals acquired by optical fiber sensing are required for model learning. However, the observed signals acquired by optical fiber sensing have a low signal-to-noise ratio (SNR) due to the inclusion of a large amount of optical noise (e.g., highly white shot noise and quantum noise). As a result, there are few observed signals that are effective for model training, making it difficult to perform model training using the observed signals and generate a deep learning model.
[0007] Therefore, there is a demand for a technology that can recognize events using observed signals acquired by optical fiber sensing without performing model learning.
[0008] In view of the above-mentioned problems, an object of the present disclosure is to provide an event recognition device, an event recognition method, and a program that are capable of recognizing events without performing model learning using observation signals acquired by optical fiber sensing. [Means for solving the problem]
[0009] An event recognition apparatus according to one aspect includes: an acquisition unit that acquires an observation signal indicative of sound generated at a point along the optical fiber and detected by the optical fiber sensing; a deep learning model that is trained using an acoustic signal acquired by an acoustic sensor, and that receives an input signal indicating a sound occurring at the location and outputs a recognition result of an event occurring at the location; a recognition unit that inputs the observation signal to the deep learning model as the input signal, obtains the recognition result as an output of the deep learning model, and outputs the recognition result; and a first processing unit that performs a first process on an intermediate signal of a time-frequency structure obtained in an intermediate layer inside the deep learning model when the observed signal is input to the deep learning model, to improve the signal-to-noise ratio of the intermediate signal.
[0010] An event recognition method according to one aspect includes: An event recognition method executed by an event recognition device, comprising: acquiring an observation signal indicative of acoustics generated at a point along the optical fiber and detected by the optical fiber sensing; inputting the observed signal as an input signal to a deep learning model that is trained using an acoustic signal acquired by an acoustic sensor and that receives an input signal indicating an acoustic occurring at the location and outputs a recognition result of an event occurring at the location; performing a first process for improving a signal-to-noise ratio of an intermediate signal having a time-frequency structure obtained in an intermediate layer inside the deep learning model when the observation signal is input to the deep learning model; obtaining the recognition result as an output of the deep learning model when the observed signal is input, and outputting the recognition result; Includes:
[0011] In one aspect, the program comprises: On the computer, acquiring observed signals indicative of acoustics generated at points along the optical fiber and detected by the optical fiber sensing; a step of inputting the observed signal as an input signal to a deep learning model that is a model trained using an acoustic signal acquired by an acoustic sensor, the deep learning model receiving an input signal indicating an acoustic occurring at the location and outputting a recognition result of an event occurring at the location; a step of performing a first process on an intermediate signal having a time-frequency structure obtained in an internal intermediate layer of the deep learning model when the observation signal is input to the deep learning model, the first process improving the signal-to-noise ratio of the intermediate signal; obtaining the recognition result as an output of the deep learning model when the observed signal is input, and outputting the recognition result; Execute the following. [Effects of the Invention]
[0012] According to the above-described aspects, it is possible to provide an event recognition device, an event recognition method, and a program that are capable of recognizing events without performing model learning using observation signals acquired by optical fiber sensing. [Brief explanation of the drawings]
[0013] [Figure 1] FIG. 10 is a diagram illustrating an example of a method for obtaining event recognition results using a deep learning model trained on acoustic signals acquired by an acoustic sensor. [Figure 2] This figure shows an example of an intermediate signal with a time-frequency structure obtained in an intermediate layer inside a deep learning model when an observation signal obtained by optical fiber sensing is input into the deep learning model. [Figure 3] 3A to 3C are diagrams illustrating an example of a process for improving the SNR of an intermediate signal shown in FIG. 2. [Figure 4] 1 is a block diagram illustrating a schematic configuration example of an event recognition device according to the present disclosure. [Figure 5] FIG. 1 is a flow diagram illustrating an example of a schematic operation flow of an event recognition device according to the present disclosure. [Figure 6] 1 is a diagram illustrating an example of an event recognition result obtained by an event recognition device according to the present disclosure, in comparison with an event recognition result obtained by a related art technique. [Figure 7] 1 is a block diagram illustrating a schematic configuration example of an event recognition device according to the present disclosure. [Figure 8] FIG. 10 is a diagram showing an example of an intermediate signal with a time-frequency structure obtained in an intermediate layer inside a deep learning model when a DAS observation signal is input to the deep learning model in an event recognition device according to the present disclosure. [Figure 9] FIG. 10 is a diagram showing an example of an intermediate signal with a time-frequency structure obtained in an intermediate layer inside a deep learning model when a silent observation signal is input to the deep learning model in an event recognition device according to the present disclosure. [Figure 10] 9 is a diagram illustrating an example of processing for improving the SNR of the intermediate signal shown in FIG. 8 in the event recognition device according to the present disclosure. [Figure 11] FIG. 1 is a flow diagram illustrating an example of a schematic operation flow of an event recognition device according to the present disclosure. [Figure 12] 1 is a block diagram illustrating a schematic configuration example of an event recognition device according to the present disclosure. [Figure 13] FIG. 1 is a flow diagram illustrating an example of a schematic operation flow of an event recognition device according to the present disclosure. [Figure 14] 1 is a block diagram illustrating a schematic configuration example of an event recognition device according to the present disclosure. [Figure 15] FIG. 1 is a block diagram illustrating a schematic hardware configuration example of a computer that realizes an event recognition device according to the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0014] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. Note that the following description and drawings have been omitted and simplified as appropriate for clarity of explanation. In addition, in each of the following drawings, the same elements are given the same reference numerals, and duplicate explanations are omitted as necessary. Furthermore, specific numerical values etc. shown below are merely examples to facilitate understanding of the present disclosure, and are not limited thereto.
[0015] <Outline of each embodiment> Before describing the details of each embodiment of the present disclosure, an overview of each embodiment will be described. As mentioned above, the observed signals acquired by optical fiber sensing have a low SNR due to the inclusion of a large amount of optical noise. Therefore, there are few observed signals that are effective for model training, making it difficult to perform model training using the observed signals and generate a deep learning model.
[0016] Meanwhile, in recent years, deep learning models trained using acoustic signals acquired by acoustic sensors such as microphones have become common on the market. Therefore, by using deep learning models that are available on the market and have been trained using acoustic signals acquired by acoustic sensors, model training using observed signals acquired by optical fiber sensing becomes unnecessary.
[0017] FIG. 1 is a diagram illustrating an example of a method for obtaining an event recognition result using a deep learning model trained on acoustic signals acquired by an acoustic sensor. As shown in Fig. 1, a deep learning model 80 trained using an acoustic signal acquired by an acoustic sensor is prepared. The deep learning model 80 is a model that learns events occurring at a certain point using, as input, an acoustic signal that indicates sound generated at that point and that is acquired by the acoustic sensor.
[0018] To obtain a recognition result of an event occurring at a point along an optical fiber, an observation signal representing the sound occurring at that point is input to the deep learning model 80. For example, the observation signal may represent sound such as a dog barking, a person talking, or the sound of a car driving.
[0019] Then, an event recognition result is obtained as an output of the deep learning model 80. For example, the event may be a dog barking, a person talking, a car driving, etc.
[0020] However, as mentioned above, the observed signals acquired by optical fiber sensing have a low SNR due to the inclusion of a large amount of optical noise, whereas the acoustic signals acquired by acoustic sensors such as microphones have a high SNR because they do not include optical noise.
[0021] Therefore, there is a gap in SNR between the observed signal and the acoustic signal, and therefore, when the observed signal is input to the deep learning model 80 trained using the acoustic signal, there is a concern that the event recognition performance may be degraded due to the low SNR of the observed signal. Therefore, even when an observed signal is input to the deep learning model 80, in order to improve the event recognition performance, it is necessary to fill the SNR gap between the observed signal and the acoustic signal.
[0022] Therefore, the inventors focus on the intermediate signal of the time-frequency structure obtained in the internal intermediate layer of the deep learning model 80, and perform processing to improve the SNR of this intermediate signal of the time-frequency structure.
[0023] FIG. 2 is a diagram showing an example of an intermediate signal with a time-frequency structure obtained in an intermediate layer inside deep learning model 80 when an observation signal acquired by optical fiber sensing is input to deep learning model 80.
[0024] The intermediate signal shown in Figure 2 is a signal with a time-frequency structure, where the horizontal axis is time and the vertical axis is frequency, and it shows the time variation of the frequency component of the observed signal. embed is the embedding dimension of the hidden layer inside the deep learning model 80. embed This shows an example where the value is "3".
[0025] FIG. 3 is a diagram illustrating an example of an SNR improvement process for improving the SNR of the intermediate signal shown in FIG. As shown in Fig. 3, SNR improvement processing is performed to improve the SNR of intermediate signals with a time-frequency structure obtained in an internal intermediate layer of the deep learning model 80. This SNR improvement processing is performed on each intermediate signal arranged in the embedding direction. This SNR improvement processing may also be noise suppression processing that suppresses noise in the intermediate signals.
[0026] This allows the domain of the observed signal to be adapted to the domain of the acoustic signal, thereby bridging the SNR gap between the observed signal and the acoustic signal, thereby improving the event recognition performance even when the observed signal is input to the deep learning model 80. Hereinafter, each embodiment of the present disclosure will be described.
[0027] <First Embodiment> FIG. 4 is a block diagram showing a schematic configuration example of the event recognition device 10. As shown in FIG. As shown in FIG. 4, the event recognition device 10 includes an acquisition unit 11 and a recognition unit 12.
[0028] The acquisition unit 11 acquires an observation signal indicating sound (e.g., dog barking, human talking, car running sound, etc.) generated at a point along the optical fiber and detected by optical fiber sensing. In the first embodiment, the acquisition unit 11 acquires a DAS observation signal from a DAS device as the observation signal. For example, the DAS observation signal is a time domain signal R that indicates the time variation of the intensity of sound generated at a point along the optical fiber. T Alternatively, the DAS observation signal may be expressed as its time domain signal R T The frequency domain signal C after Fourier transform of T×F It may be.
[0029] The recognition unit 12 holds a deep learning model 13. The deep learning model 13 is a model available on the market that is trained using acoustic signals acquired by an acoustic sensor such as a microphone. In detail, the deep learning model 13 is a model that learns events occurring at a certain location using, as input, an acoustic signal that indicates a sound generated at that location and that is acquired by an acoustic sensor.
[0030] Furthermore, the deep learning model 13 is a model that, when an input signal indicating a sound generated at a certain point is input to the input layer, outputs a recognition result of an event occurring at that point from the output layer. For example, the deep learning model 13 may output events such as a dog barking, a person talking, or a car driving as the event recognition result. Alternatively, the deep learning model 13 may output the probability that at least one event has occurred at that point, and output the event with the highest probability among these as the event recognition result.
[0031] The recognition unit 12 inputs the DAS observation signal acquired by the acquisition unit 11 as an input signal to the deep learning model 13. Then, the recognition unit 12 obtains a recognition result of an event occurring at that point as an output of the deep learning model 13. Furthermore, when the recognition unit 12 obtains the event recognition result, it outputs the recognition result to the outside of the event recognition device 10.
[0032] The recognition unit 12 also includes a first processing unit 14. The first processing unit 14 performs SNR improvement processing (see FIG. 3) on an intermediate signal (see FIG. 2) having a time-frequency structure obtained in an intermediate layer inside the deep learning model 13 when the DAS observation signal is input to the deep learning model 13, to improve the SNR of the intermediate signal. For example, the SNR improvement processing may be noise suppression processing that suppresses noise in the intermediate signal. Furthermore, for example, the noise suppression processing may be well-known noise suppression processing such as Wiener filtering, spectral subtraction, frequency filtering, or cepstrum analysis.
[0033] Next, a schematic example of the operation of the event recognition device 10 will be described. FIG. 5 is a flow diagram illustrating an example of a schematic operation flow of the event recognition device 10. As shown in FIG.
[0034] As shown in FIG. 5, the acquisition unit 11 acquires, from the DAS device, a DAS observation signal indicating sound generated at a point along the optical fiber and detected by optical fiber sensing (step S11).
[0035] Next, the recognition unit 12 inputs the DAS observation signal acquired by the acquisition unit 11 to the deep learning model 13 as an input signal (step S12). Next, the first processing unit 14 performs an SNR improvement process on the intermediate signal of the time-frequency structure obtained in the intermediate layer inside the deep learning model 13 to improve the SNR of the intermediate signal (step S13).
[0036] Thereafter, the recognition unit 12 obtains a recognition result of the event occurring at the above-mentioned location as an output of the deep learning model 13, and outputs the recognition result to the outside of the event recognition device 10 (step S14).
[0037] Next, the results of event recognition by the event recognition device 10 will be described. FIG. 6 is a diagram illustrating an example of an event recognition result by the event recognition device 10 in comparison with an event recognition result by the related art.
[0038] The top row of Figure 6 shows four simulated signals that simulate the DAS observation signal, which indicates the same sound (here, bird song) generated at the same point along the optical fiber. The four simulated signals are the frequency domain signal C T×F In each diagram, the horizontal axis represents time and the vertical axis represents frequency. The four simulated signals are more similar to the DAS observation signal (i.e., less similar to the acoustic signal acquired by the acoustic sensor) and have a lower SNR as they move to the right in the diagram. In other words, the four simulated signals are more dissimilar to the DAS observation signal (i.e., more similar to the acoustic signal) and have a higher SNR as they move to the left in the diagram.
[0039] The middle part of Figure 6 shows the event recognition results when the four simulated signals are input as input signals to a deep learning model according to the related technology. The related technology inputs the input signals to the deep learning model, and then obtains the output of the deep learning model as the event recognition result without performing SNR improvement processing on the intermediate signals in the intermediate layer of the deep learning model.
[0040] 6 shows the event recognition results when the four simulated signals are input as input signals to the deep learning model 13 of the event recognition device 10 according to the present disclosure. As described above, the event recognition device 10 inputs the input signals to the deep learning model 13, then performs SNR improvement processing on the intermediate signals in the intermediate layer of the deep learning model 13, and obtains the output of the deep learning model 13 as the event recognition result.
[0041] In addition, in the middle and lower diagrams of Figure 6, the horizontal axis indicates the event class, and the vertical axis indicates the probability that an event of each event class has occurred. In this way, the deep learning model 13 according to the present disclosure in Figure 6 outputs, as an event recognition result, the probability that an event of each event class has occurred at that point, and further outputs the event with the highest probability among these. The same is true for deep learning models according to related technologies.
[0042] As shown in Figure 6, the related technology is able to recognize the event as "chirping birds" in the case of the two simulated signals on the left side of the figure, which are dissimilar to the DAS observation signals and have a high SNR. However, in the case of the two simulated signals on the right side of the figure, which are similar to the DAS observation signals but have a low SNR, the related technology recognizes the event as "rain" and fails to recognize the event as "birds chirping."
[0043] In contrast, the event recognition device 10 was able to recognize that the event for all four simulated signals was "birds singing." In this way, it can be seen that the event recognition device 10 can improve the event recognition performance.
[0044] As described above, according to the first embodiment, the acquisition unit 11 acquires a DAS observation signal indicating sound generated at a point along an optical fiber and detected by optical fiber sensing. The recognition unit 12 inputs the DAS observation signal as an input signal to a deep learning model 13 that has been trained using the acoustic signal acquired by the acoustic sensor. The first processing unit 14 performs SNR improvement processing on an intermediate signal with a time-frequency structure obtained in an intermediate layer inside the deep learning model 13 when the DAS observation signal is input to the deep learning model 13, thereby improving the SNR of the intermediate signal. The recognition unit 12 obtains a recognition result of an event occurring at the above-mentioned point as an output of the deep learning model 13, and outputs the recognition result.
[0045] As described above, according to the first embodiment, the deep learning model 13 trained using the acoustic signal acquired by the acoustic sensor is used, which eliminates the need for model training using the observation signal acquired by optical fiber sensing.
[0046] Furthermore, according to the first embodiment, an SNR improvement process is performed on an intermediate signal having a time-frequency structure obtained in an internal intermediate layer of the deep learning model 13, to improve the SNR of the intermediate signal. This makes it possible to fill the SNR gap between the DAS observation signal and the acoustic signal. Therefore, even when the DAS observation signal is input to the deep learning model 13, it is possible to improve the event recognition performance.
[0047] <Embodiment 2> FIG. 7 is a block diagram showing a schematic configuration example of the event recognition device 10A. As shown in FIG. 7, the event recognition device 10A includes an acquisition unit 11A and a recognition unit 12A.
[0048] The acquisition unit 11A acquires, from the DAS device, a DAS observation signal that indicates sound generated at a point along the optical fiber and detected by optical fiber sensing. Furthermore, the acquisition unit 11A acquires, from the DAS device, a silent observation signal, which is a DAS observation signal at an arbitrary point in a silent state. The DAS observation signal in the second embodiment is a signal when some kind of sound is generated at the above-mentioned point, and is to be distinguished from a silent observation signal.
[0049] The recognition unit 12A holds a deep learning model 13A. The deep learning model 13A is similar to the deep learning model 13.
[0050] The recognition unit 12A inputs the DAS observation signal acquired by the acquisition unit 11A to the deep learning model 13A as an input signal. The recognition unit 12A also inputs the silent observation signal acquired by the acquisition unit 11A to the deep learning model 13A as an input signal. The silent observation signal is used in SNR improvement processing by the first processing unit 14A, which will be described later. The recognition unit 12A then obtains a recognition result of an event occurring at that point as an output of the deep learning model 13A, and outputs the recognition result to the outside of the event recognition device 10A.
[0051] The recognition unit 12A also includes a first processing unit 14A. The first processing unit 14A uses an intermediate signal of a time-frequency structure obtained in an intermediate layer inside the deep learning model 13 when a silent observation signal is input to the deep learning model 13A to perform SNR improvement processing on the intermediate signal of a time-frequency structure obtained in an intermediate layer inside the deep learning model 13 when a DAS observation signal is input to the deep learning model 13A.
[0052] Here, the SNR improvement process performed by the first processing unit 14A will be described in detail. As described above, in the second embodiment, the DAS observation signal and the silent observation signal are input to the deep learning model 13A by the recognition unit 12A.
[0053] Fig. 8 is a diagram showing an example of an intermediate signal with a time-frequency structure obtained in an internal intermediate layer of deep learning model 13A when a DAS observation signal is input to deep learning model 13A. Fig. 9 is a diagram showing an example of an intermediate signal with a time-frequency structure obtained in an internal intermediate layer of deep learning model 13A when a silent observation signal is input to deep learning model 13A.
[0054] Here, when a DAS observation signal is input, the signal in a specific frequency band of the intermediate signal contains a large amount of noise components. For example, in the case of the intermediate signal shown in Figure 8, the signal in the high frequency band contains a large amount of noise components.
[0055] Therefore, the first processing unit 14A replaces the signal in a specific frequency band that contains a large amount of noise components among the intermediate signals when a DAS observation signal is input with the signal in a specific frequency band among the intermediate signals when a silent observation signal is input.
[0056] FIG. 10 is a diagram illustrating an example of SNR improvement processing for improving the SNR of the intermediate signal shown in FIG. 8 in the first processing unit 14A. As described above, the intermediate signal shown in Fig. 8 is assumed to contain many noise components in the high frequency band signal. Therefore, the high frequency band signal among the intermediate signals shown in Fig. 8 becomes the signal in the specific frequency band to be replaced.
[0057] Therefore, the first processing unit 14A replaces the signal in a specific frequency band (here, the high frequency band) that contains a large amount of noise components among the intermediate signals shown in Figure 8 with the signal in a specific frequency band among the intermediate signals shown in Figure 9.
[0058] 8 to 10, the specific frequency band containing a large amount of noise components is described as a high frequency band, but this is not limiting. The specific frequency band containing a large amount of noise components changes dynamically depending on the location where an event is recognized and the conditions of that location (for example, weather, season, time of day, etc.), and may be a low frequency band or an intermediate frequency band between a high frequency band and a low frequency band. Therefore, it is preferable that the specific frequency band be dynamically set in the first processing unit 14A depending on the location where an event is recognized and the conditions of that location. The specific frequency band may be set in the first processing unit 14A manually by the user, for example.
[0059] Next, a schematic operation example of the event recognition device 10A will be described. FIG. 11 is a flow diagram illustrating an example of a schematic operation flow of the event recognition device 10A.
[0060] 11, the acquisition unit 11A acquires, from the DAS device, a DAS observation signal indicating sound generated at a point along the optical fiber and detected by optical fiber sensing. Furthermore, the acquisition unit 11A acquires, from the DAS device, a silent observation signal which is a DAS observation signal at an arbitrary point in a silent state (step S21).
[0061] Next, the recognition unit 12A inputs the DAS observation signal and the silent observation signal acquired by the acquisition unit 11A to the deep learning model 13A as input signals (step S22). Next, the first processing unit 14A uses the intermediate signal obtained in the intermediate layer when the silent observation signal is input to the deep learning model 13A to perform SNR improvement processing on the intermediate signal obtained in the intermediate layer when the DAS observation signal is input to the deep learning model 13A (step S23).
[0062] Specifically, the first processing unit 14A replaces a signal in a specific frequency band that contains a large amount of noise components among the intermediate signals when a DAS observation signal is input with a signal in a specific frequency band among the intermediate signals when a silent observation signal is input. Thereafter, the process of step S24, which is the same as step S14 in FIG. 5, is carried out.
[0063] As described above, according to the second embodiment, the acquisition unit 11A acquires a DAS observation signal indicating sound generated at a point along the optical fiber and detected by optical fiber sensing, and also acquires a silent observation signal at an arbitrary point in a silent state. The recognition unit 12A inputs the DAS observation signal and the silent observation signal as input signals to the deep learning model 13A, which has been trained using the acoustic signal acquired by the acoustic sensor. The first processing unit 14A uses an intermediate signal obtained in the intermediate layer when the silent observation signal is input to the deep learning model 13A, and performs SNR improvement processing on the intermediate signal obtained in the intermediate layer when the DAS observation signal is input to the deep learning model 13A. Other operations of the second embodiment are the same as those of the first embodiment described above.
[0064] As described above, the second embodiment differs from the first embodiment in the operation relating to the SNR improvement process for the intermediate signal, but other operations are the same as those of the first embodiment. Therefore, the second embodiment can achieve the same effects as the first embodiment described above.
[0065] <Third Embodiment> FIG. 12 is a block diagram showing a schematic configuration example of the event recognition device 10B. As shown in FIG. 12, the event recognition device 10B includes an acquisition unit 11B, a second processing unit 15, and a recognition unit 12B.
[0066] The acquisition unit 11B is similar to the acquisition unit 11. The second processing unit 15 performs SNR improvement processing on the DAS observation signal acquired by the acquisition unit 11B to improve the SNR of the DAS observation signal. For example, the SNR improvement processing may be noise suppression processing that suppresses noise in the intermediate signal. Furthermore, for example, the noise suppression processing may be well-known noise suppression processing such as Wiener filtering, spectral subtraction, frequency filtering, cepstrum analysis, etc.
[0067] The recognition unit 12B holds a deep learning model 13B. The deep learning model 13B is similar to the deep learning model 13. The recognition unit 12B also includes a first processing unit 14B. The first processing unit 14B is similar to the first processing unit 14.
[0068] The recognition unit 12B inputs the DAS observation signal, the SNR of which has been improved by the second processing unit 15, to the deep learning model 13B as an input signal. Other operations of the recognition unit 12B are the same as those of the recognition unit 12.
[0069] Next, a schematic operation example of the event recognition device 10B will be described. FIG. 13 is a flow diagram illustrating an example of a schematic operation flow of the event recognition device 10B.
[0070] As shown in FIG. 13, first, the process of step S31, which is the same as step S11 in FIG. 5, is performed. Next, the second processing unit 15 performs an SNR improvement process on the DAS observation signal acquired by the acquisition unit 11B to improve the SNR of the DAS observation signal (step S32).
[0071] Next, the recognition unit 12B inputs the DAS observation signal, the SNR of which has been improved by the second processing unit 15, to the deep learning model 13B as an input signal (step S33). Thereafter, the processes of steps S34 and S35, which are similar to steps S13 and S14 in FIG. 5, are carried out.
[0072] As described above, according to the third embodiment, second processing unit 15 performs SNR improvement processing on the DAS observation signal acquired by acquisition unit 11B to improve the SNR of the DAS observation signal. Recognition unit 12B inputs the DAS observation signal, the SNR of which has been improved by second processing unit 15, to deep learning model 13B as an input signal. Other operations of the third embodiment are the same as those of the first embodiment described above.
[0073] As described above, according to the third embodiment, compared to the first embodiment, an operation is added in which an SNR improvement process is performed on the DAS observation signal, and the DAS observation signal with the improved SNR is input to the deep learning model 13 B. This makes it possible to further close the SNR gap between the DAS observation signal and the acoustic signal, and therefore, when the DAS observation signal is input to the deep learning model 13 B, it is possible to further improve the event recognition performance. Other effects of the third embodiment are the same as those of the first embodiment described above.
[0074] <Fourth Embodiment> The fourth embodiment corresponds to an embodiment that is a superordinate concept of the first to third embodiments described above. FIG. 14 is a block diagram showing a schematic configuration example of the event recognition device 10C. As shown in FIG. 14, an event recognition device 10C includes an acquisition unit 11C and a recognition unit 12C.
[0075] The acquisition unit 11C acquires observation signals indicative of sounds generated at points along the optical fiber and detected by optical fiber sensing. The recognition unit 12C holds a deep learning model 13C. Deep learning model 13C is a model trained using acoustic signals acquired by an acoustic sensor, and is a model that takes an input signal indicating the sound generated at the above-mentioned location as input and outputs a recognition result of an event occurring at the above-mentioned location.
[0076] The recognition unit 12C inputs the observation signal acquired by the acquisition unit 11C to the deep learning model 13C as an input signal, obtains an event recognition result as an output of the deep learning model 13C, and outputs the recognition result.
[0077] The recognition unit 12C also includes a first processing unit 14C. The first processing unit 14C performs a first process to improve the SNR of an intermediate signal having a time-frequency structure obtained in an intermediate layer inside the deep learning model 13C when an observed signal is input to the deep learning model 13C.
[0078] As described above, according to the fourth embodiment, the deep learning model 13C trained using the acoustic signal acquired by the acoustic sensor is used, which eliminates the need for model training using the observation signal acquired by optical fiber sensing.
[0079] Furthermore, according to the fourth embodiment, a first process is performed on an intermediate signal having a time-frequency structure obtained in an internal intermediate layer of the deep learning model 13C to improve the SNR of the intermediate signal. This makes it possible to fill the SNR gap between the observed signal and the acoustic signal. Therefore, even when an observed signal is input to the deep learning model 13C, it is possible to improve the event recognition performance.
[0080] In addition, the deep learning model 13C may be a model trained using an acoustic signal acquired by a microphone as an acoustic sensor. Furthermore, as the first processing, the first processing unit 14C may perform processing on the intermediate signal to suppress noise in the intermediate signal.
[0081] The acquisition unit 11C may further acquire a silent observation signal, which is an observation signal at an arbitrary point in a silent state. The recognition unit 12C may further input the silent observation signal as an input signal to the deep learning model 13C. The first processing unit 14C may perform, as the first process, a process of replacing a signal in a specific frequency band among intermediate signals obtained when the observation signal is input to the deep learning model 13C with a signal in a specific frequency band among intermediate signals obtained when the silent observation signal is input to the deep learning model 13C.
[0082] The event recognition device 10C may further include a second processing unit that performs a second process on the observed signal to improve the SNR of the observed signal. The recognition unit 12C may input the observed signal, the SNR of which has been improved by the second process, as an input signal to the deep learning model 13C. The second processing unit may perform a process on the observed signal to suppress noise in the observed signal as the second process.
[0083] Furthermore, the deep learning model 13C may be a model that outputs the probability that at least one event occurs at the above-mentioned location as an event recognition result. The acquisition unit 11C may also acquire the observation signal from a DAS device.
[0084] <Hardware configuration of the event recognition device> FIG. 15 is a block diagram showing an example of a schematic hardware configuration of a computer 90 that realizes the event recognition devices 10, 10A, 10B, and 10C.
[0085] 15, a computer 90 includes a processor 91, a memory 92, a storage 93, an input / output interface (input / output I / F) 94, and a communication interface (communication I / F) 95. The processor 91, the memory 92, the storage 93, the input / output interface 94, and the communication interface 95 are connected by a data transmission path for transmitting and receiving data to and from each other.
[0086] The processor 91 is, for example, an arithmetic processing device such as a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit). The memory 92 is, for example, a memory such as a RAM (Random Access Memory) or a ROM (Read Only Memory). The storage 93 is, for example, a storage device such as an HDD (Hard Disk Drive), an SSD (Solid State Drive), or a memory card. The storage 93 may also be a memory such as a RAM or a ROM.
[0087] A program is stored in the storage 93. When the program is loaded into the computer, it includes a set of instructions (or software code) that causes the computer 90 to perform one or more functions of the event recognition devices 10, 10A, 10B, and 10C described above. The components of the event recognition devices 10, 10A, 10B, and 10C described above may be realized by the processor 91 reading and executing the program stored in the storage 93. Furthermore, the storage function of the event recognition devices 10, 10A, 10B, and 10C described above may be realized by the memory 92 or the storage 93.
[0088] The above-described programs may also be stored on non-transitory computer-readable media or tangible storage media. By way of example and not limitation, computer-readable media or tangible storage media include RAM, ROM, flash memory, SSD or other memory technology, CD (Compact Disc)-ROM, DVD (Digital Versatile Disc), Blu-ray® disc or other optical disk storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices. The programs may also be transmitted on a transitory computer-readable medium or communication medium. By way of example and not limitation, transitory computer-readable media or communication media include electrical, optical, acoustic, or other forms of propagated signals.
[0089] The input / output interface 94 is connected to a display device 941, an input device 942, a sound output device 943, etc. The display device 941 is a device that displays a screen corresponding to drawing data processed by the processor 91, such as an LCD (Liquid Crystal Display), a CRT (Cathode Ray Tube) display, or a monitor. The input device 942 is a device that accepts operational inputs from an operator, such as a keyboard, a mouse, or a touch sensor. The display device 941 and the input device 942 may be integrated and realized as a touch panel. The sound output device 943 is a device that outputs sound corresponding to audio data processed by the processor 91, such as a speaker.
[0090] The communication interface 95 transmits and receives data to and from an external device. For example, the communication interface 95 communicates with the external device via a wired communication path or a wireless communication path.
[0091] Although the present disclosure has been described above with reference to the embodiments, the present disclosure is not limited to the above-described embodiments. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present disclosure within the scope of the present disclosure. Furthermore, each embodiment can be combined with other embodiments as appropriate.
[0092] Furthermore, each drawing is merely an example for describing one or more embodiments. Each drawing may relate not only to one particular embodiment, but also to one or more other embodiments. As will be understood by those skilled in the art, various features or steps described with reference to any one drawing can be combined with features or steps shown in one or more other drawings to create, for example, an embodiment not explicitly shown or described. Not all features or steps shown in any one drawing are necessarily required to describe an exemplary embodiment, and some features or steps may be omitted. The order of steps described in any drawing may be changed as appropriate.
[0093] Furthermore, some or all of the above-described embodiments can be described as, but are not limited to, the following supplementary notes. (Appendix 1) an acquisition unit that acquires an observation signal indicative of sound generated at a point along the optical fiber and detected by the optical fiber sensing; a deep learning model that is trained using an acoustic signal acquired by an acoustic sensor, and that receives an input signal indicating a sound occurring at the location and outputs a recognition result of an event occurring at the location; a recognition unit that inputs the observation signal to the deep learning model as the input signal, obtains the recognition result as an output of the deep learning model, and outputs the recognition result; a first processing unit that performs a first processing on an intermediate signal of a time-frequency structure obtained in an intermediate layer inside the deep learning model when the observation signal is input to the deep learning model, to improve the signal-to-noise ratio of the intermediate signal; Event recognizer. (Appendix 2) The deep learning model is a model trained using an acoustic signal acquired by a microphone as the acoustic sensor. 2. The event recognizer of claim 1. (Appendix 3) the first processing unit performs, as the first processing, a process of suppressing noise in the intermediate signal. 2. The event recognizer of claim 1. (Appendix 4) the acquisition unit further acquires a silent observation signal that is an observation signal at an arbitrary point in a silent state; The recognition unit further inputs the silent observation signal to the deep learning model as the input signal; the first processing unit performs, as the first processing, a process of replacing a signal in a specific frequency band among the intermediate signals obtained when the observation signal is input to the deep learning model with the signal in the specific frequency band among the intermediate signals obtained when the silent observation signal is input to the deep learning model; 2. The event recognizer of claim 1. (Appendix 5) a second processing unit that performs a second processing on the observation signal to improve a signal-to-noise ratio of the observation signal; The recognition unit inputs the observation signal, the signal-to-noise ratio of which has been improved by the second processing, to the deep learning model as the input signal. 2. The event recognizer of claim 1. (Appendix 6) the second processing unit performs, as the second processing, a process of suppressing noise in the observed signal. 6. The event recognizer of claim 5. (Appendix 7) The deep learning model is a model that outputs, as the recognition result, a probability that at least one event occurs at the location. 2. The event recognizer of claim 1. (Appendix 8) The acquisition unit acquires the observation signal from a DAS (Distributed Acoustic Sensing) device. 2. The event recognizer of claim 1. (Appendix 9) An event recognition method executed by an event recognition device, comprising: acquiring an observation signal indicative of acoustics generated at a point along the optical fiber and detected by the optical fiber sensing; inputting the observed signal as an input signal to a deep learning model that is trained using an acoustic signal acquired by an acoustic sensor and that receives an input signal indicating an acoustic occurring at the location and outputs a recognition result of an event occurring at the location; performing a first process for improving a signal-to-noise ratio of an intermediate signal having a time-frequency structure obtained in an intermediate layer inside the deep learning model when the observation signal is input to the deep learning model; obtaining the recognition result as an output of the deep learning model when the observed signal is input, and outputting the recognition result; An event recognition method comprising: (Appendix 10) On the computer, acquiring observed signals indicative of acoustics generated at points along the optical fiber and detected by the optical fiber sensing; a step of inputting the observed signal as an input signal to a deep learning model that is a model trained using an acoustic signal acquired by an acoustic sensor, the deep learning model receiving an input signal indicating an acoustic occurring at the location and outputting a recognition result of an event occurring at the location; a step of performing a first process on an intermediate signal having a time-frequency structure obtained in an internal intermediate layer of the deep learning model when the observation signal is input to the deep learning model, the first process improving the signal-to-noise ratio of the intermediate signal; obtaining the recognition result as an output of the deep learning model when the observed signal is input, and outputting the recognition result; A program that executes the following.
[0094] Note that some or all of the elements (e.g., configurations and functions) described in Supplementary Notes 2 to 8 that are dependent on Supplementary Note 1 may also be dependent on Supplementary Notes 9 and 10 in the same dependency relationship as Supplementary Notes 2 to 8. Some or all of the elements described in any Supplementary Note may be applied to various hardware, software, recording means for recording software, systems, and methods. [Explanation of symbols]
[0095] 10, 10A, 10B, 10C Event recognition device 11,11A,11B,11C Acquisition part 12,12A,12B,12C recognition part 13, 13A, 13B, 13C, 80 Deep Learning Models 14, 14A, 14B, 14C First processing section 15 Second Processing Section 90 Computer 91 processors 92 memory 93 Storage 94 Input / Output Interface 941 Display device 942 Input Device 943 Sound Output Device 95 Communication Interface
Claims
1. an acquisition unit that acquires an observation signal indicative of sound generated at a point along the optical fiber and detected by the optical fiber sensing; a deep learning model that is trained using an acoustic signal acquired by an acoustic sensor, and that receives an input signal indicating a sound occurring at the location and outputs a recognition result of an event occurring at the location; a recognition unit that inputs the observation signal to the deep learning model as the input signal, obtains the recognition result as an output of the deep learning model, and outputs the recognition result; a first processing unit that performs a first process on an intermediate signal having a time-frequency structure obtained in an intermediate layer inside the deep learning model when the observation signal is input to the deep learning model, to improve the signal-to-noise ratio of the intermediate signal; Event recognizer.
2. The deep learning model is a model trained using an acoustic signal acquired by a microphone as the acoustic sensor. The event recognition device according to claim 1 .
3. the first processing unit performs, as the first processing, a process of suppressing noise in the intermediate signal. The event recognition device according to claim 1 .
4. the acquisition unit further acquires a silent observation signal that is an observation signal at an arbitrary point in a silent state; The recognition unit further inputs the silent observation signal to the deep learning model as the input signal; the first processing unit performs, as the first processing, a process of replacing a signal in a specific frequency band among the intermediate signals obtained when the observation signal is input to the deep learning model with the signal in the specific frequency band among the intermediate signals obtained when the silent observation signal is input to the deep learning model; The event recognition device according to claim 1 .
5. a second processing unit that performs a second processing on the observation signal to improve a signal-to-noise ratio of the observation signal; The recognition unit inputs the observation signal, the signal-to-noise ratio of which has been improved by the second processing, to the deep learning model as the input signal. The event recognition device according to claim 1 .
6. the second processing unit performs, as the second processing, a process of suppressing noise in the observation signal. The event recognition device according to claim 5 .
7. The deep learning model is a model that outputs, as the recognition result, a probability that at least one event occurs at the location. The event recognition device according to claim 1 .
8. The acquisition unit acquires the observation signal from a Distributed Acoustic Sensing (DAS) device. The event recognition device according to claim 1 .
9. An event recognition method executed by an event recognition device, comprising: acquiring an observation signal indicative of acoustics generated at a point along the optical fiber and detected by the optical fiber sensing; inputting the observed signal as an input signal to a deep learning model that is trained using an acoustic signal acquired by an acoustic sensor and that receives an input signal indicating an acoustic occurring at the location and outputs a recognition result of an event occurring at the location; performing a first process for improving a signal-to-noise ratio of an intermediate signal having a time-frequency structure obtained in an internal intermediate layer of the deep learning model when the observation signal is input to the deep learning model; obtaining the recognition result as an output of the deep learning model when the observed signal is input, and outputting the recognition result; An event recognition method comprising:
10. On the computer, acquiring observed signals indicative of acoustics generated at points along the optical fiber and detected by the optical fiber sensing; a step of inputting the observed signal as an input signal to a deep learning model that is a model trained using an acoustic signal acquired by an acoustic sensor, the deep learning model receiving an input signal indicating an acoustic occurring at the location and outputting a recognition result of an event occurring at the location; a step of performing a first process on an intermediate signal having a time-frequency structure obtained in an internal intermediate layer of the deep learning model when the observation signal is input to the deep learning model, the first process improving the signal-to-noise ratio of the intermediate signal; obtaining the recognition result as an output of the deep learning model when the observed signal is input, and outputting the recognition result; A program that executes the following.
Citation Information
Patent Citations
Detection and localization of city-scale acoustic impulses
JP2023538196A
Cited By
Event recognition apparatus, event recognition method, and non-transitory computer-readable medium
US12730000B2