Signal processing device, signal processing method, and program

The signal processing device addresses the trade-off in conventional analysis by using an AI model inspired by the human auditory nervous system to analyze time-series signals, enhancing both time and frequency resolution.

JP7736879B1Active Publication Date: 2025-09-09SOFTBANK CORPORATION
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024128643
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-08-05
Publication Date
2025-09-09
Estimated Expiration
2044-08-05

AI Technical Summary

Technical Problem

Conventional time-series signal analysis techniques face a trade-off between time and frequency analysis resolution, and the selection of window width and function depends on specialized knowledge, limiting their effectiveness.

Method used

A signal processing device that models a neuronal network to extract different sound features using an AI model, incorporating multiple buffering times to analyze both time and frequency domains, mimicking the human auditory nervous system.

Benefits of technology

Enables simultaneous and accurate analysis of both time and frequency domains of time-series signals, improving analysis resolution and reducing reliance on specialized knowledge.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007736879000001_ABST
    Figure 0007736879000001_ABST
Patent Text Reader

Abstract

Conventionally, in the analysis of time-series signals, there has been a problem in that increasing the resolution of frequency analysis reduces the resolution of time analysis, and vice versa. [Solution] A signal processing device 1 comprising a reception unit 12 that receives a time series signal, an information acquisition unit 131 that acquires information on the time length according to the buffering time from the time series signal for each of two or more buffering times, a frequency conversion unit 132 that performs frequency conversion on the information acquired by the information acquisition unit 131 for each of two or more buffering times to acquire an image, a signal transmission unit 133 that passes the image acquired by the frequency conversion unit 132 to a neural network for each of two or more buffering times and acquires output information based on the signal output from the neural network, and an information output unit 135 that outputs the output information, allows both time domain analysis and frequency domain analysis to be performed appropriately.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a signal processing device or the like that processes a time-series signal, acquires information, and outputs the information. [Background technology]

[0002] In the field of signal processing, conventional techniques such as analyzing time series data using an AI model to perform frequency analysis, speech recognition, and sound direction estimation are widely known. It is also known that the analytical accuracy of an AI model is higher when the time series data is converted into frequency data rather than inputted directly. Therefore, it is common to apply a "window function" with a certain buffering time width to the time series data to extract it, and then input the Fourier-transformed frequency data to the AI ​​model. This Fourier transform is called the short-time Fourier transform (STFT) (see Non-Patent Document 1). The AI ​​model here is typically a neural network.

[0003] Furthermore, in order to improve the efficiency of speech learning, there is a speech recognition device characterized by comprising a creation means for creating multiple signal sequences by performing signal analysis on an input speech signal from different times, and a learning means for performing learning for each of the multiple signal sequences created by the creation means (see Patent Document 1). [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2009-25480 [Non-patent literature]

[0005] [Non-Patent Document 1] "Wikipedia:Short-time Fourier transform", [online], [Retrieved July 24, 2024], Internet [URL: https: / / ja.wikipedia.org / wiki / Short_time_Fourier_transform] Summary of the Invention [Problem to be solved by the invention]

[0006] However, in conventional time-series signal analysis techniques, there is a trade-off between time analysis and frequency analysis due to Heisenberg uncertainty. In other words, in conventional techniques, increasing the resolution of frequency analysis reduces the resolution of time analysis, and vice versa. Furthermore, in conventional time-series signal analysis techniques, the selection of the window width and window function, which have a significant impact on the analysis results, depends on the signal data being analyzed, requiring the analyst's specialized knowledge and experience.

[0007] In view of the above problems, the inventor was inspired by the biological fact that three cells in the human auditory nervous system (tufted cells, stellate cells, and octopus cells) extract sound features on different time scales, and wondered whether this could be reproduced using an AI model. As a result, the present invention proposes a signal processing device that can solve the above problems by modeling a neuronal network that can extract different sound features, and that has an AI model for analyzing general-purpose time-series signals that can be used for a variety of tasks. [Means for solving the problem]

[0008] The signal processing device of the first invention is a signal processing device comprising: a receiving unit that receives a time series signal; an information acquiring unit that acquires information of a time length according to the buffering time from the time series signal for each of two or more buffering times; a frequency conversion unit that performs frequency conversion on the information acquired by the information acquiring unit for each of two or more buffering times to acquire an image; a signal transmission unit that passes the image acquired by the frequency conversion unit to a neural network for each of two or more buffering times and acquires output information based on the signal output from the neural network; and an information output unit that outputs the output information.

[0009] With this configuration, it is possible to appropriately analyze both the time domain and the frequency domain of a time series signal.

[0010] Furthermore, in contrast to the first invention, the signal processing device of the second invention further includes a learning unit that, for each of two or more buffering times, passes the image acquired by the frequency conversion unit to a sub-neural network corresponding to the buffering time, updates each of the two or more sub-neural networks, and merges the two or more updated sub-neural networks to form a single neural network, and the signal transmission unit passes the image acquired by the frequency conversion unit to one neural network for each of two or more buffering times, and acquires output information output from the one neural network.

[0011] With this configuration, it is possible to appropriately analyze both the time domain and the frequency domain of a time series signal.

[0012] Furthermore, in the signal processing device of the third invention, compared to the first invention, the signal transmission unit passes the image acquired by the frequency conversion unit to a neural network for each of two or more buffering times, and acquires output information output from each of the two or more neural networks, and the information output unit outputs the two or more pieces of output information acquired by the signal transmission unit, or one piece of output information based on the two or more pieces of output information acquired by the signal transmission unit.

[0013] With this configuration, it is possible to appropriately analyze both the time domain and the frequency domain of a time series signal.

[0014] Furthermore, the signal processing device of the fourth invention differs from the signal processing device of the third invention in that the signal transmission unit passes the image acquired by the frequency conversion unit to a sub-neural network for each of two or more buffering times, acquires output information output from each of the two or more sub-neural networks, passes the two or more pieces of output information to one neural network, acquires the output information output from the neural network, and the information output unit outputs the output information acquired by the signal transmission unit.

[0015] With this configuration, it is possible to appropriately analyze both the time domain and the frequency domain of a time series signal.

[0016] Furthermore, the signal processing device of the fifth invention is a signal processing device according to any one of the first to fourth inventions, in which the reception unit receives two or more time series signals, the information acquisition unit acquires information of a time length according to the buffering time from the time series signals for each of the two or more time series signals and for each of the two or more buffering times, and the frequency conversion unit performs frequency conversion on the information acquired by the information acquisition unit for each of the two or more buffering times, thereby acquiring an image.

[0017] With this configuration, it is possible to appropriately analyze both the time domain and the frequency domain of a time series signal.

[0018] Furthermore, the signal processing device of the sixth invention is a signal processing device in which, compared to any one of the first to fifth inventions, the two or more buffering times are three or more buffering times, and three of the three or more buffering times correspond to long-term, medium-term, and short-term time scales, respectively, like the tufted cells, stellate cells, and octopus cells contained in the human auditory nerve.

[0019] With this configuration, by utilizing a model based on the auditory nervous system of a living organism, it is possible to more appropriately analyze both the time domain and the frequency domain of a time series signal.

[0020] Furthermore, a signal processing device of the seventh invention is a signal processing device according to any one of the first to sixth inventions, wherein the time series signal is an audio signal, and the output information includes any one of a frequency analysis result for the audio signal, a speech recognition result for the audio signal, a sound source separation result for the audio signal, and a sound source direction estimation result for the audio signal.

[0021] With this configuration, it is possible to output any one of an appropriate frequency analysis result, an appropriate speech recognition result, an appropriate sound source separation result, and an appropriate sound source direction estimation result for a time-series signal of speech. [Effects of the Invention]

[0022] The signal processing device according to the present invention can appropriately perform both time domain analysis and frequency domain analysis of a time series signal. [Brief explanation of the drawings]

[0023] [Figure 1] Block diagram of a signal processing device 1 according to a first embodiment [Figure 2] A flowchart illustrating an example of the operation of the signal processing device 1. [Figure 3] A flowchart illustrating an example of the image acquisition process [Figure 4] A flowchart illustrating an example of the learning process [Figure 5] 10 is a flowchart illustrating a first example of the output information acquisition process. [Figure 6] 10 is a flowchart illustrating a second example of the output information acquisition process. [Figure 7] 10 is a flowchart illustrating a third example of the output information acquisition process. [Figure 8] FIG. 1 is a schematic diagram illustrating a specific example of the operation of the signal processing device 1. [Figure 9] FIG. 1 is a schematic diagram illustrating a specific example of the operation of the signal processing device 1. [Figure 10] Block diagram of the computer system DETAILED DESCRIPTION OF THE INVENTION

[0024] Hereinafter, embodiments of a signal processing device and the like will be described with reference to the drawings. Note that components with the same reference numerals in the embodiments perform similar operations, and therefore repeated description may be omitted.

[0025] (Embodiment 1) In this embodiment, a signal processing device is described that receives a time series signal, acquires information from the time series signal according to each of two or more buffering times, frequency-converts each of the acquired information, acquires an image, passes each of the two or more images to a neural network, acquires output information based on the output from the neural network, and outputs the output information.

[0026] In this embodiment, we will describe a signal processing device with a learning function that passes an image corresponding to each of two or more buffering times to a sub-neural network corresponding to the buffering time, and combines the two or more sub-neural networks to form a single neural network.

[0027] In this embodiment, a signal processing device is described that passes an image to a sub-neural network corresponding to each of two or more buffering times, obtains an output from each of the two or more sub-neural networks, and outputs output information based on the output from each of the two or more sub-neural networks.

[0028] In this embodiment, a signal processing device that receives two or more time-series signals, obtains output information based on the two or more time-series signals, and outputs the output information will be described.

[0029] In this embodiment, a signal processing device having a configuration that mimics the human auditory nerve will be described.

[0030] In this specification, information X being associated with information Y means that information Y can be obtained from information X, or information X can be obtained from information Y, and the method of association is not important. Information X and information Y may be linked, may exist in the same buffer, information X may be included in information Y, or information Y may be included in information X, etc.

[0031] Furthermore, in this specification, selecting or determining information Z means obtaining information Z, obtaining a pointer to information Z, obtaining the ID of information Z, setting a flag on information Z, etc., and it is sufficient if information Z can be accessed.

[0032] FIG. 1 is a block diagram of a signal processing device 1 according to this embodiment. The signal processing device 1 is a device that receives a time-series signal and outputs output information. The signal processing device 1 may be a terminal, or a server that receives a time-series signal from a terminal device (not shown) and transmits output information to the terminal device or another device. When the signal processing device 1 is a terminal, the signal processing device 1 may be, for example, a personal computer, a smartphone, or a tablet terminal, but the type does not matter. When the signal processing device 1 is a server, the signal processing device 1 may be, for example, a cloud server or an ASP server, but the type does not matter.

[0033] The signal processing device 1 includes a storage unit 11, a reception unit 12, a processing unit 13, and an output unit 14. The storage unit 11 includes an SNN storage unit 111, an NN storage unit 112, and two or more buffers 113. The processing unit 13 includes an information acquisition unit 131, a frequency conversion unit 132, a signal transmission unit 133, and a learning unit 134. The output unit 14 includes an information output unit 135.

[0034] The storage unit 11 constituting the signal processing device 1 is a storage area. The storage unit 11 may be composed of two or more types of recording media (for example, a non-volatile recording medium and a volatile recording medium). Various types of information are stored in the storage unit 11. The various types of information are, for example, two or more sub-neural networks (described later) and a neural network (described later).

[0035] The SNN storage unit 111 stores two or more sub-neural networks (hereinafter referred to as "SNNs" as appropriate). Each of the two or more sub-neural networks corresponds to a different buffering time. The SNN also has a neural network structure. A neural network may also be called a neural network. For example, two or more SNNs are merged to form the neural network in the NN storage unit 112.

[0036] The buffering time is the time length of information extracted from a time-series signal, and is, for example, the time width of a window function.

[0037] The NN storage unit 112 stores a neural network (hereinafter referred to as "NN" where appropriate).

[0038] The activation functions and structures of the nodes in NNs and SNNs are not important. The edges in NNs and SNNs correspond to weights, for example. NNs and SNNs typically have an input layer, an intermediate layer (hidden layer), and an output layer. The type of neural network in NNs and SNNs is not important, and may be, for example, a feedforward neural network (FFNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a spiking neural network (SNN), etc. NNs and SNNs are typically neural networks that can be used in so-called deep learning. The neural networks described above may also be included as neural networks that can be used in deep learning.

[0039] NNs and SNNs are networks that typically accept image feature vectors at their input layer and output output information (described later) or information that serves as the basis for the output information from their output layer.

[0040] Furthermore, NNs and SNNs are AI models that accept, for example, image feature vectors (described later) based on audio signals, which are time-series signals, and output output information or a signal that is the basis for the output information, such as frequency analysis results for the audio signals, speech recognition results for the audio signals, sound source separation results for the audio signals, or sound source direction estimation results for the audio signals.

[0041] Furthermore, NNs and SNNs are AI models that accept, for example, image feature vectors (described later) based on a video signal, which is a time-series signal, and output either a video summary for the video signal or the recognition results of objects in the video, or a signal that serves as the basis for the output information.

[0042] The storage unit 11 has buffer 113(1), buffer 113(2), ... buffer 113(N) [N is a natural number greater than or equal to 2]. Each of the two or more buffers 113 corresponds to a buffering time. The buffer 113 may be a volatile recording medium such as a memory or a cache, or may be a non-volatile recording medium such as a hard disk or an SSD. Each of the two or more buffers 113 in the storage unit 11 corresponds to a different buffering time. Note that two or more buffers 113 may correspond to one buffering time.

[0043] The reception unit 12 receives various instructions and information. The various instructions are, for example, a learning instruction and an output instruction. A learning instruction is to perform learning and update the SNN or NN. The learning instruction may include a time series signal. An output instruction is an instruction to obtain output information based on the time series signal and output the output information. The output instruction may include a time series signal.

[0044] The receiving unit 12 receives a time-series signal. The receiving unit 12 may receive two or more time-series signals.

[0045] A time-series signal is data that changes over time. Examples of time-series signals include audio data, video, economic data such as stock prices and exchange rates, meteorological data such as temperature and precipitation, physiological data such as heart rate and blood pressure, and industrial data such as sensor data and machine vibration data. Note that the time-series signal referred to here is not limited. Note that audio data may also be called audio signals. Video may also be called image signals, video signals, or video signals.

[0046] The receiving unit 12 receives, for example, the following two or more time-series signals: The receiving unit 12 receives, for example, the following two or more time-series signals in the same time period: The two or more time-series signals are audio data from a plurality of microphones, video from a plurality of cameras, trend data of stock prices of a plurality of companies, trend data of exchange rates between one currency and two or more other currencies, and trend data of multiple physiological data of one person (for example, body temperature, heart rate, and blood pressure).

[0047] The two or more time-series signals received by the receiving unit 12 are, for example, audio signals generated from two or more sound sources installed at different positions.

[0048] Here, reception is a concept that includes reception of information input from input devices such as a microphone, keyboard, mouse, or touch panel, reception of information transmitted via a wired or wireless communication line, and reception of information read from recording media such as an optical disk, magnetic disk, or semiconductor memory.

[0049] The processing unit 13 performs various types of processing, such as processing performed by the information acquisition unit 131, the frequency conversion unit 132, the signal transmission unit 133, or the learning unit .

[0050] The information acquisition unit 131 acquires, for each of two or more buffering times, information on the time length according to the buffering time from the time-series signal accepted by the acceptance unit 12. The information acquisition unit 131 typically temporarily stores the acquired information for each of two or more buffering times in the buffer 113 corresponding to the buffering time. For example, the information acquisition unit 131 extracts information from the time-series signal for each of two or more buffering times using a window function of the time width indicated by the buffering time. Note that two or more buffering times may be considered to be equivalent to two or more buffers 113. In this case, two or more buffers 113 may be associated with one buffering time.

[0051] The information acquiring unit 131 may acquire, for each of the two or more time-series signals received by the receiving unit 12 and for each of the two or more buffering times, information on the time length according to the buffering time from the time-series signals.

[0052] The two or more buffering times are preferably, for example, three or more buffering times. For example, it is preferable that three of the three or more buffering times have time lengths corresponding to long-term, medium-term, and short-term time scales, respectively. The time lengths corresponding to the long-term, medium-term, and short-term time scales are, for example, time lengths corresponding to the time scales of tufted cells, astrocytes, and octopus cells contained in the auditory nerve of living organisms, including humans. The time scale of tufted cells is in the millisecond range (usually within a few milliseconds). The time scale of astrocytes is tens of milliseconds to hundreds of milliseconds. The time scale of octopus cells is hundreds of microseconds to several milliseconds.

[0053] The frequency conversion unit 132 performs frequency conversion on each piece of information acquired by the information acquisition unit 131 for each of two or more buffering times to acquire an image. Two or more images may be acquired for each buffering time. The frequency conversion algorithm may be, for example, a fast Fourier transform (FFT), a short-time Fourier transform (STFT), a wavelet transform, a Wigner-Ville distribution, or the like, but is not limited thereto.

[0054] The signal transmission unit 133 transmits the image acquired by the frequency conversion unit 132 for each of two or more buffering times, which corresponds to the buffering time, to the neural network and acquires output information output from the neural network. The output information output from the neural network may be output information based on a signal output from the neural network. Note that the signal may be considered as information.

[0055] The signal transmission unit 133 acquires, for example, at each of two or more buffering times, a feature vector of the image acquired by the frequency conversion unit 132, provides the feature vector to the input layer of the neural network, propagates a signal within the neural network, and acquires output information based on the signal output from the output layer of the neural network. Note that passing an image to a neural network usually means passing the feature vector of the image to the neural network. Furthermore, signal propagation processing within a neural network is a well-known technique, and therefore a detailed description thereof will be omitted.

[0056] Note that an image feature vector is a set of image features. The image features that make up the feature vector may be, for example, color features, shape features, texture features, etc., but are not limited thereto. The technology for acquiring a feature vector from an image is well known.

[0057] The signal transmission unit 133 passes the image acquired by the frequency conversion unit 132 to one neural network for each of two or more buffering times, and acquires output information based on the signal output from the one neural network. Note that the neural network in this case is the neural network in the NN storage unit 112.

[0058] For example, the signal transmission unit 133 passes the image acquired by the frequency conversion unit 132 to a sub-neural network for each of two or more buffering times, and acquires output information based on the signals output from each of the two or more sub-neural networks. Here, the signal transmission unit 133 may acquire one piece of output information based on the output information from each of the two or more sub-neural networks. Note that the sub-neural network is a neural network in the SNN storage unit 111.

[0059] The output information is information to be output. The output information includes, for example, one or more of a frequency analysis result for a speech signal, a speech recognition result for a speech signal, a sound source separation result for a speech signal, and a sound source direction estimation result for a speech signal. However, the output information is not limited.

[0060] For each of the two or more buffering times, the learning unit 134 passes the image acquired by the frequency conversion unit 132, which corresponds to the buffering time, to the sub-neural network corresponding to the buffering time, and updates each of the two or more sub-neural networks. The learning unit 134, for example, stores the two or more updated sub-neural networks in the SNN storage unit 111.

[0061] Furthermore, the learning unit 134 may, for example, merge two or more updated sub-neural networks to form one neural network. The learning unit 134 may, for example, store the formed neural network in the NN storage unit 112.

[0062] The sub-neural network is updated by a neural network learning process. The update of the sub-neural network is, for example, a process of changing the weight of an edge or a process of changing the probability that a node will fire. The more signals pass through an edge, the greater the weight of the edge. If there are many opportunities or periods when no signals pass through an edge, for example, the weight of the edge will be smaller. The more a node fires, the greater the probability that the node will fire. If there are many opportunities or periods when the node does not fire, for example, the probability that the node will fire will be smaller. The processing of the learning unit 134 is a known technique. The processing of the learning unit 134 is, for example, a deep learning process.

[0063] Merging two or more sub-neural networks is usually a process of generating edges to connect adjacent nodes of two or more adjacent sub-neural networks, and combining the two or more sub-neural networks. Note that the method for combining the two or more sub-neural networks is not important.

[0064] The output unit 14 outputs various types of information. The various types of information are output information, which will be described later. Here, output is a concept that includes, for example, displaying on a display, projecting using a projector, printing using a printer, outputting sound, transmitting to an external device, storing in a recording medium, and transferring processing results to another processing device, another program, etc.

[0065] The information output unit 135 outputs one or more pieces of output information acquired by the signal transmission unit 133. The information output unit 135 outputs, for example, two or more pieces of output information acquired by the signal transmission unit 133, or one piece of output information based on the two or more pieces of output information acquired by the signal transmission unit 133.

[0066] The storage unit 11, the SNN storage unit 111, the NN storage unit 112, and the buffer 113 are preferably non-volatile recording media, but may also be realized as volatile recording media.

[0067] There is no restriction on the process by which information is stored in the storage unit 11 etc. For example, information may be stored in the storage unit 11 etc. via a recording medium, information transmitted via a communication line etc. may be stored in the storage unit 11 etc., or information input via an input device may be stored in the storage unit 11 etc.

[0068] The reception unit 12 is preferably realized by a device driver for an input means such as a microphone, touch panel, or keyboard, control software for a menu screen, or a wireless or wired communication means, but can also be realized by a means for receiving broadcasts, etc.

[0069] The processing unit 13, information acquisition unit 131, frequency conversion unit 132, signal transmission unit 133, and learning unit 134 can usually be realized by a processor, memory, etc. The processing procedures of the processing unit 13, etc. are usually realized by software, and the software is recorded on a recording medium such as a ROM. However, they may also be realized by hardware (dedicated circuit). The processor may be a CPU, MPU, GPU, etc., and the type is not important.

[0070] The output unit 14 and the information output unit 135 can be realized by driver software for an output device such as a display or a speaker, or by a combination of driver software for an output device and the output device, etc. The output unit 14 may be realized by wireless or wired communication means.

[0071] Next, an example of the operation of the signal processing device 1 will be described with reference to the flowchart of FIG.

[0072] (Step S201) The reception unit 12 determines whether or not a learning instruction has been received. If a learning instruction has been received, the process proceeds to step S202, and if not, the process proceeds to step S207.

[0073] (Step S202) The reception unit 12 determines whether or not a time-series signal for the learning process has been received. If a time-series signal has been received, the process proceeds to step S203; if not, the process proceeds to step S205. Here, it is assumed that the reception unit 12 sequentially receives time-series signals after receiving a learning instruction, for example.

[0074] (Step S203) The processing unit 13 performs image acquisition processing for each of the two or more buffers 113. An example of the image acquisition processing will be described with reference to the flowchart of FIG.

[0075] (Step S204) The learning unit 134 performs the learning process. Return to step S202. An example of the learning process will be described with reference to the flowchart in FIG.

[0076] (Step S205) The processing unit 13 determines whether or not a timeout has occurred. If a timeout has occurred, the process proceeds to step S206, and if not, the process returns to step S202. Note that, for example, if a predetermined time or more has passed since the time-series signal was received, the processing unit 13 determines that a timeout has occurred.

[0077] (Step S206) The learning unit 134 merges two or more SNNs updated by the learning process to form one NN, and stores the NN in the NN storage unit 112. The process returns to step S201.

[0078] (Step S207) The reception unit 12 determines whether or not an output instruction has been received. If an output instruction has been received, the process proceeds to step S208, and if not, the process returns to step S201.

[0079] (Step S208) The reception unit 12 determines whether or not a time-series signal for the output information acquisition process has been received. If a time-series signal has been received, the process proceeds to step S209; if not, the process proceeds to step S213. Here, it is assumed that the reception unit 12 sequentially receives time-series signals after receiving an output instruction, for example.

[0080] (Step S209) The processing unit 13 performs image acquisition processing for each buffer 113. An example of the image acquisition processing will be described with reference to the flowchart of FIG.

[0081] (Step S210) The processing unit 13 performs an output information acquisition process. An example of the output information acquisition process will be described with reference to the flowcharts of FIGS.

[0082] (Step S211) The information output unit 135 outputs the output information.

[0083] (Step S212) The processing unit 13 determines whether or not to end the processing. If the processing is to be ended, the process returns to step S201, and if not, the process returns to step S208. The processing is ended, for example, when an end instruction is received.

[0084] (Step S213) The processing unit 13 determines whether or not a timeout has occurred. If a timeout has occurred, the process returns to step S201, and if not, the process returns to step S208. Note that, for example, if a predetermined time or more has passed since the time-series signal was received, the processing unit 13 determines that a timeout has occurred.

[0085] In the flowchart of FIG. 2, the process ends when the power is turned off or an interrupt occurs to end the process.

[0086] Next, an example of the image acquisition process in steps S203 and S209 will be described with reference to the flowchart in FIG.

[0087] (Step S301) The information acquisition unit 131 assigns 1 to a counter i.

[0088] (Step S302) The information acquisition unit 131 determines whether the i-th buffer 113 exists. If the i-th buffer 113 exists, the process proceeds to step S303, and if it does not exist, the process returns to the upper level process. Note that the determination of whether the i-th buffer 113 exists may be considered to be the same as the determination of whether the i-th buffering time exists.

[0089] (Step S303) The information acquisition unit 131 acquires the buffering time corresponding to the i-th buffer 113.

[0090] (Step S304) The information acquisition unit 131 assigns 1 to a counter j.

[0091] (Step S305) The information acquisition unit 131 determines whether or not the j-th information having the length indicated by the buffering time acquired in step S303 can be acquired from the accepted time-series signal. If the j-th information can be acquired, the process proceeds to step S306; if not, the process proceeds to step S309.

[0092] (Step S306) The information acquiring unit 131 acquires the j-th information having the length indicated by the buffering time acquired in step S303 from the received time-series signal, and temporarily stores the information in the i-th buffer 113.

[0093] (Step S307) The frequency conversion unit 132 frequency-converts the information acquired in step S306, acquires an image, and temporarily stores the image in the i-th buffer 113 or a buffer not shown.

[0094] (Step S308) The information acquisition unit 131 increments the counter j by 1. The process returns to step S305.

[0095] (Step S309) The information acquisition unit 131 increments the counter i by 1. The process returns to step S302.

[0096] Next, an example of the learning process in step S204 will be described with reference to the flowchart in FIG.

[0097] (Step S401) The learning unit 134 assigns 1 to a counter i.

[0098] (Step S402) The learning unit 134 determines whether or not the ith image exists among the images temporarily stored in the buffer.

[0099] (Step S403) The learning unit 134 acquires the buffering time corresponding to the i-th image.

[0100] (Step S404) The learning unit 134 acquires the SNN corresponding to the i-th image from the SNN storage unit 111.

[0101] (Step S405) The learning unit 134 acquires the feature vector of the i-th image.

[0102] (Step S406) The learning unit 134 passes the feature vector acquired in step S405 to the SNN acquired in step S404.

[0103] (Step S407) The learning unit 134 propagates the signal within the SNN.

[0104] (Step S408) The learning unit 134 updates the weight of each of one or more edges in the SNN in accordance with the propagation of the signal within the SNN.

[0105] (Step S409) The learning unit 134 updates the firing probability of one or more nodes in the SNN in accordance with the propagation of signals within the SNN.

[0106] (Step S410) The learning unit 134 increments the counter i by 1. The process returns to step S402.

[0107] In the flowchart of FIG. 4, it is preferable that the learning unit 134 updates the weight of one or more edges and updates the firing probability of one or more nodes while propagating a signal within the SNN.

[0108] Next, a first example of the output information acquisition process of step S210 will be described using the flowchart of Fig. 5. The first example of the output information acquisition process is a case where there is one NN that passes images corresponding to two or more buffers 113. Note that the NN here is, for example, an NN configured by the operations described using the flowchart of Fig. 4.

[0109] (Step S501) The signal transmission unit 133 assigns 1 to a counter i.

[0110] (Step S502) The signal transmission unit 133 determines whether or not the i-th image exists among the temporarily stored images. If the i-th image exists, the process proceeds to step S503, and if not, the process proceeds to step S510.

[0111] (Step S503) The signal transmission unit 133 acquires the NN from the NN storage unit 112.

[0112] (Step S504) The signal transmission unit 133 acquires the feature vector of the i-th image.

[0113] (Step S505) The signal transmission unit 133 passes the feature vector acquired in step S504 to the NN acquired in step S503.

[0114] (Step S506) The signal transmission unit 133 propagates the signal within the NN acquired in step S503.

[0115] (Step S507) The signal transmission unit 133 acquires output information based on the signal output from the output layer of the NN.

[0116] (Step S508) The signal transmission unit 133 increments the counter i by 1. The process returns to step S502.

[0117] (Step S509) The signal transmission unit 133 composes output information to be output based on the two or more pieces of output information acquired in step S507, and returns to the upper level processing.

[0118] 5, the signal transmission unit 133 acquires output information for each image acquired by the frequency conversion unit 132. Acquisition of output information for each image may be considered as acquisition of output information for each buffering time corresponding to the image.

[0119] Next, a second example of the output information acquisition process of step S210 will be described using the flowchart in Fig. 6. The second example of the output information acquisition process is a case where output information is configured from output information based on signals from two or more SNNs.

[0120] (Step S601) The signal transmission unit 133 assigns 1 to a counter i.

[0121] (Step S602) The signal transmission unit 133 determines whether or not the i-th image exists among the temporarily stored images. If the i-th image exists, the process proceeds to step S603, and if not, the process proceeds to step S610.

[0122] (Step S603) The signal transmission unit 133 acquires the buffering time corresponding to the i-th image.

[0123] (Step S604) The signal transmission unit 133 acquires from the SNN storage unit 111 the SNN corresponding to the buffering time acquired in step S603.

[0124] (Step S605) The signal transmission unit 133 acquires the feature vector of the i-th image.

[0125] (Step S606) The signal transmission unit 133 passes the feature vector acquired in step S605 to the SNN acquired in step S604.

[0126] (Step S607) The signal transmission unit 133 propagates the signal within the SNN acquired in step S604.

[0127] (Step S608) The signal transmission unit 133 acquires output information corresponding to the i-th image based on a signal from the output layer resulting from signal propagation within the SNN.

[0128] (Step S609) The signal transmission unit 133 increments the counter i by 1. The process returns to step S602.

[0129] (Step S610) The signal transmission unit 133 composes output information to be output based on the two or more pieces of output information acquired in step S608. The process returns to the upper level processing. Note that the output information to be output may be any information based on the two or more pieces of output information acquired in step S608. The output information to be output may be, for example, information that includes the two or more pieces of output information as they are, information in which the two or more pieces of output information are merged into one piece of information, or information obtained by providing the two or more pieces of output information to a function and executing the function.

[0130] Next, a third example of the output information acquisition process of step S210 will be described using the flowchart of Fig. 7. In the flowchart of Fig. 7, explanations of the same steps as in the flowchart of Fig. 6 will be omitted. The third example of the output information acquisition process is a case where output information based on two or more SNNs is given to one NN and the output information to be output is acquired.

[0131] (Step S701) The signal transmission unit 133 acquires the NN from the NN storage unit 112.

[0132] (Step S702) The signal transmission unit 133 passes the two or more pieces of output information acquired in step S608 to the NN acquired in step S701. Note that the method for passing the two or more pieces of output information to the NN does not matter. The signal transmission unit 133 may pass the two or more pieces of output information to the NN sequentially, or the signal transmission unit 133 may pass to the NN a single piece of information based on the two or more pieces of output information. Note that passing a single piece of information based on two or more pieces of output information to the NN means, for example, passing to the NN a single piece of information obtained by merging the two or more pieces of output information, passing to the NN a single piece of information obtained by calculating the two or more pieces of output information, passing to the NN a single piece of information obtained by searching a database (not shown) using the two or more pieces of output information as a key, etc.

[0133] (Step S703) The signal transmission unit 133 propagates a signal within the NN acquired in step S701.

[0134] (Step S704) The signal transmission unit 133 acquires output information based on a signal from the output layer as a result of signal propagation within the NN, and returns to the upper level processing.

[0135] In the flowcharts of FIGS. 5 to 7, when the signal transmission unit 133 propagates a signal within an NN or SNN, learning processing such as updating edge weights and node firing probabilities may also be performed.

[0136] An example of a specific outline of the operation of the signal processing device 1 according to this embodiment will be described below.

[0137] (Example 1) A specific example of the operation of the signal processing device 1 will be described with reference to the schematic diagram of the specific operation of the signal processing device 1 in FIG.

[0138] First, the receiving unit 12 of the signal processing device 1 receives a time-series signal (801 in FIG. 8). Here, the storage unit 11 of the signal processing device 1 has four buffers 113 with buffering times "T=10, 5, 1, 0.01" (802).

[0139] Then, the information acquisition unit 131 acquires information from the time-series signal using a window function according to the buffering time for each of the four buffers 113. It is preferable that the processing for each of the four buffers 113 be performed in parallel.

[0140] Next, for each of the four buffers 113, the frequency conversion unit 132 performs frequency conversion on the information sequentially acquired by the information acquisition unit 131, and sequentially acquires images (803). Note that it is preferable that the frequency conversion process for each buffer 113 (each buffering time) is also performed in parallel.

[0141] Next, the signal transmission unit 133 acquires the SNN corresponding to each of the four buffers 113 from the SNN storage unit 111, sequentially provides the feature vector of the image acquired by the frequency conversion unit 132 to the SNN corresponding to each buffer 113, propagates the signal to the SNN, and obtains output information for each SNN (804). Note that it is preferable that the signal propagation process within the SNN in 804 of Fig. 8 is performed in parallel.

[0142] Next, the signal transmission unit 133 acquires one NN from the NN storage unit 112, passes output information that is output from each of two or more SNNs to the one NN, propagates signals within the one NN, and acquires output information based on the signals from the one NN (805). Note that it is preferable that the signal propagation processing within the NN in 805 of Figure 8 is performed in parallel.

[0143] Next, the information output unit 135 outputs the output information based on the output from the signal transmission unit 133 to one NN (806).

[0144] (Example 2) A specific example of the operation of the signal processing device 1 will be described with reference to the schematic diagram of the specific operation of the signal processing device 1 in FIG.

[0145] First, the receiving unit 12 of the signal processing device 1 receives a time-series signal. Here, the storage unit 11 of the signal processing device 1 has three buffers 113 for buffering times (T1 simulating tufted cells, T2 simulating astrocytes, and T3 simulating octopus cells).

[0146] Then, the information acquisition unit 131 acquires information from the time-series signal using a window function according to the buffering time for each of the three buffers 113. It is preferable that the processing for each of the three buffers 113 be performed in parallel.

[0147] Next, for each of the three buffers 113, the frequency conversion unit 132 performs frequency conversion on the information sequentially acquired by the information acquisition unit 131, and sequentially acquires images. Note that it is preferable that the frequency conversion process for each buffer 113 (each buffering time) is also a parallel process. An image of the above process is shown in 901 in FIG. 9.

[0148] Next, the signal transmission unit 133 acquires the SNN (903 in FIG. 9) corresponding to each of the three buffers 113 from the SNN storage unit 111, sequentially provides the feature vectors of the images acquired by the frequency conversion unit 132 to the SNNs corresponding to each buffer 113, propagates the signal to the SNNs, and obtains output information for each SNN (903).

[0149] Thereafter, the information output unit 135 may output the three pieces of output information for each SNN as they are, or may output one piece of output information based on the three pieces of output information. Also, as described in specific example 1, the signal transmission unit 133 may acquire one NN from the NN storage unit 112, pass the output information that is the output from each of the three SNNs to the one NN, propagate signals within the one NN, acquire the output information from the one NN, and the information output unit 135 may output the output information.

[0150] As described above, according to this embodiment, it is possible to appropriately perform both time domain analysis and frequency domain analysis of a time series signal.

[0151] Furthermore, according to this embodiment, by using a model based on the auditory nervous system of a living organism, it is possible to more appropriately analyze both the time domain and the frequency domain of a time series signal.

[0152] Furthermore, according to this embodiment, any one of a frequency analysis result, a speech recognition result, a sound source separation result, and a sound source direction estimation result can be output for a time-series signal of speech.

[0153] The processing in this embodiment may be implemented by software. This software may be distributed by software download or the like. Furthermore, this software may be recorded on a recording medium such as a CD-ROM and distributed. This also applies to other embodiments in this specification. The software implementing the signal processing device 1 in this embodiment is the following program. That is, this program causes a computer to function as: a receiving unit that receives a time-series signal; an information acquiring unit that acquires, from the time-series signal, information of a time length according to each of two or more buffering times; a frequency converting unit that performs frequency conversion on the information acquired by the information acquiring unit and acquires an image, for each of the two or more buffering times; a signal transmitting unit that passes the image acquired by the frequency converting unit to a neural network for each of the two or more buffering times and acquires output information based on the signal output from the neural network; and an information output unit that outputs the output information.

[0154] FIG. 10 is a block diagram of a computer system 300 that executes the programs described in this specification to realize the signal processing devices 1 and the like according to the various embodiments described above.

[0155] In FIG. 10, computer system 300 includes computer 301 including a CD-ROM drive, keyboard 302, mouse 303, and monitor 304.

[0156] 10, computer 301 includes, in addition to CD-ROM drive 3012, MPU 3013, bus 3014 connected to CD-ROM drive 3012 etc., ROM 3015 for storing programs such as a boot-up program, RAM 3016 connected to MPU 3013 for temporarily storing instructions of application programs and providing temporary storage space, and hard disk 3017 for storing application programs, system programs, and data. Although not shown here, computer 301 may further include a network card for providing connection to a LAN.

[0157] A program that causes the computer system 300 to execute the functions of the signal processing device 1 and the like of the above-described embodiment may be stored on a CD-ROM 3101, inserted into the CD-ROM drive 3012, and then transferred to the hard disk 3017. Alternatively, the program may be transmitted to the computer 301 via a network (not shown) and stored on the hard disk 3017. The program is loaded into the RAM 3016 when executed. The program may also be loaded directly from the CD-ROM 3101 or the network.

[0158] The program does not necessarily include an operating system (OS) or a third-party program that causes the computer 301 to execute the functions of the signal processing device 1 of the above-described embodiment. The program only needs to include instructions that call appropriate functions (modules) in a controlled manner to achieve the desired results. How the computer system 300 operates is well known, and a detailed description thereof will be omitted.

[0159] In addition, in the above program, the steps of transmitting information and receiving information do not include processing performed by hardware, such as processing performed by a modem or interface card in the transmission step (processing that can only be performed by hardware).

[0160] The computer that executes the program may be a single computer or a plurality of computers, that is, it may perform centralized processing or distributed processing.

[0161] Furthermore, in each of the above embodiments, it goes without saying that two or more communication means present in one device may be physically realized by one medium.

[0162] Furthermore, in each of the above embodiments, each process may be realized by centralized processing in a single device, or may be realized by distributed processing in a plurality of devices.

[0163] The present invention is not limited to the above-described embodiment, and various modifications are possible, and it goes without saying that these modifications are also included within the scope of the present invention. [Industrial Applicability]

[0164] As described above, the signal processing device 1 according to the present invention has the effect of being able to appropriately perform both time domain analysis and frequency domain analysis, and is useful as, for example, a signal processing device that processes audio signals. [Explanation of symbols]

[0165] 1. Signal Processing Device 11 Storage area 12 Reception 13 Processing section 14 Output section 111 SNN storage unit 112 NN storage unit 113 buffers 131 Information Acquisition Department 132 Frequency conversion unit 133 Signal Transmission Unit 134 Learning Department 135 Information output section

Claims

1. a reception unit that receives a time series signal; an information acquisition unit that acquires information on a time length according to each of the two or more buffering times from the time series signal and temporarily stores the information; a frequency conversion unit that performs frequency conversion on each of the two or more pieces of information temporarily stored by the information acquisition unit and acquires two or more images; a signal transmission unit that transfers the two or more images acquired by the frequency conversion unit to a neural network and acquires output information based on a signal output from the neural network; and an information output unit that outputs the output information.

2. a learning unit that, for each of the two or more buffering times, passes the image acquired by the frequency conversion unit to a sub-neural network corresponding to the buffering time, updates each of the two or more sub-neural networks, and merges the two or more updated sub-neural networks to form a single neural network; The signal transmission unit The signal processing device according to claim 1 , wherein the two or more images acquired by the frequency conversion unit are passed to the one neural network, and output information output from the one neural network is acquired.

3. The signal transmission unit Passing each of the two or more images acquired by the frequency conversion unit to a neural network for each buffering time, and acquiring output information output from each of the two or more neural networks; The information output unit The signal processing device according to claim 1 , which outputs the two or more pieces of output information acquired by the signal transmission unit, or one piece of output information based on the two or more pieces of output information acquired by the signal transmission unit.

4. The signal transmission unit Passing each of the two or more images acquired by the frequency conversion unit to a sub-neural network for each buffering time, acquiring output information output from each of the two or more sub-neural networks, passing the two or more output information to one neural network, and acquiring output information output from the neural network; The information output unit The signal processing device according to claim 3 , wherein the signal transmission unit outputs the acquired output information.

5. The reception unit Accepts two or more time series signals; The information acquisition unit acquiring, for each of the two or more time-series signals and for each of the two or more buffering times, information on time lengths according to the buffering times from the time-series signals; The frequency conversion unit The signal processing device according to claim 1 , wherein, for each of the two or more buffering times, frequency conversion is performed on the information acquired by the information acquisition unit to acquire an image.

6. 5. The signal processing device according to claim 1, wherein the two or more buffering times are three or more buffering times, and three of the three or more buffering times correspond to long-term, medium-term, and short-term time scales, respectively.

7. the time-series signal is an audio signal, The signal processing device according to claim 1 , wherein the output information includes any one of a frequency analysis result for the audio signal, a speech recognition result for the audio signal, a sound source separation result for the audio signal, and a sound source direction estimation result for the audio signal.

8. A signal processing method realized by a reception unit, an information acquisition unit, a frequency conversion unit, a signal transmission unit, and an information output unit, a receiving step in which the receiving unit receives a time-series signal; an information acquisition step in which the information acquisition unit acquires information on a time length according to each of two or more buffering times from the time-series signal and temporarily stores the information; a frequency conversion step in which the frequency conversion unit performs frequency conversion on each of the two or more pieces of information temporarily stored by the information acquisition unit to acquire two or more images; a signal transmission step in which the signal transmission unit passes the two or more images acquired by the frequency conversion unit to a neural network and acquires output information based on a signal output from the neural network; an information output step in which the information output unit outputs the output information.

9. Computer, a reception unit that receives a time series signal; an information acquisition unit that acquires information on a time length according to each of the two or more buffering times from the time series signal and temporarily stores the information; a frequency conversion unit that performs frequency conversion on each of the two or more pieces of information temporarily stored by the information acquisition unit and acquires two or more images; a signal transmission unit that transfers the two or more images acquired by the frequency conversion unit to a neural network and acquires output information based on a signal output from the neural network; A program for causing the information output unit to function as an information output unit that outputs the output information.

Citation Information

Patent Citations

  • Method and system for processing images by determining rotation hypotheses

    JP2022508434A

  • Learning model creation method, learning model creation device, background sound removal method, and background sound removal device

    JP7414214B1

  • Information processing device, information processing method, training model, program, and storage medium

    WO2024135723A1

  • Speech recognition device

    JP2009025480A

  • JPP7414214B