Generation device, generation method, and generation program

The described device and method address the challenge of generating natural language explanations for changes in complex signals by training a generative model on combined signal features, effectively explaining changes in signals and improving inspection efficiency.

JP2025086200APending Publication Date: 2025-06-06HITACHI LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023200103
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-27
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Existing methods struggle to accurately generate natural language explanations of changes in general signals, such as sounds and vibrations from equipment, due to the complexity and non-local nature of these changes, making it difficult to create effective training data and generative models.

Method used

A device and method that combine features from pre-change and post-change signals, along with their differences, to train a generative model. This model generates natural language explanations by processing combined data and using a learning process to adaptively focus on relevant changes.

Benefits of technology

Enables the generation of accurate and detailed natural language explanations of changes in signals, facilitating more efficient inspections and reducing labor hours by clearly identifying significant changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025086200000001_ABST
    Figure 2025086200000001_ABST
Patent Text Reader

Abstract

To learn what is changed between signals depending on a change in condition from signals obtained on two different conditions while allowing for explanation in a natural language.SOLUTION: A generation device including a processor for executing a program and a storage device storing the program therein can access a database including a first pre-signal representing a signal before the change of a first state, a first post-signal representing a signal after the change of the first state, and a first explanatory text explaining states before and after the change by a character string. The processor performs: first connection processing of generating first connection data in which a feature quantity about the first pre-signal, a feature quantity about the first post-signal, and a first difference between the feature quantity about the first pre-signal and the feature quantity about the first post-signal are connected; and learning processing of learning a generation model which generates a character string representing signals before and after the change of the first state, on the basis of the first connection data generated by the first connection processing and the first explanatory text.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a generating device, a generating method, and a generating program for generating a character string. [Background technology]

[0002] It is important to generate a string of text in natural language that explains what has changed in the signals obtained under two different conditions as a result of the change in conditions. For example, abnormalities and signs of abnormalities in equipment or machinery can be automatically detected from the sounds of operation. However, simply showing the presence or absence of abnormalities or signs does not tell the user what to focus on in a detailed manual inspection, which requires a lot of work.

[0003] On the other hand, if it were possible to automatically and clearly present in natural language the differences between previously measured normal sounds and current sounds determined to be abnormal, this would provide clues for the inspector user to carry out a more detailed inspection, further reducing labor hours.

[0004] Non-Patent Document 1 below discloses a method for generating a string of characters that describes in natural language what has changed between two optical images. Non-Patent Document 1 states, "We present a novel Dual Dynamic Attention Model (DUDA) to perform robust Change Captioning. Our model learns to distinguish distractors from semantic changes, localize the changes via Dual Attention over "before" and "after" images, and accurately describe them in natural language via Dynamic Speaker, by adaptively focusing on the necessary visual inputs (e.g. "before" or "after" image)." [Prior art documents] [Non-patent literature]

[0005] [Non-Patent Document 1] Dong Huk Park, Trevor Darrell, and Anna Rohrbach, “Robust Change Captioning,” in arxiv, 17 April 2019. Summary of the Invention [Problem to be solved by the invention]

[0006] In Non-Patent Document 1, optical images are the subject of the detection, and therefore local changes in limited pixel regions, such as object movement, are the main detection target. Therefore, it is relatively clear from the image what changes should be noted. Meanwhile, training data is created by a human being called annotator manually adding correct explanatory strings to two images before and after the change. In the case of optical images, as described above, the changes to be noted are relatively clear, so the annotator can add appropriate explanatory strings. Therefore, appropriate training data can be created, and a generative model for character string generation can be trained based on the training data, so that highly accurate character string generation can be achieved.

[0007] However, when dealing with general signals, especially sounds and vibrations from equipment and machinery, the components are not limited locally in terms of time frequency, but the changes are also across the entire signal components, such as volume, pitch, and the appearance or disappearance of new sound sources. Since there are countless changes between the signals before and after the change, it is not clear which changes should be noted. Therefore, unless an annotator knows what to focus on among the countless changes, it is not possible to provide a desirable description. Therefore, even if appropriate training data cannot be created and a generative model for string generation is trained based on that training data, accurate string generation cannot be achieved.

[0008] The present invention aims to learn, from each of two signals obtained under different conditions, what has changed between the signals due to a change in the conditions in a manner that allows it to be explained in natural language. Another aim of the present invention is to explain, from each of two signals obtained under different conditions, what has changed between the signals due to a change in the conditions in natural language. [Means for solving the problem]

[0009] A generation device according to one aspect of the invention disclosed in the present application has a processor that executes a program and a storage device that stores the program, and is capable of accessing a database having a first pre-signal indicating a state before a change in a first state, a first posterior signal indicating a state after the change in the first state, and a first explanatory sentence that explains the state before and after the change in the form of strings of characters, and the processor executes a first combination process that generates first combined data that combines features related to the first pre-signal, features related to the first posterior signal, and a first difference between the features related to the first pre-signal and the features related to the first posterior signal, and a learning process that learns a generative model that generates strings indicating before and after the change in the first state based on the first combined data generated by the first combination process and the first explanatory sentence.

[0010] A generation device which is another aspect of the invention disclosed in the present application is a generation device having a processor which executes a program and a storage device which stores the program, and is capable of accessing a generation model which has been trained to generate a string indicating before and after a change in a state, and the processor executes a second combination process which generates second combined data which combines features relating to a second pre-event signal which indicates before the change in a second state, features relating to a second posterior signal which indicates after the change in the second state, and a second difference between the features relating to the second pre-event signal and the features relating to the second posterior signal, and a generation process which generates a string indicating before and after the change in the second state by inputting the second combined data generated by the second combination process into the generation model. Effect of the Invention

[0011] According to a representative embodiment of the present invention, it is possible to learn, from each signal obtained under two different conditions, what has changed between the signals due to a change in the conditions in a manner that can be explained in natural language. Also, it is possible to explain, from each signal obtained under two different conditions, what has changed between the signals due to a change in the conditions in natural language. Problems, configurations, and effects other than those described above will become clear from the explanation of the following examples. [Brief description of the drawings]

[0012] [Figure 1] FIG. 1 is a block diagram illustrating an example of a hardware configuration of a generating device. [Diagram 2] FIG. 2 is a block diagram illustrating an example of a functional configuration of the generating device. [Diagram 3] FIG. 3 is a block diagram illustrating an example of a functional configuration of the learning unit. [Figure 4] FIG. 4 is a flowchart illustrating an example of a learning process procedure of the learning unit. [Diagram 5] FIG. 5 is a block diagram illustrating an example of a functional configuration of the generation unit. [Figure 6] FIG. 6 is a flowchart illustrating an example of a generation process procedure of the generation unit. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0013] <Figure 1 Example of hardware configuration of generation device> FIG. 1 is a block diagram showing an example of a hardware configuration of a generating device. The generating device 100 includes a processor 101, a storage device 102, an input device 103, an output device 104, and a communication interface (communication IF) 105. The processor 101, the storage device 102, the input device 103, the output device 104, and the communication IF 105 are connected by a bus 106. The processor 101 controls the generating device 100. The storage device 102 is a working area for the processor 101. The storage device 102 is a non-transient or temporary recording medium that stores various programs and data. Examples of the storage device 102 include a ROM (Read Only Memory), a RAM (Random Access Memory), a HDD (Hard Disk Drive), and a flash memory. The input device 103 inputs data. Examples of the input device 103 include a keyboard, a mouse, a touch panel, a numeric keypad, a scanner, a microphone, and a sensor. The output device 104 outputs data. The output device 104 may be, for example, a display, a printer, or a speaker. The communication IF 105 connects to a network and transmits and receives data.

[0014] <Figure 2 Example of functional configuration of generation device 100> 2 is a block diagram showing an example of a functional configuration of the generating device 100. The generating device 100 includes a training dataset DB 201, a learning unit 202, a generative model 203, and a generating unit 204. The training dataset DB 201 is specifically stored in, for example, the storage device 102 shown in FIG. 1, or another computer capable of communicating with the generating device 100. The learning unit 202, the generative model 203, and the generating unit 204 are specifically realized by, for example, causing the processor 101 to execute a program stored in the storage device 102 shown in FIG. 1.

[0015] The training dataset DB201 is a database that stores one or more training datasets. A training dataset is a combination of training data and correct answer data. The training dataset DB201 has, as a training dataset, a set of triples u {triple 1, , triple u, , triple U} consisting of a prior signal time waveform set 2u1, a posterior signal time waveform set 2u2, and an explanatory sentence 2u3.

[0016] The prior signal time waveform set 2u1 is a set of prior signal time waveforms. The prior signal time waveform is training data indicating the time waveform of a prior signal. The prior signal is a signal under a condition before a certain state change, for example, a steady sound, a periodic sound, or a non-periodic sound of the equipment to be inspected.

[0017] The posterior signal time waveform set 2u2 is a set of posterior signal time waveforms. The posterior signal time waveform is training data showing the time waveform of a posterior signal. The posterior signal is a signal under a condition after a certain state change, for example, an abnormal sound of the inspection target device that has changed from a steady sound in a state before the change.

[0018] When there is no distinction between pre-event signal time waveforms and post-event signal time waveforms, they are referred to as signal time waveforms.

[0019] The explanatory text 2u3 is a variable-length text including onomatopoeia that expresses the change between the pre-event time waveform and the post-event time waveform.

[0020] An example of an explanatory sentence expressing a change in the bearings of a rotating body from normal to abnormal is as follows: The character strings in quotation marks are onomatopoeia added by the annotator.

[0021] "The sound changed from 'bo' to 'woo' and the pitch became higher." "The 'win-win' sound has disappeared." The buzzing and hissing sounds became higher in pitch and louder."

[0022] Here, the annotator is instructed to focus on the changes, create an explanatory text expressing the changes, and use this as ground truth data for the generative model 203, enabling the generative model 203 to explain how the steady sound differs from the sound currently determined to be abnormal.

[0023] The onomatopoeic annotation by annotators is extremely important in providing information that will serve as clues for detailed inspection by inspectors. This is because if annotators were given only the post-signal time waveform and asked to answer questions such as "what kind of sound is it?" or "what kind of sound is it?", they would only be able to obtain answers that are independent of the changes, such as "the sound of a bearing."

[0024] In addition, there is a problem that the explanation using only the plain text that does not include onomatopoeia cannot express in detail what sounds have changed and how. In other words, without the onomatopoeia expression, not only cannot the generative model 203 express in detail, but also the annotator cannot explain well, so it is not possible to create a training dataset to be used for learning the generative model 203.

[0025] Therefore, by creating explanatory text including onomatopoeia in the annotation, the annotator can describe in detail what sounds have changed and how they have changed. Using the explanatory text created in this way, a generative model 203 that can describe changes in detail can be realized.

[0026] It is also possible to have the annotator answer not with onomatopoeia but with a classification of what sound it is (for example, "the sound of a bearing"); however, such expressions tend to increase the vocabulary as the number of usage scenes increases, but since the increased vocabulary is not used in different scenes, it is difficult to obtain a general-purpose model that spans across scenes. Therefore, by focusing on the fact that onomatopoeia can be used generically across scenes, and having the annotator create an explanatory text that includes onomatopoeia in the annotation, it is possible to obtain a general-purpose generative model 203 that spans across scenes.

[0027] The learning unit 202 randomly selects a triplet u from a set of U triplets {triplet 1, , triplet u, , triplet U}. As described above, the triplet u is composed of a prior signal time waveform set 2u1, a posterior signal time waveform set 2u2, and an explanatory sentence 2u3. Furthermore, the learning unit 202 randomly selects one element from the prior signal time waveform set 2u1 of the triplet to set it as a prior signal time waveform 301, and randomly selects one element from the posterior signal time waveform set 2u2 to set it as a posterior signal time waveform 302, sets the combination of the prior signal time waveform 301 and the posterior signal time waveform 302 as an explanatory variable, and sets the explanatory sentence 2u3 to the explanatory sentence 303, which is a target variable.

[0028] The learning unit 202 uses the triplet to learn the generative model 203. Specifically, for example, the learning unit 202 calculates the value of a loss function based on the difference between the output data output as a result of inputting the set explanatory variables to the generative model 203 and the explanatory text 303, and updates the parameters of the generative model 203 so that the value of the loss function is minimized.

[0029] The generation model 203 is a language model that outputs an explanatory sentence when a signal time waveform is input. The generation model 203 is trained by the training unit 202, and the generation unit 204 generates an inference explanatory sentence 243.

[0030] The generation unit 204 inputs the prior signal time waveform 241 and the posterior signal time waveform 242 to the generative model 203, and outputs an inference explanation 243 from the generative model 203. The prior signal time waveform 241 may be a prior signal time waveform in the prior signal time waveform set 211, or may be a prior signal time waveform different from the prior signal time waveform in the prior signal time waveform set 211. The posterior signal time waveform 242 may be a posterior signal time waveform in the posterior signal time waveform set 212, or may be a posterior signal time waveform different from the posterior signal time waveform in the posterior signal time waveform set 212.

[0031] <Figure 3: Example of functional configuration of learning unit 202> 3 is a block diagram showing an example of a functional configuration of the learning unit 202. The learning unit 202 includes frame division units 311 and 312, window function multiplication units 321 and 322, frequency domain signal generation units 313 and 323, encoding units 351 and 352, a feature difference calculation unit 353, a feature combination unit 354, a decoding unit 355, an onomatopoeia-to-phoneme conversion unit 331, an onomatopoeia-to-subword conversion unit 332, and a learning processing unit 356.

[0032] The frame division units 311 and 312 divide the signal time waveform into frames. Each of the divided signal time waveforms is called a frame division signal.

[0033] The window function multiplication units 321 and 322 perform window function multiplication on the frame divided signals to convert each of the frame divided signals into a window function-multiplied signal.

[0034] The frequency domain signal generators 313 and 323 perform a short-time Fourier transform on each of the window function multiplied signals to convert them into time-frequency domain signals. The frequency domain signal generators 313 and 323 can also use a frequency transform method such as a constant Q transform (CQT) instead of the short-time Fourier transform.

[0035] The encoding units 351 and 352 calculate feature vectors based on frequency domain signals. The encoding units 351 and 352 are typically neural network encoders in which multiple convolutional layers, activation functions, and pooling layers are stacked with skip connections in between. The encoding units 351 and 352 may be recurrent neural networks having layers such as a known Transformer model, Long-Short-Term-Memory (LSTM), bidirectional LSTM, Gated recurrent unit (GRU), and bidirectional GRU.

[0036] A feature amount difference calculation unit 353 calculates a difference vector, which is the difference between the feature amount vector from the encoding unit 351 and the feature amount vector from the encoding unit 352. The difference vector is a feature amount that emphasizes changes, and by learning using this feature amount, it is possible to generate a generative model 203 that generates an explanatory sentence that emphasizes changes.

[0037] The feature amount combining unit 354 combines the feature amount vector from the encoding unit 351, the feature amount vector from the encoding unit 352, and the difference vector to generate a combined vector.

[0038] The decoding unit 355 receives the combined vector from the feature combining unit 354 as an input, and generates a variable-length text in which the onomatopoeic phonemes are converted into subwords, similar to the subword-converted explanation 343 described later. The decoding unit 355 is typically a decoder of a known Transformer model, which is a type of neural network. The decoding unit 355 may be a recurrent neural network having layers such as a Long-Short-Term-Memory (LSTM), a bidirectional LSTM, a gated recurrent unit (GRU), or a bidirectional GRU. The neural network used by the decoding unit 305 is referred to as a decoding model.

[0039] The onomatopoeia phoneme conversion unit 331 extracts character strings enclosed in parentheses as onomatopoeia from the explanatory sentence 303, converts the extracted onomatopoeia into a phoneme string, and generates onomatopoeia phoneme converted text. For example, if the onomatopoeia is "kankan," the phoneme string is / ka N ka N / . Also, if the onomatopoeia is "katakata don," the phoneme string is / katakatado: N / . Therefore, the example of the explanatory sentence 303 mentioned above, "The pitch of the sounds "boon" and "sha" has increased, and the volume has increased." is converted to "The pitch of the sounds / bu: N / and / sh a: / has increased, and the volume has increased."

[0040] The onomatopoeia subword generating unit 332 generates subwords from the onomatopoeia phoneme-converted text. Specifically, for example, the onomatopoeia subword generating unit 332 outputs partial character strings cut out for each predetermined number of characters n (gram number) while shifting the target range for the onomatopoeia phoneme string in the onomatopoeia phoneme-converted text by one character at a time.

[0041] For example, if the onomatopoeia is "katakatadon," the original phoneme ( / katakatado: N / ) is converted into the following phoneme subwords (assuming n=4):

[0042] / kata / / atak / / taka / / akat / / kata / / atad / / tado: / / ado: N /

[0043] Thereafter, each line of n=4 characters is treated as one word. In this way, a subword explanation sentence 343 is generated in which only onomatopoeia words are converted into phoneme subwords.

[0044] We will now discuss the effect of subwording onomatopoeia. Compared to normal words, onomatopoeia have a sparse frequency of occurrence. For example, "katakatadon" rarely appears in other scenes. Therefore, if it is input directly into a language model, similar onomatopoeia will be distinguished as completely different words, and training will not be possible due to a lack of training data per word. Subwording has the effect of preventing a lack of training data by breaking down "katakatadon" into high-frequency phoneme strings such as "kata" and "taka."

[0045] The learning processing unit 356 compares the variable-length text generated by the decoding unit 355 (however, like the sub-word-converted explanation sentence 343, the onomatopoeia phonemes have been sub-converted) with the sub-word-converted explanation sentence 343 generated by the onomatopoeia sub-word conversion unit 332, and updates the parameters of the neural network models of the encoding units 351, 352 and the decoding unit 355 so as to minimize the cross entropy L given below.

[0046]

number

[0047] Here, K_u is the total number of elements belonging to the prior signal time waveform set 2u1, and k is a number that uniquely identifies the element. I_u is the total number of elements belonging to the posterior signal time waveform set 2u2, and i is a number that uniquely identifies the element. T is the number of words that appear in the subword-converted explanation 343. t is a number that uniquely identifies the word. w(t) is the probability that the t-th word is correctly estimated, and can be calculated by comparing the variable-length text generated by the decoding unit 355 (however, like the subword-converted explanation 343, the onomatopoeic phonemes are subworded) with the subword-converted explanation 343. w_1:t-1 represents a sequence of words from t=1 to t=t-1. X is a combination vector. Optimization can be performed using known optimization algorithms such as SGD, Momentum SGD, AdaGrad, RMSProp, AdaDelta, and Adam.

[0048] The combination of the neural networks (encoding model) of the encoding units 351 and 352 and the neural network (decoding model) of the decoding unit 355, whose parameters have been updated, becomes the generation model 203.

[0049] <Fig. 4 Learning process procedure of the learning unit 202> FIG. 4 is a flowchart showing an example of a learning process procedure of the learning unit 202.

[0050] (Step S401) The learning processing unit 356 judges whether the value of the loss function converges. Specifically, for example, the learning processing unit 356 judges whether a convergence judgment condition is satisfied or whether the number of iterations C1 is greater than a threshold value ThC. The convergence judgment condition is, for example, a condition that the convergence judgment function becomes smaller than a predetermined threshold value.

[0051] If the convergence determination condition is not satisfied, or if the number of iterations C1 does not exceed the threshold value ThC (step S401: No), the process proceeds to step S402. If the convergence determination condition is satisfied, or if the number of iterations C1 exceeds the threshold value ThC (step S401: Yes), the process proceeds to step S419, determining that the value of the loss function has converged.

[0052] (Step S402) The learning unit 202 randomly selects a triplet u from the training data set DB 201. As described above, the triplet u is composed of a prior signal time waveform set 2u1, a posterior signal time waveform set 2u2, and an explanatory sentence 2u3. Furthermore, the learning unit 202 randomly selects one element from 2u1 of the triplet to set it as a prior signal time waveform 301, and randomly selects one element from 2u2 to set it as a posterior signal time waveform 302, sets the combination of the prior signal time waveform 301 and the posterior signal time waveform 302 as an explanatory variable, and sets the explanatory sentence 2u3 as an explanatory sentence 303, which is an objective variable.

[0053] (Step S403) The onomatopoeia-phoneme conversion unit 331 extracts onomatopoeia from the explanatory text 303, converts them into a phoneme string, and generates onomatopoeia-phoneme converted text.

[0054] (Step S404) The onomatopoeia subword generating unit 332 generates subwords from the onomatopoeia phoneme converted text converted by the onomatopoeia phoneme converting unit 331 , and generates a subword converted explanation sentence 343 .

[0055] (Step S405) The frame division unit 311 divides the advance signal time waveform into frames. The frame division signal from the frame division unit 311 is called an advance frame division signal.

[0056] (Step S406) The window function multiplication unit 321 performs window function multiplication on the pre-frame divided signals to convert each of the pre-frame divided signals into a window function multiplied signal. This window function multiplied signal is called a pre-window function multiplied signal.

[0057] (Step S407) The frequency domain signal generator 313 performs a short-time Fourier transform on each of the pre-window function multiplied signals to convert them into time-frequency domain signals, which are referred to as pre-time-frequency domain signals.

[0058] (Step S408) The encoding unit 351 calculates a feature vector based on the a priori frequency domain signal. This feature vector is called a priori feature vector.

[0059] (Step S409) The frame division unit 312 divides the post-signal time waveform into frames. The frame division signal from the frame division unit 312 is referred to as a post-frame division signal.

[0060] (Step S410) The window function multiplier 322 performs window function multiplication on the post-frame divided signals to convert each of the post-frame divided signals into a window function multiplied signal, which is referred to as a post-window function multiplied signal.

[0061] (Step S411) The frequency domain signal generator 323 performs a short-time Fourier transform on each of the post-window function multiplied signals to convert them into time-frequency domain signals, which are referred to as post-time-frequency domain signals.

[0062] (Step S412) The encoding unit 352 calculates a feature vector based on the posterior frequency domain signal. This feature vector is called a posterior feature vector.

[0063] Note that steps S409 to S412 may be executed in parallel with steps S405 to S408.

[0064] (Step S413) The feature amount difference calculation unit 353 calculates a difference vector which is the difference between the pre-feature amount vector and the posterior feature amount vector.

[0065] (Step S414) The feature combination unit 354 combines the pre-feature vector, the posterior feature vector, and the difference vector to generate a combined vector.

[0066] (Step S415) The decoding unit 355 generates a variable-length text based on the combined vector generated in step S414.

[0067] (Step S416) The learning processing unit 356 updates the parameters of the neural networks of the encoding units 351 and 352 and the decoding unit 355 using the above-mentioned equation (1).

[0068] (Step S417) The learning processing unit 356 calculates the convergence conditions.

[0069] (Step S418) The learning processing unit 356 increments the number of iterations C1, and then returns to step S404.

[0070] (Step S419) If step S404: Yes, the learning processing unit 356 saves the parameters updated in step S416 in the storage device 102 as parameters of the generative model 203.

[0071] <Fig. 5 Example of functional configuration of the generation unit 204> 5 is a block diagram showing an example of a functional configuration of the generation unit 204. The generation unit 204 has frame division units 311 and 312, window function multiplication units 321 and 322, frequency domain signal generation units 313 and 323, encoding units 351 and 352, a feature difference calculation unit 353, a feature combination unit 354, and a decoding unit 355. In other words, a part of the configuration of the learning unit 202 also functions as the generation unit 204.

[0072] <FIG. 6 Generation process procedure of the generation unit 204> FIG. 6 is a flowchart showing an example of a generation process procedure of the generation unit 204.

[0073] (Step S601) The generator 204 reads the generative model 203 .

[0074] (Steps S602 to S604) The generating unit 204 performs the same processes as steps S405 to S407 on the input prior signal time waveform 241.

[0075] (Step S605) The encoding unit 351 uses the generative model 203 to calculate a priori feature vector based on the priori time-frequency domain signal from step S604.

[0076] (Steps S606 to S608) The generating unit 204 performs the same processes as steps S409 to S411 on the input posterior signal time waveform 242.

[0077] (Step S609) The encoding unit 352 uses the generative model 203 to calculate a posterior feature vector based on the posterior time-frequency domain signal from step S608.

[0078] (Step S610) The feature difference calculation unit 353 calculates a difference vector, which is the difference between the pre-feature vector and the posterior feature vector. The difference vector is a feature that emphasizes changes, and by performing inference using this feature, it is possible to generate an explanation that emphasizes changes.

[0079] (Step S611) The feature combination unit 354 combines the pre-feature vector, the posterior feature vector, and the difference vector to generate a combined vector.

[0080] (Step S612) The decoding unit 355 uses the generative model 203 and receives the combined vector as an input to generate a variable-length text including a string of onomatopoeic subwords.

[0081] (Step S613) The decoding unit 355 converts the onomatopoeic subword string in the variable-length text generated in step S612 back into onomatopoeic text via the onomatopoeic phoneme string. In this way, the variable-length text including the onomatopoeic subwords is converted into the final inference explanation sentence 243.

[0082] First, a method of inversely converting an onomatopoeia subword string of a variable-length text into an onomatopoeia phoneme string will be described. During learning (step S403), when converting an onomatopoeia phoneme string into an onomatopoeia subword string, the onomatopoeia subword generating unit 332 creates an onomatopoeia subword string by shifting each character while overlapping the range by n-1 characters.

[0083] In the inverse conversion, the decoding unit 355 extracts only the first character v_m1 from the m-th sub-word s_m=[v_m1, ..., v_mn] for the phonemes m=1, ..., M-1 in the onomatopoeia sub-word sequence S=[s_1, ..., s_M] consisting of M onomatopoeia sub-words s_m (m=1, ..., M).

[0084] The decoding unit 355 extracts the entire character string s_M=[v_M1, ..., v_Mn] for s_M of the last sub-word m=M. That is, it generates [v_11, v_21, v_31, ..., v_{M-1}1, v_M1, ..., v_Mn] as the phoneme string of the onomatopoeia.

[0085] Next, in the reverse conversion of the onomatopoeia phoneme string to onomatopoeia text, the decoding unit 355 may convert according to the correspondence table since there is a one-to-one relationship between phonemes and katakana. In this way, the onomatopoeia text is restored, and the generation of the entire inference explanation sentence 243 is completed.

[0086] <Switching models according to the type of signal you are interested in> Signals observed as prior signals and posterior signals are roughly divided into three types, such as stationary signals, periodic signals, and non-periodic signals. Depending on the type of signal, different types of models are suitable as models (hereinafter, encoding models) for the encoding units 351 and 352. For example, for stationary signals, a network equipped with a spatial attention mechanism is suitable, and has higher accuracy than a Transformer.

[0087] Transformer is suitable for periodic and non-periodic signals, and has better accuracy than networks equipped with a spatial attention mechanism. Also, it is more accurate to prepare separate coding models for stationary, periodic, and non-periodic signals. Therefore, the generating device 100 constructs these three types of coding models as the generating model 203 as follows. As a prerequisite, a training dataset DB 201 is prepared for each type of signal.

[0088] The training dataset DB201 for stationary signals is composed of a set of pre-signal time waveforms 211 of stationary signals, a set of post-signal time waveforms 212 of stationary signals, and a set of explanations 213 in which an annotator explains the changes between them. In addition, the training dataset DB201 specialized for stationary signals can be constructed by instructing the annotator to provide an explanation "focusing on stationary signals." Then, by preparing a network equipped with a spatial attention mechanism suitable for stationary signals as described above as an encoding model, the learning unit 202 executes learning of the generative model 203.

[0089] The training dataset DB201 for periodic signals is composed of a set of prior signal time waveforms 211 of periodic signals, a set of posterior signal time waveforms 212 of periodic signals, and a set of explanations 213 in which an annotator explains the changes between them. In addition, the training dataset DB201 specialized for periodic signals can be constructed by instructing the annotator to provide an explanation "focusing on periodic signals." Then, by preparing a Transformer suitable for periodic signals as described above as an encoding model, the learning unit 202 executes learning of the generative model 203.

[0090] The training data set DB201 for non-periodic signals is composed of a set of prior signal time waveforms of non-periodic signals, a set of posterior signal time waveforms of non-periodic signals, and a set of explanations 213 in which an annotator explains the changes between them. In addition, by instructing the annotator to provide an explanation "focusing on non-periodic signals", a training data set DB201 specialized for non-periodic signals can be constructed. Then, by preparing a Transformer suitable for non-periodic signals as described above as an encoding model, the learning unit 202 executes learning of the generative model 203.

[0091] The generator 204 switches between the above three types of generative models 203. For example, in a usage scene where it is known that attention should be paid to a specific type of signal, the generator 203 of that type is specified and executed, so that the inference explanation 243 can be generated with high accuracy specialized for the specified type of signal and without being adversely affected by other noises.

[0092] The generating unit 204 may use the above three types of generative models 203 in parallel at the same time. For example, the inference explanation 243 output from the generative model 203 for stationary signals may be prefaced with a character string "For stationary signals, " the inference explanation 243 output from the generative model 203 for periodic signals may be prefaced with a character string "For periodic signals, " the inference explanation 243 output from the generative model 203 for non-periodic signals may be prefaced with a character string "For non-periodic signals, " and the three explanations may be connected and output in a differentiated manner. This allows the user to simultaneously read explanations from multiple perspectives corresponding to the three types of generative models 203, which makes it easy to compare the perspectives and to easily gain insight.

[0093] Although the generation unit 204 has been described here as executing three types of generation models 203 in parallel as an example, the generation unit 204 may execute two types of generation models 203 in parallel among the three types of generation models 203. Furthermore, if there are signals other than the above-mentioned three types, the generation unit 204 may execute four or more types of generation models 203 in parallel.

[0094] As described above, the two conditions are A and B, and for sets S_A and S_B consisting of one or more sample signals corresponding to condition A (e.g., normal (before)) and condition B (e.g., abnormal (after)), respectively, an annotator is asked to assign a string C that explains the difference between sets S_A and S_B, and the triplet set of set S_A, set S_B, and string C is used as the training dataset.

[0095] During learning, the generating device 100 inputs each element of the set S_A and the set S_B and trains the generative model 203 to output a character string C. During inference, the generating device 100 uses the generative model 203 to generate an inference explanation sentence 234 from each signal of the condition A and the condition B.

[0096] Annotators who do not know what to focus on among countless changes tend to explain only outliers rather than significant changes when comparing signal samples one-to-one. In contrast, in this embodiment, even if annotators do not have specialized knowledge, they can find significant changes between conditions A and B regardless of outliers by comparing samples in multiple different pairs with different conditions, and assign character string C.

[0097] Therefore, for example, the learning unit 202 repeatedly selects combinations of prior signal time waveforms and posterior signal time waveforms from the triplet u selected as the training data set thus created, and uses the joint vector and explanation 2u3 repeatedly generated for each selected combination to learn the generative model 203. This enables the generation unit 204 to generate an inference explanation 234 that focuses on significant changes due to changes in condition A and condition B during inference using the generative model 203.

[0098] In this way, the above-described generating device 100 can learn or infer, as if explaining in natural language, what has changed between signals due to changes in conditions from each signal obtained under two different conditions, even if the annotator does not know in advance what to focus on among the countless changes. This allows the annotator to easily identify what to focus on among the countless changes.

[0099] In the above-described embodiment, the generating device 100 has a learning unit 202 and a generating unit 204, but it is also possible for the generating device 100 to have either one of the learning unit 202 or the generating unit 204, and for the other to be possessed by another computer capable of communicating with the generating device 100.

[0100] In the above embodiment, a sound signal has been taken as an example, but the same means can be used for an ultrasonic sensor signal. The same configuration can also be used for general time-series signals such as the time waveform of an acceleration sensor or a displacement sensor, the time waveform of a current sensor, and financial indicators such as stock prices and exchange rates. In the case of the time waveform of a current sensor, and financial indicators such as stock prices and exchange rates, they are not "onomatopoeic words" but "mimetic words," and such onomatopoeic words and mimetic words can be used as onomatopoeias to apply to the onomatopoeic word-phoneme conversion unit 331 and the onomatopoeic word subword conversion unit 332.

[0101] The present invention is not limited to the above-described embodiments, and includes various modified examples and equivalent configurations within the spirit of the appended claims. For example, the above-described embodiments have been described in detail to clearly explain the present invention, and the present invention is not necessarily limited to those having all of the configurations described. Also, a part of the configuration of one embodiment may be replaced with a configuration of another embodiment. Also, a configuration of another embodiment may be added to a configuration of one embodiment. Also, a part of the configuration of each embodiment may be added, deleted, or replaced with another configuration.

[0102] Furthermore, each of the aforementioned configurations, functions, processing units, processing means, etc. may be realized in hardware, for example by designing some or all of them as an integrated circuit, or may be realized in software by a processor interpreting and executing a program that realizes each function.

[0103] Information such as programs, tables, files, etc. that realize each function can be stored in a storage device such as a memory, a hard disk, or an SSD (Solid State Drive), or in a recording medium such as an IC (Integrated Circuit) card, an SD card, or a DVD (Digital Versatile Disc).

[0104] In addition, the control lines and information lines shown are those considered necessary for the explanation, and do not necessarily show all the control lines and information lines necessary for implementation. In reality, it can be considered that almost all components are connected to each other. [Explanation of symbols]

[0105] 100 generator 101 Processor 102 Storage Devices 201 Training Dataset DB 202 Learning Department 203 Generative Model 204 Generation part 211 Pre-signal time waveform set 212 Post-signal time waveform set 213 Description collection 234 Inference Explanation 241 Pre-signal time waveform 242 Post-signal time waveform 243 Inference Explanation 303 Description 305 Decoding section 311,312 Frame division section 312,322 Window function multiplication section 313,323 Frequency domain signal generator 331 Onomatopoeia Phoneme Conversion Unit 332 Onomatopoeia Subword Generation Unit 343 Subword Description 351,352 Encoding section 353 Feature Difference Calculation Unit 354 Feature Combination Unit 355 Decoding section 356 Learning Processing Unit

Claims

1. A generating device having a processor that executes a program and a storage device that stores the program, a database having a first pre-signal indicating a state before a change in a first state, a first posterior signal indicating a state after the change in the first state, and a first description describing the state before and after the change in a character string; The processor, a first combining process for generating first combined data by combining a feature amount related to the first pre-event signal, a feature amount related to the first posterior signal, and a first difference between the feature amount related to the first pre-event signal and the feature amount related to the first posterior signal; a learning process for learning a generative model that generates a character string indicating before and after a change in the first state based on the first combined data generated by the first combination process and the first description; A generating device characterized by executing the above.

2. The generating device according to claim 1 , The processor, executing a first encoding process for outputting features related to the first a priori signal and features related to the first posterior signal by inputting the first a priori signal and the first posterior signal into an encoding model that encodes a signal and outputs features related to the signal; In the first combination process, the processor generates a feature amount related to the first prior signal output by the first encoding process, a feature amount related to the first posterior signal output by the first encoding process, and the first combined data. A generating device characterized by:

3. The generating device according to claim 2, The processor, executing a first decoding process for outputting a first decoded explanation based on a difference between the first prior signal and the first posterior signal by inputting the first combined data into a decoding model that outputs a variable-length text when a feature is input; In the learning process, the processor learns the encoding model and the generation model, which is the decoding model, based on the first explanatory sentence and a first decoded explanatory sentence output by the first decoding process. A generating device characterized by:

4. The generating device according to claim 1 , The processor, A sub-word generation process is executed to generate a plurality of sub-words based on a specific character string in the first description sentence and generate a plurality of sub-words for the first description sentence; In the learning process, the processor learns the generative model based on the first combined data and a plurality of the first explanatory sentences generated by the subword generation process. A generating device characterized by:

5. The generating device according to claim 1 , The processor, a second combining process for generating second combined data by combining a feature amount related to a second pre-event signal indicating a state before a change in a second state, a feature amount related to a second posterior signal indicating a state after a change in the second state, and a second difference between the feature amount related to the second pre-event signal and the feature amount related to the second posterior signal; a generation process for generating a character string indicating before and after the change in the second state by inputting the second combined data generated by the second combination process into the generative model; A generating device characterized by executing the above.

6. The generating device according to claim 5, The processor, executing a second encoding process that outputs features related to the second a priori signal and features related to the second posterior signal by inputting the second a priori signal and the second posterior signal into an encoding model that encodes a signal and outputs features related to the signal; In the second combining process, the processor generates a feature amount related to the second prior signal output by the second encoding process, a feature amount related to the second posterior signal output by the second encoding process, and the second combined data. A generating device characterized by:

7. 7. The generating device of claim 6, The processor, a second decoding process is executed in which the second combined data is input to a decoding model that outputs a variable-length text when a feature is input, and a second decoded explanation based on a difference between the second pre-signal and the second posterior signal is output as a character string indicating before and after the change in the second state. A generating device characterized by:

8. The generating device according to claim 1 , The first description sentence includes onomatopoeia before and after the change as the state before and after the change, A generating device characterized by:

9. The generating device according to claim 1 , The database stores a prior set having a plurality of the first prior signals and a posterior set having a plurality of the first posterior signals; In the first combining process, the processor repeatedly selects combinations of the first a priori signal and the first posterior signal from the a priori set and the posterior set, and generates the first combined data for each selected combination of the first a priori signal and the first posterior signal. A generating device characterized by:

10. The generating device according to claim 1 , The database stores a plurality of types of the first pre-signals and the first posterior signals; In the first combining process, the processor generates the first combined data for each of the plurality of types of data; In the learning process, the processor learns the generative model for each of the types. A generating device characterized by:

11. The generating device according to claim 10, The plurality of types includes at least two types of stationary, periodic, and non-periodic signal types. A generating device characterized by:

12. A generating device having a processor that executes a program and a storage device that stores the program, have access to a generative model that has been trained to generate strings of characters that represent before and after state changes; The processor, a second combining process for generating second combined data by combining a feature amount related to a second pre-event signal indicating a state before a change in a second state, a feature amount related to a second posterior signal indicating a state after a change in the second state, and a second difference between the feature amount related to the second pre-event signal and the feature amount related to the second posterior signal; a generation process for generating a character string indicating before and after the change in the second state by inputting the second combined data generated by the second combination process into the generative model; A generating device characterized by executing the above.

13. 13. The generating device of claim 12, The generative model for each of a plurality of types of the second prior signal and the second posterior signal is accessible; In the second combining process, the processor generates the second combined data for each of the types; In the generation process, the processor generates the character string by inputting the second combined data into the generative model for each of the types, and outputs a character string obtained by distinguishing and concatenating the character strings for each of the types. A generating device characterized by:

14. 14. The generating device of claim 13, The plurality of types includes at least two types of stationary, periodic, and non-periodic signal types. A generating device characterized by:

15. A generation method executed by a generation device having a processor that executes a program and a storage device that stores the program, comprising: a database having a first pre-signal indicating a state before a change in a first state, a first posterior signal indicating a state after the change in the first state, and a first description describing the state before and after the change in a character string; The processor, a first combining process for generating first combined data by combining a feature amount related to the first pre-event signal, a feature amount related to the first posterior signal, and a first difference between the feature amount related to the first pre-event signal and the feature amount related to the first posterior signal; a learning process for learning a generative model that generates a character string indicating before and after a change in the first state based on the first combined data generated by the first combination process and the first description; A generating method comprising:

16. A generation method executed by a generation device having a processor that executes a program and a storage device that stores the program, comprising: have access to a generative model that has been trained to generate strings of characters that represent before and after state changes; The processor, a second combining process for generating second combined data by combining a feature amount related to a second pre-event signal indicating a state before a change in a second state, a feature amount related to a second posterior signal indicating a state after a change in the second state, and a second difference between the feature amount related to the second pre-event signal and the feature amount related to the second posterior signal; a generation process for generating a character string indicating before and after the change in the second state by inputting the second combined data generated by the second combination process into the generative model; A generating method comprising:

17. a processor that can access a database having a first pre-event signal indicating a state before a change in a first state, a first post-event signal indicating a state after the change in the first state, and a first explanatory text that explains the states before and after the change in a character string; a first combining process for generating first combined data by combining a feature amount related to the first pre-event signal, a feature amount related to the first posterior signal, and a first difference between the feature amount related to the first pre-event signal and the feature amount related to the first posterior signal; a learning process for learning a generative model that generates a character string indicating before and after a change in the first state based on the first combined data generated by the first combination process and the first description; A generating program for causing a user to execute the above steps.

18. A processor with access to a generative model that has been trained to generate strings of characters before and after a state change. a second combining process for generating second combined data by combining a feature amount related to a second pre-event signal indicating a state before a change in a second state, a feature amount related to a second posterior signal indicating a state after a change in the second state, and a second difference between the feature amount related to the second pre-event signal and the feature amount related to the second posterior signal; a generation process for generating a character string indicating before and after the change in the second state by inputting the second combined data generated by the second combination process into the generative model; A generating program for causing a user to execute the above steps.