Authenticity determination device, authenticity determination method, and program
The system addresses the challenge of distinguishing real from generated sound data by altering sound characteristics in the analog domain and using structural parameters and machine learning to enhance authenticity determination accuracy.
Patent Information
- Application Number
- PCT/JP2025/009661
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-28
- Filing Date
- 2025-03-13
- Publication Date
- 2025-10-02
AI Technical Summary
Existing technologies struggle to accurately distinguish between real sound data and artificially generated sound data, particularly due to advancements in AI technology.
A system comprising a sound receiving unit that changes sound characteristics in the analog domain, generating restored sound signals using a restoration unit and discriminating authenticity through a discrimination unit, utilizing structural parameters and machine learning to decode and identify the original sound source.
Enhances the accuracy of determining sound signal authenticity by restoring and decoding sound signals based on unique structural parameters, making it difficult to imitate real sound data.
Smart Images

Figure JP2025009661_02102025_PF_FP_ABST
Abstract
Description
Authenticity determination device, authenticity determination method, and program
[0001] The present technology relates to an authenticity discrimination device, an authenticity discrimination method, and a program, and in particular to a technology for discriminating the authenticity of a sound signal.
[0002] Conventionally, a device has been proposed that protects speech privacy by performing masking processing of time delay and amplitude adjustment on an audio signal received by a microphone to blur the audio (for example, Patent Document 1).
[0003] Special Publication No. 2020-514819
[0004] In recent years, it has become difficult to distinguish whether digital data is data obtained by actually recording sound (hereinafter referred to as "real sound data"). For example, with the recent development of AI (Artificial Intelligence) technology, it has become difficult to distinguish counterfeit data, so-called generated data, from real sound data.
[0005] The present technology has been made in view of the above circumstances, and aims to improve the accuracy of determining the authenticity of sound signals.
[0006] The recording device according to the present technology includes a restoration unit that generates restored sound by restoring a sound signal generated by collecting sound whose characteristics have been changed by a sound receiving unit in the analog domain when collecting sound, and a discrimination unit that discriminates the authenticity of the restored sound generated by the restoration unit. This makes it possible to generate restored sound only when parameters that change the characteristics of the sound received by the sound receiving unit are known.
[0007] 1 is a diagram showing the process by which sound is collected. FIG. 2 is a diagram showing an overview of an authenticity discrimination system. FIG. 3 is a diagram showing the structure of a sound receiving unit. FIG. 4 is a diagram showing the structure of a sound collecting unit. FIG. 5 is a diagram showing masking by a mask hole. FIG. 6 is a diagram showing the structural parameters of an encoding unit and sound characteristics that change depending on those structural parameters. FIG. 7 is a diagram showing the function of a sound collecting unit. FIG. 8 is a diagram showing the structural parameters of the sound collecting unit and sound characteristics that change depending on those structural parameters. FIG. 9 is a diagram explaining the processing of a restoration unit. FIG. 10 is a diagram showing a two-dimensional map. FIG. 11 is a diagram for explaining specific processing of a restoration unit and a discrimination unit. FIG. 12 is a diagram showing examples of displaying identification results, discrimination results, and declaration results. FIG. 13 is a flowchart showing the processing procedure of a sound recording device. FIG. 14 is a flowchart showing the processing procedure of an authenticity discrimination device. FIG. 15 is a block diagram showing an example of the hardware configuration of an information processing device equivalent to an authenticity discrimination device. FIG. 16 is a diagram showing the structure of a sound receiving unit of a modified example.
[0008] Hereinafter, the embodiments will be described in the following order: <1. Authentication determination system> [1.1 General configuration of the authentication determination system] [1.2 Specific configuration of the recording device] [1.3 Specific configuration of the authentication determination device] [1.4 Processing procedure] [1.5 Hardware configuration] <2. Modifications> <3. Summary of the embodiment> <4. The present technology>
[0009] 1. Authentication Discrimination System First, the terms used in describing the authentication discrimination system according to the embodiment will be described with reference to Fig. 1. In the embodiment, authentication is performed to verify whether or not sound data is generated data, that is, whether or not the data is obtained by picking up actual sound.
[0010] Figure 1 is a diagram showing the process of collecting sound. In the process of collecting sound as shown in Figure 1, the region before the sound is collected by a microphone is referred to as the "analog region." Furthermore, the region after the sound is collected by the microphone is referred to as the "signal region."
[0011] Here, in this specification, "generated data" means data that is not actually recorded sound but is artificially generated to imitate sound. Also, in this specification, the term "actual sound data" is used as the antonym of "generated data." In other words, "actual sound data" means data that represents actual sound, unlike generated data.
[0012] 1.1 Schematic Configuration of the Authentication Discrimination System As shown in Fig. 2, the authentication discrimination system 1 includes a recording device 2 that collects sound and generates real sound data, and an authenticity discrimination device 3 that discriminates the authenticity of the acquired sound data (particularly the real sound data 100).
[0013] The sound recording device 2 includes a sound receiving unit 11 and an AD conversion unit 12. The sound receiving unit 11 includes an encoding unit 13 and a sound collection unit 14, and collects sound by changing the characteristics of the sound in the analog domain when collecting the sound, and generates a sound signal indicative of the collected sound.
[0014] The encoding unit 13 changes the characteristics of the sound in the analog domain before it is collected by the sound collection unit 14. Here, the sound whose characteristics have been changed is different from the original sound, so it can be said that the encoding unit 13 encodes and encrypts the sound.
[0015] The sound collection unit 14 changes the characteristics of the sound before it is collected in the analog domain, and collects the sound with the changed characteristics to generate a sound signal.
[0016] In this embodiment, sound characteristics are changed by both the encoding unit 13 and the sound collection unit 14. However, the parts that change the sound characteristics may be collectively formed into the encoding unit 13, and the sound collection unit 14 may only have the function of collecting the sound whose characteristics have been changed and generating a sound signal.
[0017] The AD conversion unit 12 is, for example, an AD conversion circuit, and performs analog-to-digital conversion of an analog sound signal into actual sound data 100, which is digital data. Sensor declaration information 101 is attached to the actual sound data 100. As will be described in detail later, the sensor declaration information 101 is information for identifying the sound receiving unit 11, and indicates, for example, a unique value (number) assigned to each sound receiving unit 11. The sensor declaration information 101 is attached to the actual sound data 100 by, for example, a user. Therefore, the sensor declaration information 101 can also be said to be information declared by the user.
[0018] The actual sound data 100 and sensor declaration information 101 generated by the recording device 2 are output to the authenticity discrimination device 3, for example, via a network or a recording medium. The authenticity discrimination device 3 includes a restoration unit 21 and a discrimination unit 22. The restoration unit 21 identifies the number of the sound receiving unit 11 used when the sound was picked up by the recording device 2, based on the sensor declaration information 101. The restoration unit 21 then restores the original sound by decoding the actual sound data 100, for example, based on parameters that change the characteristics of the sound of the sound receiving unit 11 corresponding to the identified number. The discrimination unit 22 discriminates the authenticity of the sound restored by the restoration unit 21.
[0019] Note that, while the following mainly describes the case where real sound data 100 is input to the authenticity discriminator 3, it is also possible that only real sound data 100 is input to the authenticity discriminator 3, and that artificially generated generated data may also be input. In such cases, the restoration unit 21 cannot restore the generated data, and it is therefore possible to determine that the data is not real sound data 100.
[0020] The sound receiving unit 11, the AD converting unit 12, the encoding unit 13, the sound collecting unit 14, the restoring unit 21 and the determining unit 22 will be described in detail below.
[0021] [1.2 Specific Configuration of Recording Device] Figure 3 is a diagram showing the structure of the sound receiving unit 11. Figure 3A is an overall perspective view of the sound receiving unit 11. Figure 3B is an exploded perspective view of the sound receiving unit 11. As shown in Figures 3A and 3B, the sound receiving unit 11 functions as an encoding unit 13 and a sound collection unit 14 by means of physical structures, and these are configured as a single assembly.
[0022] The encoding unit 13 is composed of a case 31 and a cover 32. The case 31 is formed, for example, in the shape of a hollow, approximately rectangular parallelepiped with one side open. The sound pickup unit 14 is housed inside the case 31. The open side of the case 31 is covered with the cover 32. Therefore, the case 31 and the cover 32 function as a housing frame that houses the sound pickup unit 14.
[0023] The cover 32 is formed in a substantially rectangular shape that is substantially the same shape as the open surface of the case 31. A plurality of mask holes 33 are formed in the cover 32, penetrating the cover 32 in the plate thickness direction. In the example of Figures 3A and 3B, a total of 16 mask holes 33 are provided in four rows and four columns. In the example of Figures 3A and 3B, the mask holes 33 are formed in various different shapes, such as circles, squares, triangles, and hexagons. However, one or more mask holes 33 may be formed in the same shape.
[0024] The sound collection unit 14 is composed of a microphone 41 and a sound hole 42. In the sound collection unit 14, sound that passes through the sound hole 42 is collected by the microphone 41. A plurality of microphones 41 and a plurality of sound holes 42 are provided in pairs. In the example of FIGS. 3A and 3B, a total of 16 microphones 41 and sound holes 42 are provided in four rows and four columns. Also, in the example of FIGS. 3A and 3B, the sound collection units 14 are formed in different shapes. However, one or more sound collection units 14 may be formed in the same shape.
[0025] 4 is a diagram showing the structure of the sound pickup unit 14. As shown in FIG. 4, the microphone 41 includes a diaphragm 51, a voice coil 52, and a magnet 53.
[0026] The diaphragm 51 is a circular elastic thin film that vibrates in response to input sound. One end of the voice coil 52 is connected to the diaphragm 51, and it expands and contracts in response to the vibration of the diaphragm 51. A magnet 53 is disposed on the opposite side of the voice coil 52 from the diaphragm 51. When the voice coil 52 expands and contracts in response to the vibration of the diaphragm 51, electromagnetic induction occurs. The sound pickup unit 14 acquires the current that flows due to this electromagnetic induction as a sound signal.
[0027] One end of the sound hole 42 is located in a position facing the diaphragm 51. The sound hole 42 is formed in a cylindrical shape, and sound is input from the other end that does not face the diaphragm 51. The sound that inputs to the sound hole 42 is guided to the diaphragm 51, causing it to vibrate. In this case, the sound hole 42 acts like a tuning fork, and can change the characteristics of the sound by changing the frequency or directionality of the sound that passes through.
[0028] In the above example, the number of mask holes 33 and the number of sound collection units 14 are the same, but they do not have to be the same. In the sound receiving unit 11, the mask holes 33 and the microphone 41 are spaced apart. In the sound receiving unit 11, sound that has passed through one mask hole 33 is not input to one sound collection unit 14, but rather sound that has passed through one mask hole 33 can be input to one or more sound collection units 14, and sounds that have passed through each of multiple mask holes 33 can be input to one sound collection unit 14.
[0029] Next, the parameters (structural parameters) that change the properties of the sound in the encoding unit 13 and the sound collection unit 14 will be described.
[0030] Fig. 5 is a diagram showing masking by the mask hole 33. Fig. 6 is a diagram showing the structural parameters of the encoding unit 13 and the sound characteristics that change depending on the structural parameters. As shown in Fig. 5, sound input to the sound receiving unit 11 (input sound) enters the inside of the case 31 mainly through the mask hole 33 formed in the cover 32. The sound that enters the inside of the case 31 through the mask hole 33 (passing sound) is diffracted by the mask hole 33, and the sound characteristics are changed from the input sound.
[0031] Furthermore, some of the sound that enters the interior of the case 31 does not pass through the mask hole 33 but passes through, for example, the case 31 and the cover 32, and then enters the interior of the case 31. The sound that passes through the case 31 and the cover 32 and enters the interior of the case 31 also has its sound characteristics changed from the input sound in the process of passing through the case 31 and the cover 32.
[0032] 6, possible structural parameters that change the sound characteristics in the encoding unit 13 include the shape and size of the mask holes 33, the spacing between adjacent mask holes 33, structural changes in depth, the number, etc. Note that structural changes in depth refer to how the shape, size, etc. of the holes change in the thickness direction of the cover 32.
[0033] Possible structural parameters that change the sound characteristics in the encoding unit 13 include the shape, material, thickness, and margin size of the case 31 and cover 32. The margin size refers to the outer area of the cover 32 where the mask hole 33 is not formed.
[0034] As shown in FIG. 6 , a sound signal can be represented by three sound signal parameters: signal strength, phase difference, and frequency. The phase difference includes a temporal phase difference and a spatial phase difference. The gain and transmittance characteristics related to signal strength can be changed by changing at least one of the structural parameters of the encoding unit 13. The delay characteristics related to the temporal phase difference can be changed by changing at least one of the structural parameters of the encoding unit 13. The directivity characteristics related to the spatial phase difference can be changed by changing at least one of the structural parameters of the encoding unit 13. The filter and harmonics related to frequency can be changed by changing at least one of the structural parameters of the encoding unit 13. The term "filter" refers to attenuating (blocking) or amplifying a specific frequency band. The term "harmonics" refers to attenuating (blocking) or amplifying a specific harmonic component.
[0035] In this way, in the encoding unit 13, it is possible to change the characteristics of the sound that enters the inside of the case 31 by changing the structural parameters of the case 31, the cover 32 and the mask hole 33.
[0036] Fig. 7 is a diagram showing the function of the sound pickup unit 14. Fig. 8 is a diagram showing the structural parameters of the sound pickup unit 14 and the sound characteristics that change depending on the structural parameters. As shown in Fig. 7, sound (passing sound) that enters the inside of the case 31 passes through the sound hole 42 and reaches the diaphragm 51 of the microphone 41, vibrating the diaphragm 51 and being picked up as a recorded sound. At this time, the sound characteristics are changed depending on the structural parameters of the sound hole 42 and the sound pickup unit 14.
[0037] 8, possible structural parameters that change the sound characteristics of the sound pickup unit 14 include the material, shape, size, spacing between adjacent sound holes 42, thickness, hole shape, etc. The "thickness" refers to the length of the sound hole 42 in the direction in which sound passes through.
[0038] Possible structural parameters that change the sound characteristics in the sound pickup unit 14 include the material, thickness, surface size, and distance from the sound hole 42 of the diaphragm 51. Further, possible structural parameters that change the sound characteristics in the sound pickup unit 14 include the spacing between adjacent microphones 41 and the three-dimensional position inside the case 31. Note that structural parameters that change the sound characteristics in the sound pickup unit 14 may also include the magnetic force of the magnet 53 and the elastic force of the voice coil 52.
[0039] Here, it is possible to consider making the sound signal multidimensional as a change in sound characteristics, as shown in Fig. 8. Making the sound signal multidimensional means obtaining multiple sound signals by collecting one input sound with multiple microphones 41.
[0040] In the sound collection unit 14, it is possible to change the characteristics (gain, transmittance, directivity, delay, filter, harmonics) of the sound collected by the microphone 41 by changing at least one of the structural parameters of the microphone 41 and the sound hole 42. However, the sound collection unit 14 mainly focuses on making the sound signal multidimensional, and other characteristics are changed incidentally.
[0041] In this way, the sound receiving unit 11 can change various sound characteristics mainly by changing the structural parameters of the encoding unit 13. In other words, the encoding unit 13 can conceal sound (information) by modulating the input sound.
[0042] Furthermore, in the sound receiving unit 11, the sound collection unit 14 has multiple microphones 41, so that it is possible to generate multiple sound signals representing sounds with changed characteristics. In other words, the sound collection unit 14 can make the sound signal multidimensional and increase the amount of information (capacity).
[0043] The sound signals generated by the sound collection unit 14 are output to the AD conversion unit 12, and are converted into digital data by the AD conversion unit 12. Therefore, the AD conversion unit 12 receives as many sound signals as the number of microphones 41, and converts each sound signal into digital data.
[0044] The digital data converted by the AD conversion unit 12 is output as actual sound data 100 to the authenticity discrimination device 3 via a network or a recording medium. At this time, since the changes in sound characteristics differ depending on the structural parameters of the sound receiving unit 11 (encoding unit 13 and sound collection unit 14) as described above, sensor declaration information 101 for individually identifying each sound receiving unit 11 is added to the actual sound data 100.
[0045] The sensor declaration information 101 may be added as meta information of the actual sound data 100, or may be added by being embedded in the data structure of the actual sound data 100. The sensor declaration information 101 may also be added separately from the actual sound data 100. In this case, information linking the actual sound data 100 and the sensor declaration information 101 may be added to one or both of the actual sound data 100 and the sensor declaration information 101, or may be output to the authenticity discrimination device 3 separately from the actual sound data 100 and the sensor declaration information 101.
[0046] 1.3 Specific Configuration of the Authentication Discrimination Device 3 Fig. 9 is a diagram illustrating the processing of the restoration unit 21. As shown in Fig. 9, the input sound is denoted by X, the encoding matrix of the sound from the sound receiving unit 11 is denoted by Z, and the output matrix of the actual sound data 100 is denoted by Y. Furthermore, it is assumed that the encoding matrix Z is expressed as the product of a mask characteristic matrix A determined by the structural parameters of the encoding unit 13 and a microphone characteristic matrix B determined by the structural parameters of the sound collection unit 14.
[0047] Here, the output matrix Y is expressed as in equation (1). n is the number of microphones 41 (sound signals), and k is the number of frequency divisions. For example, if there are 16 microphones 41, the maximum value of n is 16. n=1 corresponds to the first microphone 41, n=2 corresponds to the second microphone 41, and n=3 to 16 correspond to the third to sixteenth microphones 41, respectively. Furthermore, if frequencies up to 20,000 Hz are divided into 100 (divided into 200 Hz increments), the maximum value of k is 100. k=1 corresponds to 0 Hz, k=2 corresponds to 200 Hz, and k=3 to 100 correspond to frequencies obtained by sequentially adding 200 Hz from 600 Hz. Therefore, the element ynk is the spectrum value of k×200 "Hz" of the sound picked up by the nth microphone 41.
[0048] The mask characteristic matrix A is a matrix that represents how the sound characteristics are changed by the structural parameters of the encoding unit 13. The microphone characteristic matrix B is a matrix that represents how the sound characteristics are changed by the structural parameters of the sound collection unit 14. The mask characteristic matrix A is expressed as in equation (2), and the microphone characteristic matrix B is expressed as in equation (3). The encoding matrix Z is expressed as the product of the mask characteristic matrix A and the microphone characteristic matrix B. Note that n and k in the mask characteristic matrix A and microphone characteristic matrix B are the number of microphones 41 (sound signals) and the number of frequency divisions, as in equation (1).
[0049] Then, the element a of the mask characteristic matrix A nk indicates how the characteristics of the n-th frequency component in the k-th microphone 41 are changed. nk indicates the n-th frequency characteristic of the k-th microphone 41.
[0050] The mask characteristic matrix A and microphone characteristic matrix B represent the structural parameters of the encoding unit 13 and the sound collection unit 14 in a model manner, and do not actually calculate these values.
[0051] In this way, since the encoding matrix Z is a matrix representation of elements that change the characteristics of the sound depending on the structural parameters of the encoding unit 13 and the sound collection unit 14, the output matrix Y is expressed as the product of the encoding matrix Z and the input sound X, as shown in equation (4).
[0052] Also, the restoration matrix for restoring the output matrix Y, in which the sound characteristics have been changed, to the original sound is Z -1 Let X′ be the sound restored by the restoration unit 21 (hereinafter referred to as restored sound), then the restored sound X′ is expressed by equation (5).
[0053] Here, equation (5) is a matrix calculation formula, but if we try to calculate the restored sound X' using the matrix calculation formula as it is, it may become difficult to create a learning algorithm in the machine learning described below.
[0054] Therefore, in this embodiment, the output matrix Y is visualized as a two-dimensional map and machine learning for image analysis is used to simplify the processing. Also, since the input sound is collected over a relatively long period of time, two-dimensional maps are created from the output matrix Y for each sound signal divided into predetermined time intervals to improve the accuracy when calculating the restored sound X'.
[0055] 10 is a diagram showing a two-dimensional map 110. When the actual sound data 100 is input, the restoration unit 21 divides the sound signals for each of the multiple microphones 41 shown in the actual sound data 100 into predetermined time intervals (e.g., one second). The restoration unit 21 then performs a fast Fourier transform on each of the divided sound signals for each of the multiple microphones 41 to calculate a spectral value for each predetermined frequency. In this way, the restoration unit 21 calculates an output matrix Y for each predetermined time interval.
[0056] 10, the restoration unit 21 generates a two-dimensional map 110 at predetermined intervals (t=0, 1, 2, 3, ...) expressed by color shading or brightness values according to the value of each element in the output matrix Y. In other words, the two-dimensional map 110 is an image of the output matrix Y expressed by color shading or brightness values.
[0057] Here, the restoration matrix Z -1A learning model corresponding to (hereinafter referred to as a restoration model) is obtained by machine learning a combination of a correct input sound X and a two-dimensional map 110 created based on an output matrix Y indicating a sound signal obtained when the input sound X is picked up by the sound receiving unit 11. Therefore, the restoration model can be said to be a model determined by structural parameters that change the characteristics of sound from the sound receiving unit 11. Furthermore, since the restoration model is unique to each sound receiving unit 11, it is obtained for each sound receiving unit 11 (for each number of the sound receiving unit 11). The machine learning performed here can be, for example, deep learning (DL). As an algorithm used in deep learning, a known method such as a convolutional neural network (CNN) or a recurrent neural network (RNN) can be used.
[0058] The restoration unit 21 generates a restored sound X' based on the generated two-dimensional map 110 and a restoration model of the sound receiving unit 11 corresponding to the number indicated in the sensor declaration information 101. However, since the sensor declaration information 101 is not always correct, there are cases where the restoration unit 21 generates a restored sound X' using a restoration model of the sound receiving unit 11 corresponding to a different number by referring to the similarity between the sound receiving units 11.
[0059] 11 is a diagram for explaining specific processes of the restoration unit 21 and the determination unit 22. As shown in FIG. 11, the restoration unit 21 functions as a sensor identification unit 61 and a label generation unit 62.
[0060] The sensor identification unit 61 generates a restored sound X' based on the generated two-dimensional map 110 and a decoding model of the sound receiving unit 11 corresponding to the number indicated in the sensor declaration information 101. The sensor identification unit 61 then determines whether the generated restored sound X' is a restored version of the input sound X. Specifically, the sensor identification unit 61 calculates the correlation between the restored sounds X' for each microphone 41, and if the correlation coefficient is equal to or greater than a predetermined threshold, determines that the restored sound X' is a correct restoration of the input sound X. The sensor identification unit 61 also determines that the sensor declaration information 101 indicated the correct number. In this case, the label generation unit 62 outputs to the discrimination unit 22 a declaration result indicating that the identification of the sound receiving unit 11 that changed the sound characteristics has been completed (identification complete) and that the identification was as declared, and an identification result indicating the number of the sound receiving unit 11 at the time of the determination.
[0061] On the other hand, if the correlation coefficient is less than a predetermined threshold, the sensor identification unit 61 determines that the restored sound X' is not a correct restoration of the input sound X. In other words, the sensor identification unit 61 determines that the input actual sound data 100 was not collected by the sound receiving unit 11 of the number indicated in the sensor declaration information 101. In this case, the sensor identification unit 61 acquires restored models of the sound receiving units 11 of other numbers that are highly similar to the sound receiving unit 11 of the number indicated in the sensor declaration information 101, in order of similarity.
[0062] Here, the similarity between different sound receiving units 11 is known in advance. Therefore, the sensor identification unit 61 can obtain the similarity of sound receiving units 11 with other numbers based on the number indicated in the sensor declaration information 101.
[0063] The sensor identification unit 61 regenerates the restored sound X' based on the restored models acquired in descending order of similarity and the two-dimensional map 110. The sensor identification unit 61 also determines whether the generated restored sound X' is a restored version of the input sound X. The sensor identification unit 61 then repeats this process until it determines that the restored sound X' is a restored version of the input sound X.
[0064] Assume that the sensor identification unit 61 determines that the restored sound X' generated based on the restoration model of the sound receiving unit 11 corresponding to a number different from the number indicated in the sensor declaration information 101 is the restored sound of the input sound X. In this case, the label generation unit 62 outputs to the discrimination unit 22 a declaration result indicating that the identification of the sound receiving unit 11 that has changed the characteristics of the sound has been completed (identification completed) and that the result is different from the declaration, and an identification result indicating the number of the sound receiving unit 11 at the time of the determination.
[0065] On the other hand, if it is not determined that the restored sound X' is a restoration of the input sound X even after using a predetermined number of restoration models in descending order of similarity, the label generation unit 62 outputs information to the discrimination unit 22 indicating that the identification of the sound receiving unit 11 that has changed the characteristics of the sound could not be completed (identification is not possible) and that identification is not possible.
[0066] When the information indicating the completion of the identification is input from the restoration unit 21, the determination unit 22 determines whether the restored sound X′ is the actual sound data 100.
[0067] A possible method for discriminating the restored sound X' is to use a machine-learned AI model. For example, the "Discriminator" part of the GAN can be used. The "Discriminator" of the GAN performs machine learning using the input sound X and the correct restored sound X', making it possible to derive the accuracy probability (%) of the restored sound X', which is actual sound data.
[0068] Using this type of discrimination method, the discrimination unit 22 derives the accuracy probability (%) of the restored sound X' as the discrimination result.
[0069] Thereafter, if the discrimination unit 22 has input thereto a declaration result indicating that the discrimination has been completed and that the result is as declared, and an identification result indicating the number of the sound receiving unit 11 at the time of the judgment, the discrimination unit 22 outputs the discrimination result, the discrimination result, and the declaration result indicating that the result is as declared. A display unit such as a display may be used as a destination for these outputs. For example, as shown in FIG. 12 , the discrimination unit 22 displays the discrimination result (the number of the sound receiving unit 11 (e.g., No. 0.10)), the discrimination result (correctness probability = 80%), and the fact that the declaration is as declared on the display unit 77 (see FIG. 15 ). This allows the user to understand that the restored sound X' is the restored input sound picked up by the sound receiving unit 11 declared.
[0070] In addition, when the discrimination unit 22 receives input indicating that the identification of the sound receiving unit 11 that has changed the sound characteristics has been completed (identification completed), a declaration result indicating that the result is different from the declaration, and an identification result indicating the number of the sound receiving unit 11 at the time of identification, the discrimination unit 22 discriminates the authenticity of the restored sound X' as described above.
[0071] Then, the discrimination unit 22 outputs the identification result, the discrimination result, and a message indicating that the result is different from the declaration. For example, the discrimination unit 22 displays the identification result (the number of the sound receiving unit 11), the discrimination result (correct answer probability = 80%), and a message indicating that the result is different from the declaration on the display unit 77. This allows the user to understand that the restored sound X' is a restored version of input sound that was picked up by a sound receiving unit 11 different from the sound receiving unit 11 that made the declaration.
[0072] Furthermore, when the discrimination unit 22 receives information indicating that the discrimination of the sound receiving unit 11 that has changed the sound characteristics could not be completed (indistinguishable) and that discrimination is not possible, the discrimination unit 22 outputs information indicating that discrimination is not possible and that discrimination is not possible without discriminating the authenticity of the restored sound X'. For example, the discrimination unit 22 displays the discrimination result (indistinguishable) or the discrimination result (indistinguishable) on the display unit. This allows the user to understand that the restored sound X' was not picked up by the wrong sound receiving unit 11.
[0073] In this way, by outputting the results in three patterns, it is possible to easily understand not only the identification results and discrimination results, but also whether discrimination was even possible in the first place.
[0074] [1.4 Processing Procedure] Next, we will explain the processing procedures of the above-mentioned sound recording device 2 and authenticity discriminator 3. Fig. 13 is a flowchart showing the processing procedures of the sound recording device 2. As shown in Fig. 13, when processing starts in the sound recording device 2, in step S1, the sound receiving unit 11 picks up input sound and generates a sound signal.
[0075] In the next step S2, the AD conversion unit 12 converts the sound signal into digital data, that is, actual sound data 100. Then, in step S3, the AD conversion unit 12 adds sensor declaration information 101 to the actual sound data 100 and outputs it to the authenticity discrimination device 3.
[0076] 14 is a flowchart showing the processing procedure of the authenticity discriminator 3. As shown in Fig. 14, when the authenticity discriminator 3 starts processing, in step S11 the restoration unit 21 acquires actual sound data 100 and sensor declaration information 101. In step S12, the restoration unit 21 generates a time-divided two-dimensional map 110 based on the actual sound data 100.
[0077] In step S13, the restoration unit 21 selects one sound receiving unit 11 in descending order of similarity to the sound receiving unit 11 having the number indicated in the sensor declaration information 101. Here, the sound receiving unit 11 having the number indicated in the sensor declaration information 101 is selected first. In step S14, the restoration unit 21 generates a restored sound X' based on all the time-divided two-dimensional maps 110 and the restoration model of the selected sound receiving unit 11.
[0078] In step S15, the restoration unit 21 determines whether the restored sound X' is a restoration of the input sound X (whether the restored sound X' has been matched). If the restored sound X' has not been matched (No in step S15), the process returns to step S13. On the other hand, if the restored sound X' has been matched (Yes in step S15), the discrimination unit 22 determines the authenticity of the restored sound X' in step S16. Then, in step S16, the discrimination unit 22 outputs the identification result, the discrimination result, and the declaration result. Note that the restoration unit 21 may output the identification result and the declaration result to a display unit.
[0079] 1.5 Hardware Configuration Fig. 15 is a block diagram showing an example of the hardware configuration of an information processing device 70 corresponding to the authenticity discriminator 3. As shown in Fig. 15, the information processing device 70 includes a CPU 71, a ROM 72, and a RAM 73. The CPU 71 functions as an arithmetic processing unit that performs various processes, and executes the various processes in accordance with a program stored in the ROM 72 or a program loaded from the storage unit 79 to the RAM 73. The RAM 73 also stores data and the like required for the CPU 71 to execute the various processes, as appropriate.
[0080] The CPU 71, ROM 72, and RAM 73 are interconnected via a bus 74. An input / output interface (I / F) 75 is also connected to this bus 74.
[0081] An input unit 76 consisting of operators and operation devices is connected to the input / output interface 75. For example, the input unit 76 may be various operators and operation devices such as a keyboard, a mouse, keys, a dial, a touch panel, a touch pad, a remote controller, etc. The input unit 76 detects user operations, and the CPU 71 interprets signals corresponding to the input operations.
[0082] A display unit 77, such as an LCD (Liquid Crystal Display) or an organic EL (Electro-Luminescence) panel, and an audio output unit 78, such as a speaker, are connected integrally or separately to the input / output interface 75. The display unit 77 is used to display various types of information, and is configured as a display device provided in the housing of the information processing device 70, for example.
[0083] The display unit 77 displays images for various image processing, moving images to be processed, etc. on the display screen based on instructions from the CPU 71. The display unit 77 also displays various operation menus, icons, messages, etc., i.e., a GUI (Graphical User Interface), based on instructions from the CPU 71.
[0084] A storage unit 79 and a communication unit 80 can be connected to the input / output interface 75. The storage unit 79 is configured by a hard disk drive (HDD) or a solid state drive (SSD), and stores various types of information.
[0085] The communication unit 80 performs communication processing via a transmission path such as the Internet, and communication with various devices via wired / wireless communication, bus communication, etc. In particular, in the case of this embodiment, the communication unit 80 is capable of performing data communication with the recording device 2 described above.
[0086] A drive 81 is also connected to the input / output interface 75 as required, and a removable recording medium 82 such as a memory card or optical disk is appropriately attached thereto.
[0087] The drive 81 makes it possible to read data files such as programs used in various processes from a removable recording medium 82. The read data files are stored in a storage unit 79, and images and sounds contained in the data files are output on a display unit 77 and an audio output unit 78. Furthermore, the computer programs and the like read from the removable recording medium 82 are installed in the storage unit 79 as necessary.
[0088] Here, the information processing device 70 is not limited to being configured as a single computer device as shown in Fig. 15, but may be configured as a system of multiple computer devices. The multiple computer devices may be systemized using a LAN (Local Area Network) or the like, or may be located in a remote location using a VPN (Virtual Private Network) or the like using the Internet or the like. The multiple computer devices may include computer devices serving as a server group (cloud) available through a cloud computing service.
[0089] 2. Modifications Note that the embodiment is not limited to the specific example described above, and various modifications may be made.
[0090] For example, in the above-described authentication system 1, the recording device 2 and the authentication device 3 are provided as separate (separate) units, but they may also be provided as an integrated unit.
[0091] In the above embodiment, it is determined whether the input sound X is restored by calculating the correlation of the restored sound X' for each microphone 41. However, the determination method is not limited to this, and other methods may be used.
[0092] In the above embodiment, the discriminator 22 discriminates authenticity using the discriminator in the GAN. However, other methods may be used for discriminating authenticity.
[0093] In the above embodiment, the restored sound X' is restored based on the two-dimensional map 110 generated by the output matrix Y and the decoding model. However, the restored sound X' may be restored based on the output matrix Y and the decoding model without using the two-dimensional map 110.
[0094] In the above embodiment, the mask holes 33 and the sound collection units 14 are arranged regularly, but the mask holes 33 do not have to be arranged regularly (they may be arranged randomly), as shown in Fig. 16. Similarly, the sound collection units 14 do not have to be arranged regularly.
[0095] 3. Summary of the Embodiment As described above, the recording device 2 includes a sound receiving unit 11 that changes the characteristics of the sound in the analog domain when collecting the sound, then collects the sound, and generates a sound signal representing the collected sound. As a result, the sound signal representing the sound whose characteristics have been changed by the sound receiving unit 11 is converted into real sound data 100, which is digital data. The authenticity discrimination device 3 that receives the real sound data 100 restores the real sound data 100 using a restoration model of the sound receiving unit 11 that collected the sound signal. Here, the decoded model is a model determined by structural parameters that change the characteristics of the sound in the sound receiving unit 11. Therefore, when restoring the real sound data 100, the restoration is based on the structural parameters that change the characteristics of the sound in the sound receiving unit 11. Furthermore, because the structural parameters of the sound receiving unit 11 are difficult to imitate, it is also possible to make it difficult to imitate the real sound data 100 (sound signal).
[0096] The sound receiving unit 11 includes a sound collection unit 14 having a plurality of microphones 41 that collect sounds with altered characteristics. This makes it possible to collect sounds with altered characteristics using the plurality of microphones 41, thereby increasing the amount of information (capacity). Furthermore, by setting the plurality of microphones 41 to the same structural parameters, redundancy can be achieved, thereby improving the reliability of authenticity determination.
[0097] The system includes an encoding unit 13 having multiple structural parameters that change the characteristics of sound before it is picked up. As shown in FIG. 6, the encoding unit 13 has multiple structural parameters, such as the shape and size of the mask hole 33 and the shapes and materials of the case 31 and cover 32. This allows the multiple structural parameters to act in a complex manner to change the characteristics of the sound, making it difficult to imitate the encoding (change in characteristics). Furthermore, because the sound characteristics are changed in the analog domain, it is possible to make it difficult to tamper with data in the digital domain.
[0098] The system includes a converter (AD converter 12) that converts the sound signal into digital data. By converting the sound signal, whose sound characteristics have been changed in the analog domain, into digital actual sound data 100, the encoding (change in characteristics) can be made difficult to imitate.
[0099] The encoding unit 13 is formed with a plurality of mask holes 33 that allow sound to pass through. By providing a plurality of mask holes 33 that allow sound to pass through in this way, sounds that have been changed to have different characteristics in the respective mask holes 33 can be picked up by the microphone 41. This makes it more difficult to imitate the encoding (change in characteristics).
[0100] The encoding unit 13 changes the characteristics of the sound by changing at least one of the shape, size, and depth of the mask holes 33, and the spacing between the mask holes. In this way, by changing the structural parameters of the mask holes 33, the characteristics of the sound can be easily changed.
[0101] The encoding unit 13 includes a housing frame (case 31, cover 32) that houses the microphone 41, and changes the sound characteristics by changing at least one of the shape, material, thickness, and margin size of the housing frame. In this way, the sound characteristics can be easily changed by changing the structural parameters of the microphone 41.
[0102] The sound collection unit 14 is provided with a cylindrical sound hole 42 for passing sound, which is provided for each of the plurality of microphones 41. This allows the sound holes 42 to change the characteristics of the sound, such as changing the frequency or directionality of the passing sound.
[0103] The sound pickup unit 14 changes the sound characteristics by changing at least one of the material, shape, size, thickness, and spacing between the sound holes 42. In this way, by changing the structural parameters of the sound holes 42, the sound characteristics can be easily changed.
[0104] The microphone 41 includes a diaphragm 51, a voice coil 52, and a magnet 53, and the sound pickup unit 14 changes the sound characteristics by changing at least one of the material, thickness, surface size, and distance from the sound hole 42 of the diaphragm 51. In this way, by changing the structural parameters of the diaphragm 51, the sound characteristics can be easily changed.
[0105] The system includes a restoration unit 21 that generates restored sound X' by restoring digital data (actual sound data 100) based on the structural parameters of the sound receiving unit 11. This allows the restoration sound X' to be generated only in the authenticity discrimination device 3 for which the restoration model based on the structural parameters of the sound receiving unit 11 is known, thereby protecting the confidentiality of the actual sound data.
[0106] The system includes a discrimination unit 22 that discriminates the authenticity of the restored sound X' restored by the restoration unit 21. Since the restoration unit 21 generates the restored sound X' using a decoding model, it is easy to determine the authenticity.
[0107] The authenticity discrimination system 1 also includes a sound receiving unit 11 that changes the characteristics of the sound in the analog domain when collecting the sound and then collects the sound to generate a sound signal indicative of the collected sound, a restoration unit 21 that generates a restored sound X' by restoring the sound signal generated by the sound receiving unit 11 based on parameters that change the characteristics of the sound generated by the receiving unit, and a discrimination unit 22 that determines the authenticity of the restored sound restored by the restoration unit 21. With this authenticity discrimination system 1, it is possible to obtain the same functions and effects as those of the above-mentioned embodiment.
[0108] The authenticity discrimination device 3 includes a restoration unit 21 that generates a restored sound X' by restoring a sound signal generated by collecting sound whose characteristics have been changed by the sound receiving unit 11 in the analog domain when collecting the sound, and a discrimination unit 22 that determines the authenticity of the restored sound X' generated by the restoration unit 21. The authenticity discrimination device 3 restores the sound signal based on the restoration model of the sound receiving unit 11. Therefore, it is possible to generate the restored sound X' only when the restoration model of the sound receiving unit 11 is known, thereby enabling the sound signal to be concealed (encrypted). Furthermore, because parameters that change the characteristics of the sound in the sound receiving unit 11, i.e., the restoration model, are used, it is possible to easily identify counterfeit data and improve the accuracy of authenticity discrimination.
[0109] The sound signal is provided with sensor declaration information 101 for identifying the sound receiving unit 11 that picked up the sound. This allows the authenticity discrimination device 3 to easily identify the sound receiving unit 11 that picked up the sound, and can significantly reduce the time required to identify the restoration model (parameters of the sound receiving unit 11) to be used in the restoration process by the restoration unit 21.
[0110] The restoration unit 21 generates restored sound based on the parameters (restoration model) of the sound receiving unit 11 indicated in the sensor declaration information 101. This allows the authenticity discrimination device 3 to easily generate restored sound based on the restoration model of the sound receiving unit 11 indicated in the sensor declaration information 101. Furthermore, if there is no fraud, the sound characteristics have been changed by the sound receiving unit 11 indicated in the sensor declaration information 101, so there is a high possibility that the restored sound reproduces the sound before it was picked up, and the accuracy of authenticity discrimination can also be improved.
[0111] The restoration unit 21 determines whether the generated restored sound X' is a restored version of the sound before it was picked up. This allows for excluding restored sounds X' that will not be restored to the original sound even if a sound signal representing a sound picked up by a different sound receiving unit 11 is restored.
[0112] If the discrimination unit 22 determines that the generated restored sound X' is a restored version of the sound before it was picked up, it discriminates the authenticity of the restored sound. As a result, the discrimination unit 22 discriminates the authenticity of only the restored sound X' that has been correctly restored, thereby eliminating unnecessary authenticity discrimination and reducing the processing load.
[0113] The similarity of parameters that change the characteristics of sound is determined for multiple sound receiving units 11, and the restoration unit 21 generates restored sound based on parameters of other sound receiving units 11 that have a high similarity to the sound receiving unit 11 indicated in the sensor declaration information 101. This makes it possible to generate restored sound X' using optimal parameters (restoration model), thereby improving the restoration accuracy of restored sound X'.
[0114] The restoration unit 21 generates restored sound X' based on parameters of the sound receiving unit 11 in descending order of similarity to the sound receiving unit 11 indicated in the sensor declaration information 101, and determines whether the generated restored sound X' is a restoration of the sound before it was picked up. As a result, the restored sound X' is generated based on parameters (restoration model) of the sound receiving unit 11 that are highly similar to the sound receiving unit 11 indicated in the sensor declaration information 101, and therefore there is a high possibility that it is a restoration of the sound before it was picked up. Therefore, the number of times the restored sound X' is generated can be reduced, and the processing speed can be improved.
[0115] When the restoration unit 21 determines that the restored sound X' generated based on the parameters of the sound receiving unit 11 indicated in the sensor declaration information 101 is a restoration of the sound before it was picked up, it outputs an identification result indicating the sound receiving unit 11 indicated in the sensor declaration information 101 and a declaration result indicating that the declaration is correct. This makes it possible to notify the user that the declaration is correct and the number of the sound receiving unit 11 as the identification result.
[0116] When the restoration unit 21 determines that the restored sound X' generated based on parameters of another sound receiving unit 11 that is highly similar to the sound receiving unit 11 indicated in the sensor declaration information 101 is a restoration of the sound before it was picked up, it outputs an identification result indicating the sound receiving unit 11 that has the parameters from which the restored sound X' was generated, and a declaration result indicating that the declaration is different. This makes it possible to notify the user that the restored sound X' was generated using parameters (restoration model) of a sound receiving unit 11 that is different from the sound receiving unit 11 indicated in the sensor declaration information 101, and to notify the user which sound receiving unit 11 it was.
[0117] The restoration unit 21 generates restored sound X' based on parameters of a predetermined number of sound receiving units 11 in descending order of similarity to the sound receiving units 11 indicated in the sensor declaration information 101, and if it determines that the generated restored sound X' is not a restoration of the sound before it was picked up, it outputs information indicating that it is indistinguishable. This makes it possible to notify the user that restored sound X' that restored the sound before it was picked up was not generated. It can also be notified that no authenticity determination has been performed thereafter.
[0118] The discrimination unit 22 outputs the result of the authenticity determination, thereby notifying the user of the result of the authenticity determination (probability of correctness).
[0119] The display unit displays the identification result indicating the sound receiving unit 11 having the parameters from which the restored sound was generated, the declaration result indicating whether the sound is as declared or different from the declaration, and the authenticity result of the authenticity determination. This allows the user to easily check the identification result, declaration result, and authenticity result.
[0120] The restoration unit 21 generates a two-dimensional map 110 based on the sound signal (actual sound data 100), and generates restored sound based on the generated two-dimensional map 110. This makes it possible to use existing image analysis algorithms when performing machine learning on the restored sound, thereby improving the learning speed and reducing the effort required to generate new algorithms.
[0121] The restoration unit 21 divides the sound signal in time series and generates a two-dimensional map for each divided sound signal. This makes it possible to generate the restored sound X' using multiple two-dimensional maps 110, thereby improving the generation accuracy of the restored sound X'.
[0122] The restoration unit 21 divides the sound signals generated by the plurality of sound receiving units 11 in time series, and generates a two-dimensional map 110 for each divided sound signal. This makes it possible to generate restored sound X' using an even greater number of two-dimensional maps 110, thereby improving the accuracy of generating restored sound X'.
[0123] The authenticity determination method generates a restored sound X' by restoring the sound signal generated by collecting sound whose characteristics have been changed by the sound receiving unit 11 in the analog domain when collecting the sound, based on a parameter that changes the characteristics of the sound by the sound receiving unit 11, and determines the authenticity of the generated restored sound X'. The program also causes a computer to execute a process of generating a restored sound X' by restoring the sound signal generated by collecting sound whose characteristics have been changed by the sound receiving unit 11 in the analog domain when collecting the sound, based on a parameter that changes the characteristics of the sound by the sound receiving unit 11, and determining the authenticity of the generated restored sound X'.
[0124] Such a program can be pre-recorded on a hard disk drive (HDD) or a ROM in a microcomputer having a CPU, which serves as a built-in recording medium in a computer or other device. Alternatively, the program can be temporarily or permanently stored (recorded) on a removable recording medium such as a flexible disk, a CD-ROM (Compact Disc Read Only Memory), an MO (Magneto Optical) disc, a DVD (Digital Versatile Disc), a Blu-ray Disc (registered trademark), a magnetic disk, a semiconductor memory, or a memory card. Such removable recording media can be provided as so-called packaged software. Furthermore, such a program can be installed on a personal computer or the like from a removable recording medium, or can be downloaded from a download site via a network such as a LAN (Local Area Network) or the Internet.
[0125] The effects described in this specification are merely examples and are not limiting, and other effects may also be present.
[0126] <4. The Present Technology> The present technology can also be configured as follows. (1) An authenticity discrimination device comprising: a restoration unit that generates restored sound by restoring a sound signal generated by collecting sound whose characteristics have been changed by a sound receiving unit in the analog domain when collecting sound; and a discrimination unit that discriminates the authenticity of the restored sound generated by the restoration unit. (2) The authenticity discrimination device described in (1), in which sensor declaration information for identifying the sound receiving unit that collected the sound is attached to the sound signal. (3) The authenticity discrimination device described in (2), in which the restoration unit generates the restored sound based on parameters of the sound receiving unit indicated in the sensor declaration information. (4) The authenticity discrimination device described in (3), in which the restoration unit determines whether the generated restored sound is a restoration of sound before it was collected. (5) The authenticity discrimination device described in (4), in which the discrimination unit discriminates the authenticity of the restored sound when it is determined that the generated restored sound is a restoration of sound before it was collected. (6) The authenticity discrimination device according to any one of (2) to (5), wherein similarities between parameters that change sound characteristics are determined for a plurality of the sound receiving units, and the restoration unit generates the restored sound based on parameters of the other sound receiving units that have a high similarity to the sound receiving unit indicated in the sensor declaration information. (7) The authenticity discrimination device according to (6), wherein the restoration unit generates the restored sound based on parameters of the sound receiving units in descending order of similarity to the sound receiving unit indicated in the sensor declaration information, and determines whether the generated restored sound is a restoration of the sound before it was picked up. (8) The authenticity discrimination device according to any one of (4) to (7), wherein, when the restoration unit determines that the restored sound generated based on the parameters of the sound receiving units indicated in the sensor declaration information is a restoration of the sound before it was picked up, it outputs an identification result indicating the sound receiving unit indicated in the sensor declaration information and a declaration result indicating that the declaration is correct.(9) The authenticity discrimination device according to (7) or (8), wherein, when the restoration unit determines that the restored sound generated based on parameters of another sound receiving unit having a high similarity to the sound receiving unit indicated in the sensor declaration information is a restored sound from before it was picked up, it outputs an identification result indicating the sound receiving unit having the parameters from which the restored sound was generated, and a declaration result indicating that the declaration is incorrect. (10) The authenticity discrimination device according to any of (7) to (9), wherein the restoration unit generates the restored sound based on parameters of a predetermined number of sound receiving units in descending order of similarity to the sound receiving unit indicated in the sensor declaration information, and when it determines that the generated restored sound is not a restored sound from before it was picked up, it outputs information indicating that it is indistinguishable. (11) The authenticity discrimination device according to any of (1) to (10), wherein the discrimination unit outputs an authenticity result. (12) The authenticity discrimination device according to (11), which displays on a display unit an identification result indicating the sound receiving unit having parameters from which the restored sound was generated, a declaration result indicating whether the sound is as declared or different from the declaration, and an authenticity discrimination result. (13) The authenticity determination device according to any of (1) to (12), wherein the restoration unit generates a two-dimensional map based on the sound signal and generates the restored sound based on the generated two-dimensional map. (14) The recording device according to (13), wherein the restoration unit divides the sound signal in time series and generates the two-dimensional map for each of the divided sound signals. (15) The authenticity discrimination device according to (14), wherein the restoration unit divides the sound signals generated by each of the plurality of sound receiving units in time series and generates the two-dimensional map for each of the divided sound signals. (16) A method for determining authenticity, which generates restored sound by restoring a sound signal generated by collecting sound whose characteristics have been changed by a sound receiving unit in the analog domain when collecting sound, and determines the authenticity of the generated restored sound. (17) A program that causes a computer to execute the following processes:
[0127] REFERENCE SIGNS LIST 1 Authentication discrimination system 2 Recording device 3 Authentication discrimination device 11 Sound receiving unit 12 AD conversion unit 13 Encoding unit 14 Sound collection unit 21 Restoration unit 22 Discrimination unit
Claims
1. An authenticity discrimination device comprising: a restoration unit that generates restored sound by restoring a sound signal generated by collecting sound whose characteristics have been changed by a sound receiving unit in the analog domain when collecting sound; and a discrimination unit that discriminates the authenticity of the restored sound generated by the restoration unit.
2. The authenticity determination device according to claim 1, wherein the sound signal is provided with sensor declaration information for identifying the sound receiving unit that picked up the sound.
3. The authenticity discrimination device according to claim 2, wherein the restoration unit generates the restored sound based on parameters of the sound receiving unit indicated in the sensor declaration information.
4. The authenticity determination device according to claim 3, wherein the restoration unit determines whether the generated restored sound is a restoration of the sound before it was picked up.
5. The authenticity discrimination device according to claim 4, wherein the discrimination unit discriminates the authenticity of the restored sound when it is determined that the generated restored sound is a restored version of the sound before it was picked up.
6. An authenticity discrimination device as described in claim 2, wherein the similarity of parameters that change the characteristics of the sound is determined for multiple sound receiving units, and the restoration unit generates the restored sound based on parameters of other sound receiving units that have a high similarity to the sound receiving unit indicated in the sensor declaration information.
7. The authenticity discrimination device according to claim 6, wherein the restoration unit generates the restored sound based on parameters of the sound receiving unit in descending order of similarity to the sound receiving unit indicated in the sensor declaration information, and determines whether the generated restored sound is a restoration of the sound before it was picked up.
8. The authenticity discrimination device according to claim 4, wherein, when the restoration unit determines that the restored sound generated based on the parameters of the sound receiving unit indicated in the sensor declaration information is a restoration of the sound before it was picked up, it outputs an identification result indicating the sound receiving unit indicated in the sensor declaration information and a declaration result indicating that the declaration is correct.
9. The authenticity discrimination device according to claim 7, wherein, when the restoration unit determines that the restored sound generated based on parameters of another sound receiving unit that is highly similar to the sound receiving unit indicated in the sensor declaration information is a restoration of the sound before it was picked up, it outputs an identification result indicating the sound receiving unit that has the parameters from which the restored sound was generated, and a declaration result indicating that the declaration is different.
10. The authenticity determination device according to claim 7, wherein the restoration unit generates the restored sound based on parameters of a predetermined number of the sound receiving units in descending order of similarity to the sound receiving units indicated in the sensor declaration information, and if it determines that the generated restored sound is not a restoration of the sound before it was picked up, it outputs information indicating that it is indistinguishable.
11. The authenticity determination device according to claim 1, wherein the determination unit outputs an authenticity determination result.
12. An authenticity determination device as claimed in claim 11, which displays on a display unit an identification result indicating the sound receiving unit having the parameters from which the restored sound was generated, a declaration result indicating whether the sound is as declared or different from the declaration, and an authenticity determination result.
13. The authenticity discrimination device according to claim 1, wherein the restoration unit generates a two-dimensional map based on the sound signal, and generates the restored sound based on the generated two-dimensional map.
14. The authenticity discrimination device according to claim 13, wherein the restoration unit divides the sound signal in time series and generates the two-dimensional map for each divided sound signal.
15. The authenticity discrimination device according to claim 14, wherein the restoration unit divides the sound signals generated by each of the plurality of sound receiving units in time series, and generates the two-dimensional map for each divided sound signal.
16. A method for determining authenticity, which generates restored sound by restoring a sound signal generated by collecting sound whose characteristics have been changed by a sound receiving unit in the analog domain when collecting sound, and determines the authenticity of the generated restored sound.
17. A program that causes a computer to execute the following process: generate restored sound by restoring a sound signal generated by collecting sound whose characteristics have been changed by a sound receiving unit in the analog domain when collecting sound; and determine the authenticity of the generated restored sound.