Information processing system, information processing method, recording medium, and data structure
The information processing system addresses the challenge of ensuring audio data integrity and authenticity by using ear canal biometrics to generate and apply digital watermarks, enhancing speaker verification and preventing tampering.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-12-24
- Publication Date
- 2026-03-04
AI Technical Summary
Existing biometric authentication and data tampering detection technologies are inadequate in ensuring the integrity and authenticity of audio data, particularly in preventing unauthorized copying and tampering, and in accurately verifying the speaker's identity.
An information processing system that acquires biometric information from a user's ear canal, generates a digital watermark based on this information, and applies it to audio data to ensure authenticity and integrity, while also performing biometric authentication and timestamping.
The system ensures the integrity and authenticity of audio data by preventing fraudulent activities and accurately verifying the speaker's identity, even during silent periods, and allows for seamless playback with verification of the speaker's name and recording time.
Smart Images

Figure 0007823670000001 
Figure 0007823670000002 
Figure 0007823670000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to the technical fields of an information processing system, an information processing method, a recording medium, and a data structure. [Background technology]
[0002] Ear acoustic authentication is known as a type of biometric authentication. For example, Patent Document 1 discloses a technology in which a test signal is output from an audio device worn by a subject in the ear, and features related to the subject's ear canal are acquired from the echo signal.
[0003] There are also known techniques for detecting tampering of recorded voice data. For example, Patent Document 2 discloses a technique for verifying tampering of conversation data by attaching a digital signature or certificate with a public key to the voice data. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] International Publication No. 2021 / 130949 [Patent Document 2] Japanese Patent Application Laid-Open No. 2002-230203 Summary of the Invention [Problem to be solved by the invention]
[0005] This disclosure aims to improve upon the techniques disclosed in the prior art documents. [Means for solving the problem]
[0006] One aspect of the information processing system disclosed herein comprises a feature acquisition means for acquiring biometric information of a target, a watermark generation means for generating an electronic watermark based on the biometric information, an audio acquisition means for acquiring audio data including the target's speech, and a watermark assignment means for assigning the electronic watermark to the audio data.
[0007] One aspect of the information processing method of this disclosure is an information processing method executed by at least one computer, which acquires biometric information of a subject, generates a digital watermark based on the biometric information, acquires audio data including the subject's speech, and applies the digital watermark to the audio data.
[0008] One aspect of the recording medium of this disclosure is a computer program recorded on at least one computer that causes the computer to execute an information processing method, which includes acquiring biometric information of a subject, generating an electronic watermark based on the biometric information, acquiring audio data including the subject's speech, and applying the electronic watermark to the audio data.
[0009] One aspect of the data structure disclosed herein is a data structure of audio data acquired by an audio device, which includes metadata including personal information of the speaker of the audio data and time information related to the creation of the data, speech information related to the content of the speaker's speech, biometric authentication information indicating that authentication was performed in the audio device using the speaker's biometric information, device information of the audio device, a timestamp created based on the metadata, the speech information, the biometric authentication information, and the device information, and an electronic signature created based on the metadata, the speech information, the biometric authentication information, the device information, and the timestamp. [Brief explanation of the drawings]
[0010] [Figure 1] 1 is a block diagram showing a hardware configuration of an information processing system according to a first embodiment. [Figure 2] 1 is a block diagram showing a functional configuration of an information processing system according to a first embodiment. [Figure 3] 4 is a flowchart showing the flow of operations performed by the information processing system according to the first embodiment. [Figure 4] FIG. 3 is a sequence diagram showing an example of a recording process by the information processing system according to the first embodiment. [Figure 5]FIG. 4 is a sequence diagram showing an example of a playback process by the information processing system according to the first embodiment. [Figure 6] 2 is a conceptual diagram showing an example of a data structure of audio data handled by the information processing system according to the first embodiment. FIG. [Figure 7] FIG. 10 is a block diagram showing the functional configuration of an information processing system according to a second embodiment. [Figure 8] 10 is a flowchart showing the flow of operations performed by the information processing system according to the second embodiment. [Figure 9] FIG. 10 is a block diagram showing the functional configuration of an information processing system according to a third embodiment. [Figure 10] FIG. 11 is a conceptual diagram illustrating an example of authentication processing by an information processing system according to a third embodiment. [Figure 11] FIG. 10 is a block diagram showing the functional configuration of an information processing system according to a fourth embodiment. [Figure 12] 10 is a flowchart showing the flow of operations performed by the information processing system according to the fourth embodiment. [Figure 13] 10 is a flowchart showing the flow of a search operation by the information processing system according to the fourth embodiment. [Figure 14] FIG. 11 is a block diagram showing the functional configuration of an information processing system according to a fifth embodiment. [Figure 15] FIG. 13 is a diagram showing an example of a seek bar displayed in an information processing system according to a fifth embodiment. [Figure 16] FIG. 13 is a block diagram showing the functional configuration of an information processing system according to a sixth embodiment. [Figure 17] FIG. 13 is a diagram (part 1) showing an example of a seek bar displayed in an information processing system according to a sixth embodiment. [Figure 18] FIG. 22 is a diagram (part 2) showing an example of a seek bar displayed in the information processing system according to the sixth embodiment. [Figure 19] FIG. 13 is a block diagram showing the functional configuration of an information processing system according to a seventh embodiment. [Figure 20]13 is a flowchart showing the flow of operations performed by the information processing system according to the seventh embodiment. [Figure 21] 13 is a flowchart showing the flow of a playback operation by the information processing system according to the seventh embodiment. [Figure 22] FIG. 13 is a block diagram showing the functional configuration of an information processing system according to an eighth embodiment. [Figure 23] 13 is a flowchart showing the flow of operations performed by the information processing system according to the eighth embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0011] Hereinafter, embodiments of an information processing system, an information processing method, a recording medium, and a data structure will be described with reference to the drawings.
[0012] First Embodiment An information processing system according to a first embodiment will be described with reference to FIGS.
[0013] (Hardware configuration) First, the hardware configuration of the information processing system according to the first embodiment will be described with reference to Fig. 1. Fig. 1 is a block diagram showing the hardware configuration of the information processing system according to the first embodiment.
[0014] 1, an information processing system 10 according to the first embodiment includes a processor 11, a RAM (Random Access Memory) 12, a ROM (Read Only Memory) 13, and a storage device 14. The information processing system 10 may further include an input device 15 and an output device 16. The processor 11, RAM 12, ROM 13, storage device 14, input device 15, and output device 16 are connected via a data bus 17.
[0015] The processor 11 loads a computer program. For example, the processor 11 is configured to load a computer program stored in at least one of the RAM 12, the ROM 13, and the storage device 14. Alternatively, the processor 11 may load a computer program stored in a computer-readable storage medium using a storage medium reading device (not shown). The processor 11 may acquire (i.e., load) the computer program from a device (not shown) located outside the information processing system 10 via a network interface. The processor 11 controls the RAM 12, the storage device 14, the input device 15, and the output device 16 by executing the loaded computer program. In particular, in this embodiment, when the processor 11 executes the loaded computer program, a functional block for executing a process of adding a digital watermark to audio data is realized within the processor 11. In other words, the processor 11 may function as a controller that executes each control in the information processing system 10.
[0016] The processor 11 may be configured as, for example, a central processing unit (CPU), a graphics processing unit (GPU), a field-programmable gate array (FPGA), a demand-side platform (DSP), or an application-specific integrated circuit (ASIC). The processor 11 may be configured as one of these, or may be configured to use multiple processors in parallel.
[0017] The RAM 12 temporarily stores computer programs executed by the processor 11. The RAM 12 temporarily stores data that the processor 11 temporarily uses while the processor 11 is executing the computer programs. The RAM 12 may be, for example, a dynamic random access memory (DRAM) or a static random access memory (SRAM). Alternatively, other types of volatile memory may be used instead of the RAM 12.
[0018] The ROM 13 stores computer programs executed by the processor 11. The ROM 13 may also store fixed data. The ROM 13 may be, for example, a P-ROM (Programmable Read Only Memory) or an EPROM (Erasable Read Only Memory). Alternatively, other types of non-volatile memory may be used instead of the ROM 13.
[0019] The storage device 14 stores data that is to be saved long-term by the information processing system 10. The storage device 14 may operate as a temporary storage device for the processor 11. The storage device 14 may include, for example, at least one of a hard disk device, a magneto-optical disk device, an SSD (Solid State Drive), and a disk array device.
[0020] The input device 15 is a device that receives input instructions from a user of the information processing system 10. The input device 15 may include, for example, at least one of a keyboard, a mouse, and a touch panel. The input device 15 may be configured as a mobile terminal such as a smartphone or a tablet. The input device 15 may be, for example, a device that includes a microphone and is capable of voice input. Furthermore, the input device 15 may be configured as a hearable device that a user wears on their ears.
[0021] The output device 16 is a device that outputs information related to the information processing system 10 to the outside. For example, the output device 16 may be a display device (e.g., a display) that can display information related to the information processing system 10. The output device 16 may also be a speaker or the like that can output information related to the information processing system 10 as audio. The output device 16 may be configured as a mobile terminal such as a smartphone or a tablet. The output device 16 may also be a device that outputs information in a format other than an image. For example, the output device 16 may be a speaker that outputs information related to the information processing system 10 as audio. The input device 15 may also be configured as a hearable device that a user wears on their ear.
[0022] 1 may be provided in a device other than the information processing system 10. For example, the information processing system 10 may be configured to include only the processor 11, RAM 12, and ROM 13 described above, and the other components (i.e., the storage device 14, the input device 15, and the output device 16) may be provided in, for example, an external device connected to the information processing system 10. Furthermore, some of the calculation functions of the information processing system 10 may be realized by an external device (e.g., an external server, a cloud, etc.).
[0023] (Functional configuration) Next, the functional configuration of the information processing system 10 according to the first embodiment will be described with reference to Fig. 2. Fig. 2 is a block diagram showing the functional configuration of the information processing system according to the first embodiment.
[0024] As shown in FIG. 2, the information processing system 10 according to the first embodiment is configured to include a hearable device 50 and a processing unit 100. The hearable device 10 is a device (for example, an earphone-type device) that a user wears in their ear and is capable of inputting and outputting audio. Note that the hearable device 50 here is used to acquire biometric information of a target, and may be replaced with another device capable of acquiring biometric information. The processing unit 100 is configured to be capable of executing various processes in the information processing system 10. The hearable device 50 and the processing unit 100 are configured to be capable of transmitting and receiving information to and from each other.
[0025] The hearable device 50 includes a speaker 51, a microphone 52, a feature amount detection unit 53, and a communication unit 54 as components for realizing its functions.
[0026] The speaker 51 is configured to be able to output sound to a subject wearing the hearable device 50. The speaker 51 outputs sound corresponding to audio data played back by the device, for example. The speaker 51 is also configured to be able to output a reference sound for detecting features of the subject's ear canal. Note that a plurality of speakers 51 may be provided.
[0027] The microphone 52 is configured to be able to acquire sounds around the subject wearing the hearable device 50. For example, the microphone 52 is configured to be able to acquire sounds uttered by the subject. The microphone 52 is also configured to be able to acquire echo sounds (i.e., sounds that are generated when a reference sound emitted by the speaker 51 is echoed in the subject's ear canal) for detecting features of the subject's ear canal. Note that a plurality of microphones 52 may be provided.
[0028] The feature detection unit 53 is configured to be able to detect features of the target's ear canal using the speaker 51 and microphone 52 described above. Specifically, the feature detection unit 53 outputs a reference sound from the speaker 51 and acquires the reflected sound using the microphone 52. The feature detection unit 53 then analyzes the acquired reflected sound to detect features of the target's ear canal. Note that the feature detection unit 53 may be configured to be able to perform authentication processing (i.e., earacoustic authentication processing) using the detected ear canal features. Note that a specific method for earacoustic authentication can be appropriately adopted from existing technologies, and therefore a detailed description thereof will be omitted here.
[0029] The communication unit 54 is configured to be able to transmit and receive various data by communicating between the hearable device 50 and other devices. The communication unit 54 is configured to be able to communicate with the processing unit 100. The communication unit 54 may be able to output, for example, sound acquired by the microphone 52 to the processing unit 100. The communication unit 54 may also be able to output features of the ear canal detected by the feature detection unit 54 to the processing unit 100.
[0030] The processing unit 100 includes, as components for realizing its functions, a feature acquisition unit 110, a digital watermark generation unit 120, an audio acquisition unit 130, and a digital watermark attachment unit 140. Each of the feature acquisition unit 110, the digital watermark generation unit 120, the audio acquisition unit 130, and the digital watermark attachment unit 140 may be a functional block realized by, for example, the above-mentioned processor 11 (see FIG. 1).
[0031] The feature amount acquiring unit 110 is configured to be able to acquire the feature amount of the target ear canal detected by the feature amount detecting unit 53 in the hearable device 50. In other words, the feature amount acquiring unit 110 is configured to be able to acquire data related to the feature amount of the ear canal transmitted from the feature amount detecting unit 53 via the communication unit 54.
[0032] The digital watermark generation unit 120 generates a digital watermark from the feature of the target ear canal acquired by the feature acquisition unit 110 (in other words, the feature detected by the feature detection unit 53). The digital watermark is generated to prevent unauthorized copying or tampering of data. Note that the method for generating the digital watermark here is not particularly limited.
[0033] The voice data acquisition unit 130 is configured to be able to acquire voice data including the speech of the target. For example, the voice data acquisition unit 130 acquires voice data acquired by the microphone 52 in the hearable device 50. However, the voice data acquisition unit 130 may also acquire voice data acquired by a terminal other than the hearable device 50. For example, the voice data acquisition unit 130 may acquire voice data from a smartphone owned by the target.
[0034] The digital watermarking unit 140 is configured to be able to add (embed) the digital watermark generated by the digital watermark generating unit 120 to the audio data acquired by the audio data acquiring unit 130. As a result, the audio data including the target's speech is added with a digital watermark generated based on the features of the target's ear canal.
[0035] (Operation flow) Next, the flow of operations (particularly the process of adding a digital watermark) of the information processing system 10 according to the first embodiment will be described with reference to Fig. 3. Fig. 3 is a flowchart showing the flow of operations of the information processing system according to the first embodiment.
[0036] 3, when the operation of the information processing system 10 according to the first embodiment starts, first, the feature acquisition unit 110 acquires features of the ear canal of the target detected by the feature detection unit 53 in the hearable device 50 (step S101). The features of the ear canal of the target acquired by the feature acquisition unit 110 are output to the digital watermark generation unit 120. Thereafter, the digital watermark generation unit 120 generates a digital watermark from the features of the ear canal of the target acquired by the feature acquisition unit 110 (step S102). The digital watermark generated by the digital watermark generation unit 120 is output to the digital watermark attachment unit 140.
[0037] Next, the voice data acquisition unit 130 acquires voice data including the speech of the target (step S103). The voice data acquired by the voice data acquisition unit 130 is output to the digital watermarking unit 140. Note that the acquisition of voice data may be performed simultaneously in parallel with the above-mentioned steps S101 and S102, or may be performed one after the other. The acquisition of voice data may be started and ended in response to an operation by the target (for example, an operation of the record button). Furthermore, the acquisition of voice data may be performed when it is detected that the hearable device 50 is being worn. Alternatively, the acquisition of voice data may be started when the target utters a specific word or in response to a feature of the target's voice.
[0038] Next, the digital watermarking unit 140 adds the digital watermark generated by the digital watermark generating unit 120 to the audio data acquired by the audio data acquiring unit 130 (step S104). The audio data to which the digital watermark has been added may be stored in a database or the like. The configuration in which the information processing system 10 includes a database will be described in detail later.
[0039] (Example of recording process) Next, a more specific example of the flow of the recording process (i.e., the process of acquiring audio data and adding a digital watermark) in the information processing system 10 according to the first embodiment will be described with reference to Fig. 4. Fig. 4 is a sequence diagram showing an example of the recording process by the information processing system according to the first embodiment.
[0040] 4, when the information processing system 10 according to the first embodiment performs a recording process, first, feature data (i.e., data indicating features of the ear canal) is sent from the hearable device 50 worn by the subject to the hearable certification authority. The hearable certification authority then performs authentication on the received feature data and sends information indicating that the ear acoustic authentication was successful to the hearable device 50.
[0041] Next, the hearable device 50 starts recording the audio data. When the recording of the audio data is completed, the recorded audio data is copied (stored) in the data storage server. Here, the data creation time is written to the audio data as metadata. Note that the recorded audio data is provided with a digital watermark generated based on the features used for ear acoustic authentication as described above. The digital watermark may be provided by the hearable device or the data storage server.
[0042] Next, the data storage server sends a request for a biometric authentication certificate and a device certificate to the hearable authentication authority. In response to this request, the hearable authentication authority returns the biometric authentication certificate and the device certificate to the data storage server. Here, the name of the speaker (i.e., the target) is written as metadata into the voice data.
[0043] Next, the data storage server sends the necessary data to the TSA to request a timestamp token. In response to this request, the TSA generates a timestamp token and returns it to the data storage server. After that, the data storage server requests the hearable authentication authority for a whole digital signature. In response to this request, the hearable authentication authority returns the whole digital signature to the data storage server.
[0044] Next, the data storage server sends an electronic signature completion notification to the target. After that, when the target removes the hearable device 50, the target authentication period (i.e., the period during which the target is authenticated as having worn the hearable device 50) ends. Note that if the target is not wearing the hearable device when recording of the audio data is completed, an error may be notified and audio data with authentication may not be generated.
[0045] (Example of regeneration process) Next, a more specific example of the flow of playback processing (i.e., processing when playing back audio data with a digital watermark) in the information processing system 10 according to the first embodiment will be described with reference to Fig. 5. Fig. 5 is a sequence diagram showing an example of playback processing by the information processing system according to the first embodiment.
[0046] 5, when the information processing system 10 according to the first embodiment executes a playback process, a request for audio data is first sent to the data storage server when the playback software is started by the user. In response to this request, the data storage server sends the audio data to the user (playback software).
[0047] Next, the user obtains the public key of the hearable certification authority, decrypts the digital signature, and verifies that it has not been tampered with. After that, the user obtains the public key of the time-stamp authority, decrypts the time-stamp token, verifies that it has not been tampered with, and obtains the time information certified by the time-stamp authority.
[0048] Next, the user sends a biometric authentication and device verification request to the hearable authentication authority. In response to this request, the hearable authentication authority sends a biometric authentication and device OK (i.e., the speaker and device are authenticated) to the user.
[0049] The playback software then starts playing the audio data. When the audio data is played, it may be displayed based on the results of the above processes that the audio data has not been tampered with, and that the speaker's name, the time the data was created, the speaker, the device, and the authentication time are correct. When playing audio data, the user can freely perform operations such as fast-forwarding and rewinding.
[0050] (Data Structure) Next, the data structure of audio data (specifically, audio data with a digital watermark added) handled by the information processing system 10 according to the first embodiment will be described with reference to Fig. 6. Fig. 6 is a conceptual diagram showing an example of the data structure of audio data handled by the information processing system according to the first embodiment.
[0051] As shown in FIG. 6, the audio data to which the digital watermark is added includes metadata D1, speech data D2, a biometric authentication certificate D3, a device certificate D4, a timestamp D5, and an overall digital signature D6.
[0052] The metadata D1 is information including personal information such as the name of the authenticated speaker, and time information regarding the creation of the data.
[0053] The speech data D2 is data (for example, waveform data) including the content of the speech of the speaker. As described above, a digital watermark is added to the speech data D2.
[0054] The biometric authentication certificate D3 is information indicating that authentication was successful using biometric information of the speaker (for example, ear canal features).
[0055] The device certificate D4 is information related to the hearable device 50. The device certificate D4 may be information that certifies that the hearable device 50 that acquired the audio data is an authenticated device.
[0056] The timestamp D5 is information (e.g., information indicating that no tampering has occurred at that time) created based on the metadata D1, the speech data D2, the biometric authentication certificate D3, and the device certificate D4. The timestamp D5 may be created from hash values of the metadata D1, the biometric authentication certificate D3, and the device certificate D4, for example.
[0057] The overall digital signature D6 is a digital signature created based on the metadata D1, the speech data D2, the biometric authentication certificate D3, the device certificate D4, and the timestamp D5.
[0058] The data structure of the voice data described above is merely an example, and the information processing system 10 according to this embodiment can also handle voice data having a data structure different from the above.
[0059] (Technical Effects) Next, the technical effects obtained by the information processing system 10 according to the first embodiment will be described.
[0060] As described with reference to FIGS. 1 to 6 , the information processing system 10 according to the first embodiment generates a digital watermark from biometric features of a target person and adds the generated digital watermark to audio data containing the target person's speech. This ensures the integrity, authenticity, and non-repudiation of the audio data. This prevents fraudulent activities using audio data, such as transmitting audio with content different from the intended voice. Furthermore, by ensuring the integrity of the entire audio data, the intended voice of the target person can be more faithfully reproduced to the listener. For example, in today's news reports, it is common for only spoken parts to be extracted, giving the listener an impression different from the intended voice. Playing back audio data with this system solves this problem. Furthermore, the present invention ensures the integrity of even the nuances of the blank periods during speech when the target person is not making any sound. While voice authentication, for example, requires the target person to speak before authentication can be performed, the present invention enables authentication even during periods when the target person is not speaking. Furthermore, if the audio data includes speech content other than that of the target (for example, if the hearable device 50 also picks up speech content of other people), it is possible to provide proof that the target has heard it.
[0061] In the above embodiment, the hearable device 50 that acquires the feature quantities of the target's ear canal is used as an example, but the device that acquires the feature quantities of the target is not limited to the hearable device 50. For example, instead of the hearable device 50, the feature quantities of the target may be acquired using a device that can acquire at least one of the target's face, iris, voice, and fingerprint. For example, the target's face or iris may be acquired using a camera device. The target's fingerprint may be acquired using a device equipped with a fingerprint sensor. The target's voice may be acquired using a device equipped with a microphone.
[0062] Second Embodiment An information processing system 10 according to the second embodiment will be described with reference to Figures 7 and 8. The second embodiment differs from the first embodiment described above only in part of its configuration and operation, and other parts may be the same as those of the first embodiment. Therefore, the following will describe in detail the parts that differ from the first embodiment already described, and will omit a description of other overlapping parts as appropriate.
[0063] (Functional configuration) First, the functional configuration of the information processing system 10 according to the second embodiment will be described with reference to Fig. 7. Fig. 7 is a block diagram showing the functional configuration of the information processing system according to the second embodiment. Note that in Fig. 8, the same elements as those shown in Fig. 2 are denoted by the same reference numerals.
[0064] As shown in Fig. 7, the information processing system 10 according to the second embodiment includes a first hearable device 50a, a second hearable device 50b, and a processing unit 100. The first hearable device 50a is a device worn by a first subject, and the second hearable device 50b is a device worn by a second subject (i.e., a subject different from the first subject). The first hearable device 50a and the second hearable device 50b are each configured to be able to communicate with the processing unit 100. The first hearable device 50a and the second hearable device 50b may have the same configuration as the hearable device 50 in the first embodiment (see Fig. 2).
[0065] The processing unit 100 according to the second embodiment is configured to include, as components for realizing its functions, a feature acquisition unit 110, a digital watermark generation unit 120, a voice acquisition unit 130, a digital watermark attachment unit 140, and a voice synthesis unit 150. That is, the processing unit 100 according to the second embodiment further includes a voice synthesis unit 150 in addition to the configuration of the first embodiment described above (see FIG. 2). Note that the voice synthesis unit 150 may be a functional block realized by, for example, the processor 11 described above (see FIG. 1).
[0066] The voice synthesis unit 150 is configured to generate synthetic voice data by synthesizing the first voice data acquired from the first hearable device 50a and the second voice data acquired from the second hearable device 50b. The voice synthesis method is not particularly limited. For example, a process of overwriting quiet or noisy voice portions with the other voice data may be performed. For example, in the first voice data acquired by the first hearable device 50a, the speech of the first target is relatively loud, while the speech of the second target is relatively quiet. On the other hand, in the second voice data acquired by the second hearable device 50b, the speech of the first target is relatively quiet, while the speech of the second target is relatively loud. Therefore, by overwriting the corresponding portion of the second voice data over the speech of the second target in the first voice data, the volume difference between the speakers can be optimized.
[0067] (Operation flow) Next, the flow of operations of the information processing system 10 according to the second embodiment will be described with reference to Fig. 8. Fig. 8 is a flowchart showing the flow of operations of the information processing system according to the second embodiment. Note that in Fig. 8, the same processes as those shown in Fig. 3 are denoted by the same reference numerals.
[0068] 8, when the operation of the information processing system 10 according to the second embodiment starts, first, the feature acquisition unit 110 acquires features of the ear canal of the target detected by the feature detection unit 53 in the hearable device 50 (step S101). Note that in the second embodiment, the first hearable device 50a may acquire features of the ear canal of the first target, and the second hearable device 50b may acquire features of the ear canal of the second target.
[0069] Next, the digital watermark generation unit 120 generates a digital watermark from the feature of the ear canal of the target acquired by the feature acquisition unit 110 (step S102). Particularly in the second embodiment, a digital watermark corresponding to a first target may be generated from the feature of the ear canal of a first target, and a digital watermark corresponding to a second target may be generated from the feature of the ear canal of a second target.
[0070] Next, the voice data acquisition unit 130 acquires voice data including the target's speech (step S103). In the second embodiment, first voice data is acquired from the first hearable device 50a, and second voice data is acquired from the second hearable device 50b. Then, the voice synthesis unit 150 synthesizes the first voice data and the second voice data to generate synthetic voice data (step S201).
[0071] Next, the digital watermarking unit 140 adds the digital watermark generated by the digital watermark generation unit 120 to the synthetic audio data synthesized by the audio synthesis unit 150 (step S104). The digital watermarking unit 140 may add both the digital watermark corresponding to the first object and the digital watermark corresponding to the second object, or may add only one of them.
[0072] (Technical Effects) Next, the technical effects obtained by the information processing system 10 according to the second embodiment will be described.
[0073] 7 and 8, in the information processing system 10 according to the second embodiment, first audio data and second audio data acquired from different devices are synthesized, and a digital watermark is added to the synthesized audio data. In this way, it is possible to add a digital watermark while suppressing volume differences and noise caused by differences in the recording environment (i.e., the recording terminal).
[0074] Third Embodiment An information processing system 10 according to the third embodiment will be described with reference to Figures 9 and 10. The third embodiment differs from the first and second embodiments only in some configurations and operations, and other parts may be the same as the first and second embodiments. Therefore, the following will describe in detail the parts that differ from the embodiments already described, and will omit a description of other overlapping parts as appropriate.
[0075] (Functional configuration) First, the functional configuration of the information processing system 10 according to the third embodiment will be described with reference to Fig. 9. Fig. 9 is a block diagram showing the functional configuration of the information processing system according to the third embodiment. Note that in Fig. 9, the same elements as those shown in Fig. 2 are denoted by the same reference numerals.
[0076] As shown in FIG. 9, the information processing system 10 according to the third embodiment includes a hearable device 50 and a processing unit 100. The processing unit 100 according to the third embodiment includes, as components for realizing its functions, a feature acquisition unit 110, a digital watermark generation unit 120, a voice acquisition unit 130, a digital watermark assignment unit 140, a biometric authentication unit 160, and an authentication history storage unit 170. That is, the processing unit 100 according to the third embodiment further includes the biometric authentication unit 160 and the authentication history storage unit 170 in addition to the configuration of the first embodiment described above (see FIG. 2). The biometric authentication unit 160 may be a functional block implemented by, for example, the processor 11 described above (see FIG. 1). The authentication history storage unit 170 may be implemented by, for example, the storage device 14 described above.
[0077] The biometric authentication unit 160 is configured to be able to perform biometric authentication on a target. In particular, the biometric authentication unit 160 is configured to be able to perform biometric authentication at multiple times while recording audio data. For example, the biometric authentication unit 160 may perform biometric authentication at a predetermined cycle (for example, every few seconds or every few minutes). The biometric authentication performed by the biometric authentication unit 160 may be ear acoustic authentication. In this case, the biometric authentication unit 160 may perform biometric authentication using ear canal features acquired by the feature acquisition unit 110. However, the biometric authentication performed by the biometric authentication unit 160 may be other than ear acoustic authentication. For example, the biometric authentication unit 160 may be configured to be able to perform fingerprint authentication, face authentication, and iris authentication. In this case, the biometric authentication unit 160 may acquire features used for biometric authentication using various scanners, cameras, etc.
[0078] The authentication history storage unit 170 is configured to be able to store a result history of biometric authentication performed by the biometric authentication unit 160. Specifically, the authentication history storage unit 170 stores whether or not authentication was successful for each of multiple biometric authentications performed by the biometric authentication unit 160. The history stored in the authentication history storage unit 170 may be made available for confirmation on playback software, for example, when playing back audio data.
[0079] Here, an example has been described in which the processing unit 100 is equipped with a biometric authentication unit 160 and an authentication history storage unit 170, but at least one of the biometric authentication unit 160 and the authentication history storage unit 170 may be equipped in the hearable device 50.
[0080] (Biometric authentication operation) Next, the operation related to biometric authentication by the information processing system 10 according to the third embodiment and the result history to be stored will be described with reference to Fig. 10. Fig. 10 is a conceptual diagram showing an example of authentication processing by the information processing system according to the third embodiment.
[0081] As shown in FIG. 10 , in the information processing system 10 according to the third embodiment, the biometric authentication unit 160 performs biometric authentication at times t1, t2, t3, t4, t5, and so on. The authentication history storage unit 170 stores the results of the biometric authentication at each time. In the example shown in the figure, the following history is stored: biometric authentication successful (OK) at time t1, biometric authentication successful (OK) at time t2, biometric authentication successful (OK) at time t3, biometric authentication failed (NG) at time t4, and biometric authentication successful (OK) at time t5. The authentication history storage unit 170 also stores whether or not the subject was wearing the hearable device 50. In the example shown in the figure, the following history is stored: worn at time t1, worn at time t2, worn at time t3, not worn at time t4, and worn at time t5. From the history described above, it can be seen that the biometric authentication failed, for example, because the subject removed the hearable device 50 at time t4.
[0082] (Technical Effects) Next, the technical effects obtained by the information processing system 10 according to the third embodiment will be described.
[0083] As described with reference to FIGS. 9 and 10 , in the information processing system 10 according to the third embodiment, biometric authentication is performed at multiple times during recording, and the results are stored as a history. In this way, even if the subject is not authenticated based on whether or not the hearable device 50 is worn (for example, even if the period until the hearable device 50 is removed is not the target authentication period, as shown in FIG. 4 ), it is possible to prove from the history that the subject spoke the voice data. Furthermore, by performing biometric authentication at multiple times, it is possible to identify periods during which the subject is not authenticated. Therefore, it is possible to easily detect fraud, such as tampering. Furthermore, even if the target authentication period is set to the period while the hearable device 50 is worn, continuous authentication during that period can prevent the authentication period from being fraudulently changed by disassembling the hearable terminal 50.
[0084] <Fourth embodiment> An information processing system 10 according to the fourth embodiment will be described with reference to Figures 11 to 13. The fourth embodiment differs from the first to third embodiments described above only in part of the configuration and operation, and other parts may be the same as the first to third embodiments. Therefore, the following will describe in detail the parts that differ from the embodiments already described, and will omit a description of other overlapping parts as appropriate.
[0085] (Functional configuration) First, the functional configuration of the information processing system 10 according to the fourth embodiment will be described with reference to Fig. 11. Fig. 11 is a block diagram showing the functional configuration of the information processing system according to the fourth embodiment. In Fig. 11, the same elements as those shown in Fig. 2 are denoted by the same reference numerals.
[0086] 11, the information processing system 10 according to the fourth embodiment is configured to include a hearable device 50, a processing unit 100, and a database 200. That is, the information processing system 10 according to the third embodiment further includes the database 200 in addition to the configuration of the first embodiment (see FIG. 2).
[0087] The database 200 is configured to be capable of storing audio data to which a digital watermark has been added by the processing unit 100. The database 200 may be realized, for example, by the storage device 14 (see FIG. 1) described above. The database 200 includes a search information adding unit 210, a storage unit 220, and an extraction unit 230 as components for realizing its functions.
[0088] The search information adding unit 210 is configured to add search information (information used to search for audio data) to the audio data to which the digital watermark has been added. Specifically, the search information adding unit 210 adds at least one of keywords included in the utterance, information about the target, and the date and time of the utterance to the audio data as search information (i.e., links it to the audio data). Note that the keywords included in the utterance may be obtained, for example, by converting the audio data into text. The information about the target may be personal information such as the target's name, or may be a feature of the target (for example, a feature used for biometric authentication or a voice feature). The date and time of the utterance may be obtained, for example, from a timestamp included in the audio data (see FIG. 6).
[0089] The storage unit 220 is configured to be able to store the voice data to which search information has been added by the search information adding unit 210. The storage unit 220 stores a plurality of voice data to which search information has been added, and is configured to be able to output the voice data as needed in response to a request.
[0090] The extraction unit 230 is configured to be able to extract, from the voice data stored in the accumulation unit 220, data that matches the input search query. Information assigned as search information by the search information assignment unit 210 may be input to the extraction unit 230 as the search query. That is, the search information assignment unit 210 may be input with a search query that includes keywords included in the utterance content, information related to the target, and the date and time of the utterance. Note that the extraction unit 230 may extract only one piece of voice data that matches the search query most closely, or may extract multiple pieces of voice data that match the search query more closely than a predetermined value.
[0091] (Operation flow) Next, the flow of operations (particularly operations up to storing audio data) of the information processing system 10 according to the fourth embodiment will be described with reference to Fig. 12. Fig. 12 is a flowchart showing the flow of operations by the information processing system according to the fourth embodiment. Note that in Fig. 12, the same processes as those shown in Fig. 3 are denoted by the same reference numerals.
[0092] 12, when the operation of the information processing system 10 according to the fourth embodiment starts, first, the feature acquisition unit 110 acquires features of the target's ear canal detected by the feature detection unit 53 in the hearable device 50 (step S101). After that, the digital watermark generation unit 120 generates a digital watermark from the features of the target's ear canal acquired by the feature acquisition unit 110 (step S102).
[0093] Next, the audio data acquisition unit 130 acquires audio data including the target utterance (step S103). After that, the digital watermarking unit 140 adds the digital watermark generated by the digital watermark generation unit 120 to the audio data acquired by the audio data acquisition unit 130 (step S104).
[0094] Next, the search information adding unit 210 adds search information to the audio data to which the digital watermark has been added (step S401). Then, the storage unit 220 stores the audio data to which the search information has been added by the search information adding unit 210 (step S402). Note that the search information adding unit 210 may add the search information to the audio data after it has been stored in the storage unit 220. That is, step S401 may be executed after step S402.
[0095] (Search behavior) Next, an operation of searching for audio data in the information processing system 10 according to the fourth embodiment will be described with reference to Fig. 13. Fig. 13 is a flowchart showing the flow of the search operation by the information processing system according to the fourth embodiment.
[0096] 13, in the search operation by the information processing system 10 according to the fourth embodiment, the extraction unit 230 first receives a search query (step S411). The search query may be input as a word corresponding to the search information. Alternatively, voice (waveform data) or voice characteristics recorded on a terminal such as a smartphone may be used as the search query.
[0097] Next, the extraction unit 230 extracts speech data that matches the input search query from the multiple pieces of speech data stored in the storage unit 220 (step S412). Then, the extraction unit 230 outputs the extracted speech data as a search result (step S413). Note that if no speech data that matches the search query is found, the extraction unit 230 may output a message to that effect as the search result.
[0098] (Technical Effects) Next, the technical effects obtained by the information processing system 10 according to the fourth embodiment will be described.
[0099] As described with reference to Figures 11 to 13, in the information processing system 10 according to the fourth embodiment, search information is added to voice data and stored. In this way, it is possible to appropriately extract desired voice data from the stored voice data. Furthermore, since the search information according to this embodiment includes at least one of keywords included in the speech content, information related to the target, and the date and time of the speech, appropriate extraction can be performed even if the information related to the voice data to be extracted is somewhat vague.
[0100] Fifth Embodiment An information processing system 10 according to the fifth embodiment will be described with reference to Figures 14 and 15. The fifth embodiment differs from the fourth embodiment described above only in part of its configuration and operation, and other parts may be the same as the first to fourth embodiments. Therefore, the following will describe in detail the parts that differ from the embodiments already described, and will omit a description of other overlapping parts as appropriate.
[0101] (Functional configuration) First, the functional configuration of the information processing system 10 according to the fifth embodiment will be described with reference to Fig. 14. Fig. 14 is a block diagram showing the functional configuration of the information processing system according to the fifth embodiment. In Fig. 14, the same elements as those shown in Fig. 11 are denoted by the same reference numerals.
[0102] 14, the information processing system 10 according to the fifth embodiment is configured to include a hearable device 50, a processing unit 100, a database 200, and a playback device 300. That is, the information processing system 10 according to the fifth embodiment further includes the playback device 300 in addition to the configuration of the fourth embodiment described above (see FIG. 11).
[0103] The playback device 300 is configured as a device capable of playing back audio data stored in the database 200. The playback device 300 may be realized, for example, by the above-mentioned output device (see FIG. 1) 16. The playback device 300 includes a speaker 310 and a first display unit 320 as components for realizing its functions.
[0104] The speaker 310 is configured to be able to play back audio data acquired from the database 200. Note that the speaker 310 here may be the speaker 51 included in the hearable device 50. In other words, the hearable device 50 may have the function of the playback device 300.
[0105] The first display unit 320 is configured to be able to display a seek bar when playing back audio data. In particular, the seek bar displayed by the first display unit 320 is displayed in a manner that allows parts that match the search query to be visually recognized. The first display unit 320 may obtain information about parts that match the search query using the extraction results of the extraction unit 230. Specific examples of how the seek bar is displayed will be described in detail below.
[0106] (Example of seek bar display) Next, a display example of a seek bar by the information processing system 10 according to the fifth embodiment will be described with reference to Fig. 15. Fig. 15 is a diagram showing an example of a seek bar displayed in the information processing system according to the fifth embodiment.
[0107] As shown in Fig. 15, in the information processing system 10 according to the fifth embodiment, a seek bar is displayed on a display or the like of a device that plays back audio data. The seek bar represents the entire audio data, with the circled portion indicating the currently played portion. The circled portion gradually moves to the right as playback time passes. Therefore, the portion to the left of the circled portion is the portion that has been played, and the portion to the right of the circled portion is the portion that has not been played.
[0108] In this embodiment, in particular, the portion matching the search query is displayed on the seek bar so that it can be recognized. For example, as shown in the figure, the portion matching the search query may be displayed in a color different from the other portions. However, the portion matching the search query may also be displayed in a display manner other than that described here. The portion matching the search query may be, for example, a portion containing a word included in the search query or a portion spoken by a speaker included in the search query. Alternatively, if a search is performed using recorded audio, the portion corresponding to the recorded audio (waveform) may be determined to be the portion matching the search query.
[0109] (Technical Effects) Next, the technical effects obtained by the information processing system 10 according to the fifth embodiment will be described.
[0110] 14 and 15, in the information processing system 10 according to the fifth embodiment, the seek bar is displayed in a manner that allows the user to recognize the portion of the audio data that matches the search query. This allows the user who performed the search to visually recognize the portion of the audio data that the user wants to know. Sixth Embodiment An information processing system 10 according to the sixth embodiment will be described with reference to Figures 16 to 18. The sixth embodiment differs from the fifth embodiment described above only in part of its configuration and operation, and other parts may be the same as the first to fifth embodiments. Therefore, the following will describe in detail the parts that differ from the embodiments already described, and will omit a description of other overlapping parts as appropriate.
[0111] (Functional configuration) First, the functional configuration of the information processing system 10 according to the sixth embodiment will be described with reference to Fig. 16. Fig. 16 is a block diagram showing the functional configuration of the information processing system according to the sixth embodiment. In Fig. 16, the same elements as those shown in Fig. 14 are denoted by the same reference numerals.
[0112] As shown in FIG. 16, the information processing system 10 according to the sixth embodiment includes a hearable device 50, a processing unit 100, a database 200, and a playback device 300.
[0113] The database 200 according to the sixth embodiment includes, as components for realizing its functions, a storage unit 220 and a play count management unit 240. That is, the database 200 according to the sixth embodiment includes the play count management unit 240 instead of the search information assignment unit 210 and the extraction unit 230 of the database 200 according to the fifth embodiment (see FIG. 14). Note that the database 200 according to the sixth embodiment may be configured to include the search information assignment unit 210 and the extraction unit 230 in addition to the play count management unit 240 (that is, it may have the same search function as in the fifth embodiment).
[0114] The play count management unit 240 manages the number of times each piece of audio data has been played back, stored in the storage unit 220. Specifically, the play count management unit 240 stores the number of times each piece of audio data has been played back for each portion of the audio data. For example, the play count management unit 240 divides the audio data into multiple portions at predetermined time intervals, and stores the number of times each portion has been played back.
[0115] The playback device 300 according to the sixth embodiment includes a speaker 310 and a second display unit 330. That is, the playback device 300 according to the sixth embodiment includes the second display unit 330 instead of the first display unit 320 of the playback device 300 according to the fifth embodiment (see FIG. 14). However, the second display unit 330 may have the function of the first display unit 320 (that is, the function of displaying portions that match the search query).
[0116] The second display unit 330 is configured to be able to display a seek bar when playing audio data. In particular, the seek bar displayed by the second display unit 330 is displayed in a manner that allows frequently played portions to be visually recognized. The second display unit 330 may acquire information about frequently played portions from the play count management unit 240. Specific display examples of the seek bar will be described in detail below.
[0117] (Example of seek bar display) Next, a display example of a seek bar in the information processing system 10 according to the sixth embodiment will be described with reference to Fig. 17 and Fig. 18. Fig. 17 is a diagram (part 1) showing an example of a seek bar displayed in the information processing system according to the sixth embodiment. Fig. 18 is a diagram (part 2) showing an example of a seek bar displayed in the information processing system according to the sixth embodiment.
[0118] As shown in FIG. 17 , in the information processing system 10 according to the sixth embodiment, a seek bar is displayed on a display or the like of a device that plays back audio data. Particularly in this embodiment, a heat map indicating the number of plays may be displayed below the seek bar. In this heat map, darker areas indicate more plays and lighter areas indicate fewer plays. The heat map is generated based on information relating to the number of plays acquired from the play count management unit 240. Alternatively, the play count management unit 240 may store the play count in the form of a heat map.
[0119] As shown in Fig. 18, a graph showing the number of plays may be displayed below the seek bar. This graph shows that the higher the number of plays, the lower the number of plays. The graph is generated based on information related to the number of plays acquired from the play count management unit 240. Alternatively, the play count management unit 240 may store the number of plays in the form of a graph.
[0120] (Technical Effects) Next, the technical effects obtained by the information processing system 10 according to the sixth embodiment will be described.
[0121] 16 to 18, in the information processing system 10 according to the sixth embodiment, the seek bar is displayed in a manner that allows the user to recognize the parts of the audio data that are frequently played. In this way, it is possible to visually recognize the parts of the audio data that are of interest to other users (in other words, popular parts).
[0122] The fifth and sixth embodiments may be combined and implemented. That is, the seek bar may display information indicating the number of times the content has been played, along with the portion that matches the search query.
[0123] Seventh Embodiment An information processing system 10 according to the seventh embodiment will be described with reference to Figures 19 to 21. The seventh embodiment differs only in part of the configuration and operation from the first to sixth embodiments described above, and other parts may be the same as the first to sixth embodiments. Therefore, the following will describe in detail the parts that differ from the embodiments already described, and will omit a description of other overlapping parts as appropriate.
[0124] (Functional configuration) First, the functional configuration of the information processing system 10 according to the seventh embodiment will be described with reference to Fig. 19. Fig. 19 is a block diagram showing the functional configuration of the information processing system according to the seventh embodiment. In Fig. 19, the same elements as those shown in Fig. 14 are denoted by the same reference numerals.
[0125] As shown in FIG. 19, the information processing system 10 according to the seventh embodiment includes a hearable device 50, a processing unit 100, a database 200, and a playback device 300.
[0126] The database 200 according to the seventh embodiment includes, as components for realizing its functions, an accumulation unit 220, a specific user storage unit 250, and a user determination unit 260. That is, the database 200 according to the seventh embodiment includes the specific user storage unit 250 and the user determination unit 260 instead of the search information assignment unit 210 and the extraction unit 230 of the database 200 according to the fifth embodiment (see FIG. 14). Note that the database 200 according to the sixth embodiment may be configured to include the search information assignment unit 210 and the extraction unit 230 in addition to the specific user storage unit 250 and the user determination unit 260 (that is, it may have the same search function as in the fifth embodiment).
[0127] The specific user storage unit 250 is configured to store information about a specific user. The "specific user" here refers to a user other than the target user who has permission to play back audio data with a digital watermark. The information about the specific user is not particularly limited as long as it can identify the specific user. For example, the information may be personal information such as the specific user's name, or biometric information (e.g., feature values) of the specific user. Alternatively, the information may be an ID and password set by the specific user or automatically set by the system. Note that, as can be seen from the setting of a specific user, the audio data according to this embodiment may be intended for playback by a user other than the target user. An example of such audio data is data containing a will. In this case, the specific user may be, for example, an heir or a representative.
[0128] The user determination unit 260 is configured to be able to determine whether or not the voice data has been played back by a specific user. The user determination unit 260 is configured to be able to determine whether or not the voice data has been played back by a specific user by comparing user information (i.e., information about the user playing back the voice data) acquired by a user information acquisition unit 340 (described later) with specific user information stored in the specific user storage unit 250. For example, the user determination unit 260 may determine that the voice data has been played back by a specific user when the user information acquired by the user information acquisition unit 340 matches the specific user information. Furthermore, the user determination unit 260 may determine that the voice data has been played back by a user other than the specific user when the user information acquired by the user information acquisition unit 340 does not match the specific user information.
[0129] The playback device 300 according to the seventh embodiment includes a speaker 310 and a user information acquisition unit 340. That is, the playback device 300 according to the seventh embodiment includes the user information acquisition unit 340 instead of the first display unit 320 of the playback device 300 according to the fifth embodiment (see FIG. 14). Note that the playback device 300 according to the seventh embodiment may include the first display unit 320 (see FIG. 14) or the second display unit (see FIG. 16) in addition to the user information acquisition unit 340. That is, the playback device 300 according to the seventh embodiment may have a function of displaying the seek bar described in the fifth and sixth embodiments.
[0130] The user information acquisition unit 340 is configured to be able to acquire information about the user who reproduces the audio data (hereinafter referred to as "reproduction user information" as appropriate). The reproduction user information is acquired as information that can be compared with the specific user stored in the specific user storage unit 250. The reproduction user information may be acquired by, for example, input by the user himself or automatically acquired using a camera or the like.
[0131] (Operation flow) Next, the flow of operations (particularly operations up to storing audio data) of the information processing system 10 according to the seventh embodiment will be described with reference to Fig. 20. Fig. 20 is a flowchart showing the flow of operations by the information processing system according to the seventh embodiment. Note that in Fig. 20, the same processes as those shown in Fig. 12 are denoted by the same reference numerals.
[0132] 20, when the operation of the information processing system 10 according to the seventh embodiment starts, first, the feature acquisition unit 110 acquires features of the target's ear canal detected by the feature detection unit 53 in the hearable device 50 (step S101). After that, the digital watermark generation unit 120 generates a digital watermark from the features of the target's ear canal acquired by the feature acquisition unit 110 (step S102).
[0133] Next, the audio data acquisition unit 130 acquires audio data including the target utterance (step S103). After that, the digital watermarking unit 140 adds the digital watermark generated by the digital watermark generation unit 120 to the audio data acquired by the audio data acquisition unit 130 (step S104).
[0134] Next, the storage unit 220 stores the audio data to which the digital watermark has been added (step S402). After that, the specific user information storage unit 250 stores information on the specific users who have permission to play the stored audio data (step S701). Note that the specific user information does not need to be added to all audio data. In other words, there may be audio data that is not subject to the determination of whether or not it has been played by a specific user.
[0135] (User determination operation) Next, an operation of reproducing audio data in the information processing system 10 according to the seventh embodiment will be described with reference to Fig. 21. Fig. 21 is a flowchart showing the flow of a reproduction operation by the information processing system according to the seventh embodiment.
[0136] 21, when audio data is played back in the information processing system 10 according to the seventh embodiment, the user information acquisition unit 340 first acquires information about the user who is going to play back the audio data (i.e., playback user information) (step S711). Then, the user determination unit 260 determines whether the playback user information acquired by the user information acquisition unit 340 matches the specific user information stored in the specific user information storage unit 250 (step S712).
[0137] If the playback user information and the specific user information match (step S712: YES), the user determination unit 160 determines that the playback is by the specific user (step S713). On the other hand, if the playback user information and the specific user information do not match (step S712: NO), the user determination unit 160 determines that the playback is by a user other than the specific user (step S714).
[0138] After the above-mentioned determination, playback processing is executed for the audio data (step S715). Note that if the user performing playback is not a specific user, the audio data may not be played back. Alternatively, if the user performing playback is not a specific user, only a portion of the audio data may be played back. Alternatively, if the user performing playback is not a specific user, an alert may be output. Furthermore, the audio data may be played back regardless of whether the user performing playback is a specific user or not. However, in this case, it is preferable to record a history of playback by users other than the specific user.
[0139] If the voice data contains a will, the voice data may be stored together with the text data of the will. In this case, a process may be performed to compare the contents of the voice data with the contents of the text data, for example, when the voice data is generated or played back. If there is a discrepancy or deficiency in the contents, a notification to that effect may be provided.
[0140] (Technical Effects) Next, the technical effects obtained by the information processing system 10 according to the seventh embodiment will be described.
[0141] As described with reference to Figures 19 to 21, the information processing system 10 according to the seventh embodiment determines whether or not audio data has been played back by a specific user. This prevents audio data from being played back illegally by a user who does not have the right to play the data. Alternatively, even if the audio data has been played back illegally, this fact can be ascertained in a later verification.
[0142] Eighth Embodiment An information processing system 10 according to the eighth embodiment will be described with reference to Figures 22 and 23. The eighth embodiment differs only in part of the configuration and operation from the first to seventh embodiments described above, and other parts may be the same as the first to seventh embodiments. Therefore, the following will describe in detail the parts that differ from the embodiments already described, and will omit a description of other overlapping parts as appropriate.
[0143] (Functional configuration) First, the functional configuration of the information processing system 10 according to the eighth embodiment will be described with reference to Fig. 22. Fig. 22 is a block diagram showing the functional configuration of the information processing system according to the eighth embodiment. In Fig. 22, the same elements as those shown in Fig. 11 are denoted by the same reference numerals.
[0144] As shown in FIG. 22, the information processing system 10 according to the eighth embodiment includes a hearable device 50, a processing unit 100, and a database 200.
[0145] The database 200 according to the seventh embodiment includes, as components for realizing its functions, a storage unit 220, a common tag assigning unit 270, and a multi-search unit 280. That is, the database 200 according to the eighth embodiment includes the common tag assigning unit 270 and the multi-search unit 280 instead of the search information assigning unit 210 and the extraction unit 230 of the database 200 according to the fourth embodiment (see FIG. 11). Note that the database 200 according to the eighth embodiment may be configured to include the search information assigning unit 210 and the extraction unit 230 in addition to the common tag assigning unit 270 and the multi-search unit 280 (that is, it may have the same search function as in the fourth embodiment).
[0146] The common tag assigning unit 270 is configured to be able to assign a common tag to the audio data to which a digital watermark has been assigned and to other content data corresponding to that audio data. For example, a tag indicating the common speaker (here, the tag "Mr. A") may be assigned to data containing the same speaker (e.g., "audio data" and "video data" when Mr. A is speaking). Alternatively, a tag indicating the common location (here, "XX conference") may be assigned to data acquired at the same location (e.g., "Mr. B's audio data" and "Mr. C's audio data" when Mr. B and Mr. C are conversing at XX conference). Note that a common tag may be assigned to three or more pieces of data.
[0147] The multi-search unit 280 is configured to be able to simultaneously search for data that has been assigned a common tag using the tag assigned by the common tag assigning unit. For example, it is configured to be able to search for multiple corresponding data by simply inputting a single search query. The search targets of the multi-search unit 280 may be various types of data. Even if different types of data are the search targets, they can be simultaneously searched by using the common tag assigned to them.
[0148] (Operation flow) Next, the flow of operations (particularly operations up to storing audio data) of the information processing system 10 according to the eighth embodiment will be described with reference to Fig. 23. Fig. 23 is a flowchart showing the flow of operations by the information processing system according to the eighth embodiment. Note that in Fig. 23, the same processes as those shown in Fig. 12 are denoted by the same reference numerals.
[0149] 23, when the operation of the information processing system 10 according to the eighth embodiment starts, first, the feature acquisition unit 110 acquires features of the target's ear canal detected by the feature detection unit 53 in the hearable device 50 (step S101). After that, the digital watermark generation unit 120 generates a digital watermark from the features of the target's ear canal acquired by the feature acquisition unit 110 (step S102).
[0150] Next, the audio data acquisition unit 130 acquires audio data including the target utterance (step S103). After that, the digital watermarking unit 140 adds the digital watermark generated by the digital watermark generation unit 120 to the audio data acquired by the audio data acquisition unit 130 (step S104).
[0151] Next, it is determined whether content data corresponding to the audio data to which the digital watermark has been added is stored (step S801). This determination may be made automatically by analyzing each piece of data, or may be made manually.
[0152] If corresponding content data exists (step S801: YES), the common tag assigning unit 270 assigns a common tag to the audio data to which the digital watermark has been added and to the corresponding content data (step S802). Note that if corresponding content data does not exist (step S801: NO), the processing of step S802 described above may be omitted.
[0153] Subsequently, the storage unit 220 stores the audio data to which the digital watermark has been added (step S402). Note that the common tag adding unit 270 may add a common tag to the audio data after the audio data has been stored in the storage unit 220. That is, steps S801 and S802 may be executed after step S402.
[0154] (Technical Effects) Next, the technical effects obtained by the information processing system 10 according to the eighth embodiment will be described.
[0155] 22 and 23, in the information processing system 10 according to the eighth embodiment, a common tag is assigned to multiple corresponding pieces of content. In this way, a multi-search can be performed using the common tag as a search query. Therefore, even if audio and video are stored as separate data in the same location, it is possible to appropriately find each corresponding piece of data.
[0156] In the above-described embodiments, audio data has been described as an example, but by linking the hearable device 50 with a camera, it is possible to target not only audio data but also video data. Furthermore, by linking the hearable device 50 with another microphone, it is also possible to target stereo audio data, etc. Furthermore, by using information from the Global Positioning System (GPS) in the hearable device 50, it is possible to prove the location where a statement was made.
[0157] The information processing system 10 according to each embodiment can be used to record, for example, testimony in court, testimony in business transactions, speeches by company presidents, statements by politicians, etc. It can also be used to record not only statements by one person, but also statements by multiple people (for example, minutes of an online meeting). When people wearing hearable devices 50 are conversing with each other, audio data containing a mix of speeches from multiple people can be handled, making it possible to authenticate the conversation itself. It is also possible to synchronize multiple pieces of audio data based on time information authenticated by a timestamp.
[0158] The scope of each embodiment also includes a processing method in which a program that operates the configuration of each embodiment to realize the functions of the above-described embodiments is recorded on a recording medium, the program recorded on the recording medium is read as code, and the program is executed on a computer. In other words, a computer-readable recording medium is also included in the scope of each embodiment. Furthermore, each embodiment includes not only a recording medium on which the above-described program is recorded, but also the program itself.
[0159] Examples of recording media that can be used include floppy disks, hard disks, optical disks, magneto-optical disks, CD-ROMs, magnetic tapes, non-volatile memory cards, and ROMs. Furthermore, the scope of each embodiment is not limited to programs that execute processing by themselves, but also includes programs that execute processing by operating on an OS in cooperation with other software or functions of an expansion board. Furthermore, the program itself may be stored on a server, and part or all of the program may be downloadable from the server to a user terminal.
[0160] <Additional Notes> The above-described embodiment may be further described as follows, but is not limited to the following.
[0161] (Appendix 1) The information processing system described in Appendix 1 is an information processing system comprising: a feature acquisition means for acquiring biometric information of a target; a watermark generation means for generating an electronic watermark based on the biometric information; an audio acquisition means for acquiring audio data including the target's speech; and a watermark assignment means for assigning the electronic watermark to the audio data.
[0162] (Appendix 2) The information processing system described in Appendix 2 is the information processing system described in Appendix 1, wherein the audio acquisition means acquires first audio data from a first terminal corresponding to a first subject and acquires second audio data from a second terminal corresponding to a second subject present with the first subject, and the watermarking means adds the electronic watermark based on the biometric information acquired from at least one of the first subject and the second subject to synthesized audio data obtained by combining the first audio data and the second audio data.
[0163] (Appendix 3) The information processing system described in Appendix 3 is the information processing system described in Appendix 1 or 2, further comprising a biometric authentication means for performing biometric authentication of the target at multiple times while the voice data is being recorded, and a history storage means for storing a result history of the biometric authentication at the multiple times.
[0164] (Appendix 4) The information processing system described in Appendix 4 is an information processing system described in any one of Appendixes 1 to 3, further comprising: an audio data storage means for storing the audio data to which the electronic watermark has been added, linking it to at least one of keywords contained in the speech content, information about the target, and the date and time of the speech; and an extraction means for extracting audio data that matches the search query from the multiple audio data stored in the storage means, using a search query that includes at least one of keywords contained in the speech content, information about the target, and the date and time of the speech.
[0165] (Appendix 5) The information processing system described in Appendix 5 is the information processing system described in Appendix 4, further comprising a first display means for displaying a seek bar in a display manner that allows visual recognition of parts of the audio data that match the search query when playing back the audio data extracted by the extraction means.
[0166] (Appendix 6) The information processing system described in Appendix 6 is an information processing system described in any one of Appendixes 1 to 5, further comprising a second display means for displaying a seek bar in a manner that allows frequently played parts of the audio data to be visually recognized when playing the audio data to which the electronic watermark has been added.
[0167] (Appendix 7) The information processing system described in Appendix 7 is an information processing system described in any one of Appendixes 1 to 6, further comprising a specific user information storage means for storing information about a specific user who is a user different from the target and who has been authorized to play the audio data to which the electronic watermark has been added, and a determination means for determining whether the audio data has been played by the specific user based on the information about the specific user stored in the specific user information storage means.
[0168] (Appendix 8) The information processing system described in Appendix 8 is an information processing system described in any one of Appendixes 1 to 7, further comprising a tagging means for assigning a common tag to the audio data to which the electronic watermark has been added and to other content data corresponding to the audio data, and a search means for simultaneously searching the audio data and the other content data using the tag.
[0169] (Appendix 9) The information processing method described in Appendix 9 is an information processing method executed by at least one computer, which acquires biometric information of a subject, generates a digital watermark based on the biometric information, acquires audio data including the subject's speech, and applies the digital watermark to the audio data.
[0170] (Appendix 10) The recording medium described in Appendix 10 is a recording medium having recorded thereon a computer program that causes at least one computer to execute an information processing method, which includes acquiring biometric information of a subject, generating an electronic watermark based on the biometric information, acquiring audio data including the subject's speech, and applying the electronic watermark to the audio data.
[0171] (Appendix 11) The computer program described in Appendix 11 is a computer program that causes at least one computer to execute an information processing method of acquiring biometric information of a subject, generating a digital watermark based on the biometric information, acquiring audio data including the subject's speech, and applying the digital watermark to the audio data.
[0172] (Appendix 12) The information processing device described in Appendix 12 is an information processing device comprising: a feature acquisition means for acquiring biometric information of a target; a watermark generation means for generating an electronic watermark based on the biometric information; an audio acquisition means for acquiring audio data including the target's speech; and a watermark assignment means for assigning the electronic watermark to the audio data.
[0173] (Appendix 13) The data structure described in Appendix 13 is a data structure of audio data acquired by an audio device, and includes metadata including personal information of the speaker of the audio data and time information related to the creation of the data, speech information related to the content of the speaker's speech, biometric authentication information indicating that the audio device has authenticated the speaker using his or her biometric information, device information of the audio device, a timestamp created based on the metadata, the speech information, the biometric authentication information, and the device information, and an electronic signature created based on the metadata, the speech information, the biometric authentication information, the device information, and the timestamp.
[0174] This disclosure may be modified as appropriate within the scope that does not contradict the gist or idea of the invention that can be read from the claims and the entire specification, and information processing systems, information processing methods, recording media, and data structures that involve such modifications are also included in the technical idea of this disclosure. [Explanation of symbols]
[0175] 10 Information Processing Systems 11 processors 14 Storage device 15 Input Devices 16 Output Devices 50 Hearable Devices 51 Speaker 52 Mike 53 Feature detection unit 54 Communications Department 100 Processing section 110 Feature acquisition unit 120 Digital watermark generation unit 130 Voice Acquisition Unit 140 Digital watermarking unit 200 databases 210 Search information assignment section 220 Storage Unit 230 Extraction part 240 Playback Count Management Department 250 Specific user information storage unit 260 User Determination Unit 270 Common tagging section 280 Multi-Search Unit 300 Playback device 310 Speaker 320 1st display section 330 2nd display section 340 User information acquisition unit D1 Metadata D2 Speech data D3 Biometric Certificate D4 Device Certificate D5 Timestamp D6 Overall electronic signature
Claims
1. a feature acquisition means for acquiring biometric information of a target; a voice acquisition means for acquiring voice data including the speech of the target; a biometric authentication means for performing biometric authentication of the target at a plurality of times based on the biometric information while the voice data is being recorded; a history storage means for storing a history of the biometric authentication results at the plurality of timings; a watermark generating means for generating a digital watermark when the biometric authentication of the target is successful; a watermarking means for adding the digital watermark to the audio data; An information processing system comprising:
2. the voice acquisition means acquires first voice data from a first terminal corresponding to a first subject, and acquires second voice data from a second terminal corresponding to a second subject present with the first subject; the watermarking means adds the digital watermark based on the biometric information acquired from at least one of the first subject and the second subject to synthesized audio data obtained by synthesizing the first audio data and the second audio data; The information processing system according to claim 1 .
3. a voice data storage means for storing the voice data to which the digital watermark has been added in association with at least one of a keyword included in the speech content, information about the subject, and the date and time of the speech; an extraction means for extracting, from the plurality of pieces of voice data stored in the voice data storage means, voice data that matches the search query, using a search query that includes at least one of a keyword included in the content of the utterance, information about the target, and a date and time of the utterance; The information processing system according to claim 1 or 2, further comprising:
4. The audio data extracting device further includes a first display means for displaying a seek bar in a manner that allows a portion of the audio data that matches the search query to be visually recognized when the audio data extracted by the extraction means is played back. The information processing system according to claim 3 .
5. The audio data playback device further comprises a second display means for displaying a seek bar in a manner that allows a portion of the audio data that has been played frequently to be visually recognized when the audio data to which the digital watermark has been added is played. The information processing system according to any one of claims 1 to 4.
6. a specific user information storage means for storing information about a specific user who is a user different from the target and who is authorized to play the audio data to which the digital watermark is added; a determination means for determining whether the audio data has been reproduced by the specific user based on information about the specific user stored in the specific user information storage means; The information processing system according to claim 1 , further comprising:
7. At least one computer Acquire the subject's biometric information, acquiring voice data including the target's speech; performing biometric authentication of the target at a plurality of times based on the biometric information while the voice data is being recorded; storing a history of the biometric authentication results at the plurality of times; generating a digital watermark if the biometric authentication of the subject is successful; adding the digital watermark to the audio data; Information processing methods.
8. At least one computer Acquire the subject's biometric information, acquiring voice data including the target's speech; performing biometric authentication of the target at a plurality of times based on the biometric information while the voice data is being recorded; storing a history of the biometric authentication results at the plurality of times; generating a digital watermark if the biometric authentication of the subject is successful; adding the digital watermark to the audio data; A computer program that executes an information processing method.
Citation Information
Patent Citations
Voice recording system and voice recording service method
JP2002230203A
Information processing terminal, server and program
JP2008053824A
Individual confirmation method in e-learning learning utilizing biometrics
JP2009276950A
Video monitoring system
JP2014053717A
Audio watermark embedded device, system, and method
JP2016038455A