Apparatus and method for processing audio input recordings to obtain processed audio recordings to address privacy issues.

The apparatus and method process audio recordings to detect and modify speech based on privacy rules, effectively addressing privacy concerns and enabling compliant audio processing.

JP7851328B2Active Publication Date: 2026-04-24FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
Filing Date
2022-04-13
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing audio recording technologies fail to adequately address privacy concerns by not filtering or modifying voice recordings, relying instead on organizational measures or low-resolution recording methods that hinder useful processing.

Method used

An apparatus and method that processes audio input recordings by detecting speech and applying various rules to modify or exclude audio portions based on privacy requirements, using techniques like source separation, speaker identification, and automatic speech recognition to ensure privacy.

Benefits of technology

Effectively addresses privacy issues by modifying audio recordings to protect speech privacy, allowing useful processing while ensuring compliance with privacy laws and regulations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007851328000001
    Figure 0007851328000001
  • Figure 0007851328000002
    Figure 0007851328000002
  • Figure 0007851328000003
    Figure 0007851328000003
Patent Text Reader

Abstract

According to one embodiment, an apparatus for processing an audio input recording to obtain a processed audio recording is provided. The apparatus comprises an input interface (110) for receiving a plurality of audio input portions of the audio input recording. The apparatus further comprises a processor (120) for processing the plurality of audio input portions of the audio input recording to obtain a processed audio recording. The processor (120) is configured to determine whether an audio input portion of the plurality of audio input portions includes speech. If the processor (120) detects that the audio input portion includes speech, the processor (120) is configured to generate the processed audio recording by modifying the audio input portion to obtain a modified audio portion and by generating the processed audio recording such that the processed audio recording includes the modified audio portion instead of the audio input portion. Alternatively, if the processor (120) detects that the audio input portion includes speech, the processor (120) is configured to generate the processed audio recording such that the processed audio recording does not include the audio input portion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an apparatus and a method for processing an audio input recording to obtain a processed audio recording. In particular, the present invention relates to processing an audio input recording so as to appropriately address aspects of privacy.

Background Art

[0002] Acoustic recordings in public spaces have become a subject of debate, despite the actual need for such recordings for, for example, autonomous driving, ecological monitoring, noise monitoring, security-related facilities or production facilities. In particular, voice as a protectable entity must be specifically protected depending on the situation. While addressing privacy concerns, it is desirable to provide recording means suitable for recording external sounds (for example, recording means in the automotive field). For example, when considering an external microphone of a vehicle, the voices of pedestrians can also be recorded by such recording means for recording external sounds, for example, and should appropriately address data protection and privacy concerns. At present, there is no known prior art concept for investigating an audio recording for the presence of voices and taking measures for filtering voices in order to address privacy concerns or to make voices unintelligible. Since the prior art does not provide a technical solution, today's privacy guarantees are usually made by organizational means. (Warning signs indicating that recording is being done, declarations of consent, guarantees that no third parties are present, only researchers have access rights recognized as appropriate by an ethics committee, storing data only on a strictly protected drive) or by extensive manual post-processing. Bitzer et al. [1] have proposed a method for recording audio at a very low resolution so that intelligible speech cannot be reconstructed from such recordings. Such a method ensures privacy but also modifies portions of the audio signal that do not show speech activity, and as a result, further processing of such recordings is not useful or its use is limited. Starting from the above, improvements or enhancements are needed regarding the processing of audio input recordings to obtain processed audio recordings, so that privacy aspects are adequately addressed. [Overview of the project] [Problems that the invention aims to solve]

[0003] An apparatus is provided for processing an audio input recording in order to obtain a processed audio recording, according to one embodiment. [Means for solving the problem]

[0004] The device includes an input interface for receiving multiple audio input portions of an audio input recording. Furthermore, the device includes a processor for processing the multiple audio input portions of the audio input recording in order to obtain a processed audio recording. The processor is configured to determine whether or not one of the multiple audio input portions contains speech. If the processor detects that the audio input portion contains speech, the processor is configured to generate a processed audio recording by modifying the audio input portion and obtaining the modified audio portion, and by generating a processed audio recording such that the processed audio recording contains the modified audio portion instead of the audio input portion. Alternatively, if the processor detects that the audio input portion contains speech, the processor is configured to generate a processed audio recording such that the processed audio recording does not contain the audio input portion.

[0005] Furthermore, a method is provided for processing an audio input recording in order to obtain a processed audio recording, according to one embodiment. - Receiving multiple audio input portions of an audio input recording, - This includes processing multiple audio input portions of an audio input recording and obtaining the processed audio recording. Processing multiple audio input sections is - This includes determining whether one of the multiple audio input sections contains sound, -If it is detected that the audio input contains speech, the processed audio recording is generated by modifying the audio input to obtain the modified audio portion, and then generating the processed audio recording so that it contains the modified audio portion instead of the audio input. Alternatively, if it is detected that the audio input contains speech, the processed audio recording is generated so that it does not contain the audio input.

[0006] Furthermore, according to one embodiment, a non-temporary computer program product is provided which, when executed on a computer, includes a computer-readable medium that stores instructions for performing the method described above. Furthermore, if the method is carried out on a computer or signal processor, a computer program for carrying out the above-described method is provided. Furthermore, a microphone according to one embodiment is provided, and the above-described device is integrated into the microphone. Furthermore, an application-specific integrated circuit according to one embodiment is provided, and the above-mentioned device is integrated into the application-specific integrated circuit. Further specific embodiments are provided in the dependent claims. [Brief explanation of the drawing]

[0007] [Figure 1]This figure shows an apparatus for processing audio input recordings in order to obtain a processed audio recording, according to one embodiment. [Figure 2] This figure shows a device according to one embodiment, in which the device further includes a user interface. [Figure 3] This figure shows a device according to one embodiment, in which the device further includes memory. [Figure 4] This figure shows an apparatus according to one embodiment, further comprising an audio signaling output module. [Figure 5] This figure shows an apparatus according to one embodiment, in which the apparatus further comprises a processing signaling output module. [Figure 6] This figure shows a device according to one embodiment, in which the device further includes an input device. [Modes for carrying out the invention]

[0008] Embodiments of the present invention will now be described in more detail with reference to the accompanying drawings, where the same or similar elements are assigned the same reference numerals. Figure 1 shows an apparatus for processing audio input recordings to obtain a processed audio recording, according to one embodiment. The device includes an input interface 110 for receiving multiple audio input portions of an audio input recording. Furthermore, the device includes a processor 120 for processing multiple audio input portions of an audio input recording in order to obtain a processed audio recording. The processor 120 is configured to determine whether or not one of the multiple audio input sections contains sound. If the processor 120 detects that the audio input portion contains audio, the processor 120 is configured to generate a processed audio recording by modifying the audio input portion and obtaining the modified audio portion, and by generating a processed audio recording such that the processed audio recording includes the modified audio portion instead of the audio input portion. Alternatively, if the processor 120 detects that the audio input portion contains sound, the processor 120 is configured to generate a processed audio recording such that the processed audio recording does not include the audio input portion.

[0009] According to one embodiment, the processor 120 may output to another application, for example, the result of determining whether or not the audio input portion contains sound. In one embodiment, if the processor 120 detects that the audio input portion does not contain sound, the processor 120 may be configured to generate a processed audio recording such that the processed audio recording includes the audio input portion. According to one embodiment, the processor 120 may be configured to perform post-processing on the processed audio recording to obtain a post-processed audio recording. For example, the processor 120 may be configured to resample the processed audio recording to obtain a post-processed audio recording.

[0010] According to one embodiment, if the processor 120 detects that the audio input portion contains sound and that the audio input portion should be processed according to a first processing rule, the processor 120 may be configured to generate a processed audio recording such that the processed audio recording does not include the audio input portion. In one embodiment, if the processor 120 detects that the audio input portion contains speech and that the audio input portion should be processed according to a second processing rule, the processor 120 may be configured to modify the audio input portion, for example, so that the speech within the modified audio portion is unintelligible, and then obtain the modified audio portion. According to one embodiment, if the processor 120 detects that the audio input portion contains sound, and that the audio input portion should be processed according to a third processing rule, the processor 120 may be configured to modify the audio input portion, for example, so that the sound is filtered out of the audio input portion, and to obtain the modified audio portion. In one embodiment, the processor 120 may be configured to modify the audio input portion and obtain a modified audio portion, for example, by using the concept of source separation so that speech is filtered from the audio input portion so that only non-speech components remain in the processed portion of the audio recording.

[0011] According to one embodiment, if the processor 120 detects that an audio input portion contains speech and that the audio input portion should be processed according to a fourth processing rule, the processor 120 may be configured to modify the audio input portion and obtain the modified audio portion, for example, such that the speech in the modified audio portion remains intelligible, but the speaker of the speech can no longer be identified by analyzing the modified audio portion. In one embodiment, if the processor 120 detects that the audio input portion contains speech and the audio input portion should be processed according to a fifth processing rule, the processor 120 may be configured to generate an audio recording processed by using speaker identification and / or automatic speech recognition and / or voice filtering, such that, for example, if the speech is emitted from a previously identified speaker or a speaker trained on a device, the speech remains intelligible in the modified audio portion; otherwise, the processed audio recording does not contain the audio input portion, or the modified audio portion is generated using voice filtering so that only speech from a previously identified speaker or a speaker trained on a device is intelligible. Alternatively, the processor 120 may be configured to produce a processed audio recording that has been processed by, for example, using speaker identification and / or automatic speech recognition and / or voice filtering, such that if the speech is emitted from a previously identified speaker or a speaker trained on a device, the processed audio recording does not include the audio input portion, or the modified audio portion is generated using a voice filter, so that the speech from the previously identified speaker or the speaker trained on the device is incomprehensible, or the speech remains incomprehensible in the modified audio portion.

[0012] According to one embodiment, when the processor 120 detects that the audio input portion contains voice, and when the audio input portion is to be processed according to the sixth processing rule, the processor 120 may be configured to generate an audio recording processed by using automatic speech recognition such that the processed audio recording includes the audio input portion only when the voice in the audio input portion includes a predefined first keyword. And / or the processor 120 may be configured to generate an audio recording processed by using automatic speech recognition such that the processed audio recording includes the audio input portion only when the voice in the audio input portion does not include a predefined second keyword. And / or the processor 120 may be configured to generate an audio recording processed by using automatic speech recognition such that the processed audio recording includes the audio input portion only when the voice in the audio input portion does not include a name.

[0013] In one embodiment, when the processor 120 detects that the audio input portion contains voice, and when the audio input portion is to be processed according to the seventh processing rule, the processor 120 may be configured to determine, for example, a value indicating the degree of understanding of the voice in the audio input portion, and the processor 120 may be configured to generate a processed audio recording such that the processed audio recording includes the audio input portion according to the value indicating the degree of understanding. According to one embodiment, the processor 120 may be configured to execute a threshold test that compares the value with a threshold to determine whether to generate a processed audio recording such that the processed audio recording includes the audio input portion. In one embodiment, the processor 120 may be configured to process the audio input portion, for example, according to the first one of a group of processing rules, the group of processing rules may include at least two of, for example, a first processing rule, a second processing rule, a third processing rule, a fourth processing rule, a fifth processing rule, a sixth processing rule, and a seventh processing rule. The processor 120 may be configured to process another audio input portion among a plurality of audio input portions, for example, according to the second one of the group of processing rules, and the second one of the group of processing rules may be different from the first one of the group of processing rules, for example.

[0014] FIG. 2 shows an apparatus according to an embodiment, the apparatus further comprising a user interface 115. The user interface 115 is configured to provide means for a user to select a processing rule from a group of processing rules, the group of processing rules may include at least two of, for example, a first processing rule, a second processing rule, a third processing rule, a fourth processing rule, a fifth processing rule, a sixth processing rule, and a seventh processing rule. The processor 120 is configured to process the audio input portion according to the processing rule selected by the user. According to one embodiment, the group of processing rules may include at least three of, for example, a first processing rule, a second processing, a third processing rule, a fourth processing rule, a fifth processing rule, a sixth processing rule, and a seventh processing rule. In one embodiment, the group of processing rules may include at least four of, for example, a first processing rule, a second processing, a third processing rule, a fourth processing rule, a fifth processing rule, a sixth processing rule, and a seventh processing rule. According to one embodiment, the group of processing rules may include at least five of, for example, a first processing rule, a second processing, a third processing rule, a fourth processing rule, a fifth processing rule, a sixth processing rule, and a seventh processing rule. In one embodiment, the group of processing rules may include, for example, at least six of the following: a first processing rule, a second processing rule, a third processing rule, a fourth processing rule, a fifth processing rule, a sixth processing rule, and a seventh processing rule. According to one embodiment, the group of processing rules may include, for example, a first processing rule, a second processing rule, a third processing rule, a fourth processing rule, a fifth processing rule, a sixth processing rule, and a seventh processing rule. In one embodiment, the processor 120 may be configured to determine whether or not the audio input portion contains speech, for example, by using machine learning speech activity detection. According to one embodiment, the processor 120 may be configured to store the processed audio recording in the memory 130, for example.

[0015] Figure 3 shows an apparatus according to one embodiment, the apparatus further comprising a memory 130. In one embodiment, the processor 120 may be configured to store, for example, the audio input portion in the memory 130. The processor 120 may also be configured to process the audio input portion according to, for example, a first processing rule, or a second processing rule, or a third processing rule, or a fourth processing rule, or a fifth processing rule, or a sixth processing rule, or a seventh processing rule, and the processor 120 may be configured to, for example, replace the audio input portion in the memory 130 with a modified audio portion, or remove the audio input portion from the memory 130 without replacement, depending on the processing. According to one embodiment, the processor 120 may store, for example, information in the memory that can indicate, for example, whether or not sound is present in the audio input portion. According to one embodiment, the processor 120 may be configured to determine metadata, for example, the number of speakers present in the audio input portion, and / or the metadata indicating whether the speaker is male or female, and / or the metadata indicating whether background noise is present, and / or the metadata indicating what kind of background noise is present, and / or the metadata indicating the deleted or separated portion of the audio input recording. In one embodiment, the metadata indicates the reason why the deleted or separated portion of the audio input recording was deleted or separated.

[0016] Figure 4 shows an apparatus according to one embodiment, which further comprises an audio signaling output module 140 configured to signal whether or not sound has been detected by using a display, and / or by using an acoustic signal, and / or by using an optical signal, and / or by using a tactile signal, and / or by using an electronic signal.

[0017] Figure 5 shows an apparatus according to one embodiment, further including a processing signaling output module 150 configured to signal whether a processing rule for processing an audio input recording is applicable, and / or which of a plurality of processing rules for processing an audio input recording is applicable, and / or which of a plurality of processing rules for processing an audio input recording is not applicable. The processing signaling output module 150 is configured to use a display, and / or use an acoustic signal, and / or use an optical signal, and / or use a tactile signal, and / or use an electronic signal for signaling.

[0018] Figure 6 shows an apparatus according to one embodiment, which further comprises an input device 118 configured to allow the user to input what steps should be taken to ensure privacy when the modified audio recording is stored. In one embodiment, the device may be adapted for use in a public environment, for example. The following describes specific embodiments of the present invention. For example, a machine learning (ML) based model is used to determine whether a recording, such as one from a microphone or solid-borne sound sensor, contains speech. In other words, this is voice activity detection (VAD) or speech activity detection (SAD). This information is used to control or modify the recording. For example, according to one embodiment, voice activity detection, such as ML voice activity detection, is performed for audio recording. If no sound is detected, the audio recording is saved for further use.

[0019] If sound is detected, one of the following embodiments will be applied. According to the first embodiment, the portion of the audio recording in which sound is detected is not stored, resulting in a gap in the stored audio recording. According to the second embodiment, the portion of the audio recording in which speech is detected is modified so that the audio portion becomes unintelligible (for example, by applying one of the concepts proposed by Bitzer et al. [1]) and the reconstruction of spoken language becomes impossible. According to the third embodiment, in the portion of the audio recording where sound is detected, the sound is filtered out, for example by using the sound source separation concept, leaving only non-sound components in the processed portion of the audio recording. According to the fourth embodiment, the portion of the audio recording in which speech is detected is modified such that the speech remains intelligible, but it is no longer possible to identify the speaker from the processed portion of the audio recording.

[0020] According to a fifth embodiment, speaker identification and / or automatic speech recognition and / or voice filtering are used, and as a result, if the speech is emitted from a previously identified speaker or a speaker trained on the device, the portion of the audio recording in which the speech is detected is recorded so that the speech remains intelligible; otherwise, the portion of the audio recording in which the speech is detected is not recorded, or is recorded using voice filtering so that only the speech of a predefined speaker is intelligible. In a further embodiment, the speech is stored only if the audio portion is not emitted from a predefined speaker. According to the sixth embodiment, an automatic speech recognition device is used such that speech is stored only if it contains a predefined keyword (for example, if the audio recording includes machine commands). Other parts of the speech are not stored or are modified so as not to be understood. In one embodiment, all speech components other than names are stored. In a further embodiment, speech is stored only if the speech part does not contain a keyword (confidentiality). According to the seventh embodiment, the level of speech comprehension is determined, for example, by performing model calculations or model estimations. The model estimates the current level of comprehension and uses (e.g., non-binary) thresholds to determine whether to store the audio data.

[0021] Further specific embodiments are used below. According to further embodiments, the audio recording is fully stored, and after a predefined period, one of the embodiments described above is applied. For example, the complete audio recording may be recorded to an edge or the cloud, for example, as it may be required for automatic speech recognition. The portion of the audio recording containing the voice is then deleted or modified according to the described embodiments. In another embodiment, the device includes an interface for selecting one of the embodiments described above, for example, depending on the application scenario. In further embodiments, metadata related to the audio recording is determined and / or stored. For example, the metadata may indicate, for example, how many speakers are present and / or, for example, whether the speakers are male or female and / or, for example, whether background noise is present and / or what kind of background noise is present. In one embodiment, the metadata may be determined and / or stored to indicate, for example, deleted or separated portions of the audio recording. For example, the metadata may indicate or allow determination of, for example, why a deleted or separated portion of the audio recording was deleted or separated.

[0022] According to another embodiment, a recording device is provided that signals (e.g., in real time) whether or not sound has been detected using a display and / or an acoustic signal and / or an optical signal and / or a tactile signal and / or an electronic signal. In another embodiment, a recording device is provided that signals whether one of the above embodiments for modifying an audio recording is applicable, and / or which of the above embodiments for modifying an audio recording is applicable, and / or which of the above embodiments for modifying an audio recording is not applicable. For example, such information may be provided to the user, for example, using a display and / or using an acoustic signal and / or using an optical signal and / or using a tactile signal and / or an electronic signal.

[0023] In a further embodiment, a device is provided that allows the user to input (for example, by using a button and / or a switch) what steps should be taken to ensure privacy if, for example, the user detects that voice activity was not detected in error, and / or the user does not want to rely on the voice activity detection decision. In another embodiment, user input is used to improve one or more of the concepts (e.g., used) for voice activity detection. For example, post-trained concepts and / or reinforcement learning can be used. Embodiments of the present invention enable or support compliance with the law.

[0024] In some embodiments, the above-described embodiments may be integrated into, for example, a microphone, or implemented by, for example, an application-specific integrated circuit (ASIC), enabling the application of the audio technology to a public environment. Embodiments of the present invention are used and essential for all applications that utilize microphones installed in public environments or workplaces, particularly when clear audio signal recording is required. Embodiments of the present invention can be used, for example, in a vehicle measuring device comprising one or more sensors, such as one or more microphones. Furthermore, embodiments of the present invention can be used, for example, in recording devices used in factories. Furthermore, embodiments of the present invention can be used, for example, in smart speakers or voice control assistance devices. Furthermore, embodiments of the present invention can be used, for example, in a dosimeter for measuring noise that does not evaluate speech. Furthermore, embodiments of the present invention can be used (for example, in real time or offline) in software products for modifying audio recordings, which can be implemented, for example, as a standalone software product or as a plug-in, for example, in an audio editor or for example, in a digital audio workstation.

[0025] While some embodiments have been described in the context of the apparatus, it is clear that these embodiments also represent descriptions of the corresponding methods, where blocks or devices correspond to method steps or features of method steps. Similarly, embodiments described in the context of method steps also represent descriptions of the corresponding blocks, items, or features of the corresponding apparatus. Some or all of the method steps may be performed by (or using) hardware devices such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps may be performed by such devices. Depending on specific implementation requirements, embodiments of the present invention may be implemented in hardware or software, or at least partially in hardware or at least partially in software. Implementation may be carried out using a digital storage medium storing electronically readable control signals, such as a floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or FLASH memory, which may (or may) cooperate with a programmable computer system to perform the respective methods. Therefore, the digital storage medium may be computer-readable.

[0026] Some embodiments of the present invention include a data carrier having an electronically readable control signal that can cooperate with a programmable computer system so that one of the methods described herein is performed. In general, embodiments of the present invention can be implemented as a computer program product having program code, the program code operates to perform one of the methods when the computer program product is executed on a computer. The program code can be stored, for example, in a machine-readable carrier. Other embodiments include a computer program stored in a machine-readable carrier for performing one of the methods described herein. In other words, one embodiment of the method of the present invention is a computer program having program code for performing one of the methods described herein when the computer program is executed on a computer. Accordingly, a further embodiment of the method of the present invention is a data carrier (or digital storage medium, or computer-readable medium) containing a computer program for performing one of the methods described herein. The data carrier, digital storage medium, or recorded medium is typically tangible and / or non-temporary. Accordingly, a further embodiment of the method of the present invention is a data stream or signal sequence representing a computer program for performing one of the methods described herein. The data stream or signal sequence may be configured to be transmitted, for example, over a data communication connection, such as the Internet.

[0027] Further embodiments include processing means configured or adapted to perform one of the methods described herein, such as a computer or a programmable logic device. Further embodiments include a computer on which a computer program for performing one of the methods described herein is installed. Further embodiments of the present invention include an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may include, for example, a file server for transferring the computer program to the receiver. In some embodiments, a programmable logic device (e.g., a field-programmable gate array) can be used to perform some or all of the functions of the method herein. In some embodiments, a field-programmable gate array can cooperate with a microprocessor to perform one of the methods herein. Generally, the method is preferably performed by any hardware device.

[0028] The embodiments described above are merely illustrative of the principles of the present invention. Modifications and variations of the configurations and details described herein will be apparent to those skilled in the art. Therefore, it is intended that the invention be limited only by the imminent claims and not by the specific details presented as part of the description and explanation of the embodiments herein. Although each claim refers to only one single claim, this disclosure also encompasses any possible combination of claims. References [1] Bitzer, J., Kissner, S. & Holube, I.: Privacy-Aware Acoustic Assessments of Everyday Life. JAES 64(6), pp.395-404.

Claims

1. A device for processing an audio input recording in order to obtain a processed audio recording, wherein the device is An input interface for receiving multiple audio input portions of the aforementioned audio input recording, The system includes a processor for processing multiple audio input portions of the aforementioned audio input recording and obtaining the processed audio recording. The processor is configured to determine whether or not one of the multiple audio input portions contains sound. The processor is configured to generate a processed audio recording by, when it detects that the audio input portion contains sound, modifying the audio input portion and obtaining the modified audio portion, and by generating the processed audio recording such that the processed audio recording includes the modified audio portion instead of the audio input portion. If the processor detects that the audio input portion contains speech and that the audio input portion should be processed according to specific processing rules, the processor is configured to generate the processed audio recording by using automatic speech recognition and / or speaker identification such that, if the speech is emitted from a previously identified speaker or the speaker who trained the device, the speech remains intelligible in the modified audio portion; otherwise, the processed audio recording does not contain the audio input portion, or the modified audio portion is generated using a voice filter such that only speech from the previously identified speaker or the speaker who trained the device is intelligible. Device.

2. If the processor detects that the audio input portion does not contain sound, the processor is configured to generate the processed audio recording such that the processed audio recording includes the audio input portion. The apparatus according to claim 1.

3. If the processor detects that the audio input portion contains sound, and that the audio input portion should be processed according to a first processing rule, the processor is configured to generate the processed audio recording such that the processed audio recording does not include the audio input portion. The apparatus according to claim 1.

4. If the processor detects that the audio input portion contains sound and that the audio input portion should be processed according to a second processing rule, the processor is configured to modify the audio input portion so that the sound in the modified audio portion is incomprehensible and to obtain the modified audio portion. The apparatus according to claim 1.

5. If the processor detects that the audio input portion contains sound and that the audio input portion should be processed according to a third processing rule, the processor is configured to modify the audio input portion so that the sound is filtered out of the audio input portion and to obtain the modified audio portion. The apparatus according to claim 1.

6. The processor is configured to modify the audio input portion and obtain the modified audio portion by using the concept of source separation so that the audio is filtered from the audio input portion, so that only non-voice components remain in the processed portion of the audio recording. The apparatus according to claim 5.

7. If the processor detects that the audio input portion contains speech and that the audio input portion should be processed according to a fourth processing rule, the processor is configured to modify the audio input portion and obtain the modified audio portion such that the speech in the modified audio portion remains intelligible, but it is no longer possible to identify the speaker of the speech by analyzing the modified audio portion. The apparatus according to claim 1.

8. If the processor detects that the audio input portion contains sound and that the audio input portion should be processed according to a sixth processing rule, the processor is configured to generate the processed audio recording by using automatic speech recognition such that the processed audio recording includes the audio input portion only if the sound in the audio input portion contains a predefined first keyword. The apparatus according to claim 1.

9. If the processor detects that the audio input portion contains sound and that the audio input portion should be processed according to a sixth processing rule, the processor is configured to generate the processed audio recording by using automatic speech recognition such that the processed audio recording includes the audio input portion only if the sound in the audio input portion does not contain a predefined second keyword. The apparatus according to claim 1.

10. If the processor detects that the audio input portion contains speech and that the audio input portion should be processed according to a sixth processing rule, the processor is configured to generate the processed audio recording by using automatic speech recognition such that the processed audio recording includes the audio input portion only if the speech in the audio input portion does not contain a name. The apparatus according to claim 1.

11. The processor is configured to detect that the audio input portion contains sound and that the audio input portion should be processed according to a seventh processing rule, to determine a value indicating the intelligibility of the sound in the audio input portion, and to generate the processed audio recording such that the processed audio recording includes the audio input portion, according to the value indicating the intelligibility. The apparatus according to claim 1.

12. The processor is configured to perform a threshold test comparing the value to a threshold to determine whether or not to generate the processed audio recording such that the processed audio recording includes the audio input portion. The apparatus according to claim 11.

13. The processor is configured to process the audio input portion according to the first of a group of processing rules, the group of processing rules includes at least two of the first processing rule, the second processing rule, the third processing rule, the fourth processing rule, the fifth processing rule, the sixth processing rule, and the seventh processing rule, The processor is configured to process another of the plurality of audio input portions according to the second of the group of processing rules, wherein the second of the group of processing rules differs from the first of the group of processing rules. The processor is configured to generate the processed audio recording such that the processed audio recording does not include the audio input portion, in accordance with the first processing rule described above. The processor is configured to obtain the modified audio portion by modifying the audio input portion in accordance with the second processing rule described above, such that the sound in the modified audio portion is incomprehensible. In accordance with the third processing rule described above, the processor is configured to modify the audio input portion so that the sound is filtered from the audio input portion and to obtain the modified audio portion. The processor is configured to modify the audio input portion to obtain the modified audio portion in accordance with the fourth processing rule, such that the speech in the modified audio portion remains intelligible, but the speaker of the speech can no longer be identified by analyzing the modified audio portion. In accordance with the fifth processing rule, which is a specific processing rule, the processor is configured to generate the processed audio recording by using the automatic speech recognition and / or speaker identification such that, if the speech is emitted from the previously identified speaker or the speaker who trained the device, the speech remains intelligible in the modified audio portion; otherwise, the processed audio recording does not include the audio input portion, or the modified audio portion is generated using the voice filter such that only speech from the previously identified speaker or the speaker who trained the device is intelligible. In accordance with the sixth processing rule, the processor is configured to generate the processed audio recording by using automatic speech recognition such that the processed audio recording includes the audio input portion only if the audio in the audio input portion includes a predefined first keyword. The processor is configured to determine a value indicating the intelligibility of the audio in the audio input portion in accordance with the seventh processing rule, and the processor is configured to generate the processed audio recording such that the processed audio recording includes the audio input portion, according to the value indicating the intelligibility. The apparatus according to claim 1.

14. The apparatus includes a user interface, and the user interface is configured to provide the user with means for selecting processing rules from a group of processing rules that includes at least two of the first processing rules, the second processing rules, the third processing rules, the fourth processing rules, the fifth processing rules, the sixth processing rules, and the seventh processing rules. The processor is configured to process the audio input portion according to the processing rules selected by the user, The processor is configured to generate the processed audio recording such that the processed audio recording does not include the audio input portion, in accordance with the first processing rule described above. The processor is configured to obtain the modified audio portion by modifying the audio input portion in accordance with the second processing rule described above, such that the sound in the modified audio portion is incomprehensible. In accordance with the third processing rule described above, the processor is configured to modify the audio input portion so that the sound is filtered from the audio input portion and to obtain the modified audio portion. The processor is configured to modify the audio input portion to obtain the modified audio portion in accordance with the fourth processing rule, such that the speech in the modified audio portion remains intelligible, but the speaker of the speech can no longer be identified by analyzing the modified audio portion. In accordance with the fifth processing rule, which is a specific processing rule, the processor is configured to generate the processed audio recording by using the automatic speech recognition and / or speaker identification such that, if the speech is emitted from the previously identified speaker or the speaker who trained the device, the speech remains intelligible in the modified audio portion; otherwise, the processed audio recording does not include the audio input portion, or the modified audio portion is generated using the voice filter such that only speech from the previously identified speaker or the speaker who trained the device is intelligible. In accordance with the sixth processing rule, the processor is configured to generate the processed audio recording by using automatic speech recognition such that the processed audio recording includes the audio input portion only if the audio in the audio input portion includes a predefined first keyword. The processor is configured to determine a value indicating the intelligibility of the audio in the audio input portion in accordance with the seventh processing rule, and the processor is configured to generate the processed audio recording such that the processed audio recording includes the audio input portion, according to the value indicating the intelligibility. The apparatus according to claim 1.

15. The group of processing rules includes at least three of the first processing rule, the second processing rule, the third processing rule, the fourth processing rule, the fifth processing rule, the sixth processing rule, and the seventh processing rule, The apparatus according to claim 13.

16. The processor is configured to determine whether the audio input portion contains speech using machine learning speech activity detection. The apparatus according to claim 1.

17. The processor is configured to store the processed audio recording in memory. The apparatus according to claim 1.

18. The device includes the memory, The apparatus according to claim 17.

19. The processor is configured to store the audio input portion in memory, The processor is configured to process the audio input portion in accordance with a first processing rule, or a second processing rule, or a third processing rule, or a fourth processing rule, or a fifth processing rule, or a sixth processing rule, or a seventh processing rule. The processor is configured to replace the audio input portion in the memory with the modified audio portion, or to remove the audio input portion from the memory without replacement, in accordance with the processing. The processor is configured to generate the processed audio recording such that the processed audio recording does not include the audio input portion, in accordance with the first processing rule described above. The processor is configured to obtain the modified audio portion by modifying the audio input portion in accordance with the second processing rule described above, such that the sound in the modified audio portion is incomprehensible. In accordance with the third processing rule described above, the processor is configured to modify the audio input portion so that the sound is filtered from the audio input portion and to obtain the modified audio portion. The processor is configured to modify the audio input portion to obtain the modified audio portion in accordance with the fourth processing rule, such that the speech in the modified audio portion remains intelligible, but the speaker of the speech can no longer be identified by analyzing the modified audio portion. In accordance with the fifth processing rule, which is a specific processing rule, the processor is configured to generate the processed audio recording by using automatic speech recognition and / or speaker identification such that the speech remains intelligible in the modified audio portion when the speech is emitted from the previously identified speaker or the speaker who trained the device; otherwise, the processed audio recording does not include the audio input portion, or the modified audio portion is generated using a voice filter such that only speech from the previously identified speaker or the speaker who trained the device is intelligible. In accordance with the sixth processing rule, the processor is configured to generate the processed audio recording by using automatic speech recognition such that the processed audio recording includes the audio input portion only if the audio in the audio input portion includes a predefined first keyword, and The processor is configured to determine a value indicating the intelligibility of the audio in the audio input portion in accordance with the seventh processing rule, and the processor is configured to generate the processed audio recording such that the processed audio recording includes the audio input portion, according to the value indicating the intelligibility. The apparatus according to claim 1.

20. The processor is configured to determine metadata, the metadata indicating the number of speakers present in the audio input portion, and / or whether the speakers are male or female, and / or whether background noise is present, and / or what type of background noise is present, and / or whether the metadata indicates deleted or separated portions of the audio input recording. The apparatus according to claim 1.

21. The metadata indicates the reason why the deleted or separated portion of the audio input recording was deleted or separated. The apparatus according to claim 20.

22. The device includes an audio signaling output module configured to signal whether or not sound has been detected by using a display, and / or by using an acoustic signal, and / or by using an optical signal, and / or by using a tactile signal, and / or by using an electronic signal. The apparatus according to claim 1.

23. The device includes a processing signaling output module configured to signal whether a processing rule for processing the audio input recording is applied, and / or which of a plurality of processing rules for processing the audio input recording is applied, and / or which of a plurality of processing rules for processing the audio input recording is not applied, The processing signaling output module is configured to use a display, and / or an acoustic signal, and / or an optical signal, and / or a tactile signal, and / or an electronic signal for the signaling. The apparatus according to claim 1.

24. A method for processing an audio input recording in order to obtain a processed audio recording, wherein the method is The device's input interface receives multiple audio input portions of the aforementioned audio input recording, The process includes processing multiple audio input portions of the audio input recording using the processor of the device and obtaining the processed audio recording, The processor processes the multiple audio input portions, This includes determining whether or not one of the aforementioned plurality of audio input portions contains sound, If it is detected that the audio input portion contains sound, the audio input portion is modified to obtain the modified audio portion, and the processed audio recording is generated such that the processed audio recording includes the modified audio portion instead of the audio input portion. If it is detected that the audio input portion contains speech and that the audio input portion should be processed according to specific processing rules, the generation of the processed audio recording is performed by using automatic speech recognition and / or speaker identification so that the speech remains intelligible in the modified audio portion if the speech is emitted from a previously identified speaker or the speaker who trained the device; otherwise, the processed audio recording does not contain the audio input portion, or the modified audio portion is generated using a voice filter so that only speech from the previously identified speaker or the speaker who trained the device is intelligible. ,method.

25. When the method described above is performed on a computer, a computer program for performing the method described in claim 24.

26. It is a microphone, A microphone in which the apparatus described in claim 1 is integrated.

27. An integrated circuit for specific applications, An application-specific integrated circuit incorporating the apparatus described in claim 1.

Citation Information

Patent Citations

  • Filtering device

    JP2006201496A

  • Audio signal recording device and electronic file

    JP2008309959A