Speech signal processing device, speech signal playback system and method for outputting an unemotional speech signal
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
- Filing Date
- 2022-08-01
- Publication Date
- 2026-04-23
AI Technical Summary
Existing systems fail to address the challenge of comprehending speech signals enriched with suprasegmental features, such as intonation and emotion, which can lead to misinterpretation and communication difficulties for individuals with autism, cognitive impairments, or non-native speakers, especially in emotionally charged situations.
A speech signal processing device and method that utilizes neural networks or artificial intelligence to separate and transcribe emotion information into additional word information, producing a de-emotionalized speech signal that is tailored to individual needs, either by removing or modifying suprasegmental features.
Facilitates clear communication by transforming emotionally charged speech into a form that is easily understandable, improving comprehension for individuals with autism, cognitive impairments, and non-native speakers, and adapting to individual auditory needs and environments.
Description
[0001] The present invention relates to a speech signal processing device for outputting a de-emotionalized speech signal in real time or after a period of time, a speech signal playback system, a method for outputting a de-emotionalized speech signal in real time or after a period of time, and a computer-readable storage medium.
[0002] To date, no technical system is known that solves a fundamental problem of speech-based communication. The problem lies in the fact that spoken language is always enriched with so-called suprasegmental features (SM) such as intonation, speech rate, duration of pauses, intensity, and loudness. Different dialects of a language can also lead to phonetic variations in the spoken language, which can cause comprehension problems for outsiders. A case in point is the difference between a North German dialect and a South German dialect. Suprasegmental features of language are a phonological property that, across sounds, indicates feelings, illnesses, and individual characteristics (see also Wikipedia on "suprasegmental feature"). These SM convey emotions, in particular, but also aspects that alter the content of the speech to the listener. However, not everyone is able to handle these SM appropriately.to interpret them correctly.
[0003] For example, people with autism have significantly more difficulty accessing the emotions of others. The term autism is used here very broadly for the sake of simplicity. In fact, there are very different forms and degrees of autism (i.e., a so-called autism spectrum). However, this distinction is not necessarily important for understanding the invention. Emotions, as well as changes in content embedded in language via social interaction, are often unrecognizable to them and / or confusing, sometimes leading to a rejection of verbal communication and the use of alternatives such as written language or picture cards.
[0004] Cultural differences or communication in a foreign language can also hinder the information gained from social media communication (SM) or lead to misinterpretations. The situation a speaker is in (e.g., firefighters directly at the scene of a fire) can also lead to highly emotionally charged communication (e.g., with the incident commander), which makes managing the situation more difficult. A similar problem arises with particularly complex language that is difficult for cognitively impaired individuals to understand and may further complicate comprehension in SM.
[0005] The problem has been addressed in various ways, or it has simply remained unresolved. For individuals with autism, various alternative communication methods are used (purely text-based interaction, for example, by writing messages on a tablet, using picture cards, etc.). For individuals with cognitive impairments, so-called "plain language" is sometimes used today, for example, in written announcements or in special news broadcasts. However, a solution that modifies spoken language in real time so that it is understandable for the target groups outlined above is not known.
[0006] The publication by Zhou Kun et al.: "Seen and unseen Emotional Style Transfer for Voice Conversion with a new Emotional Speech Dataset", ICASSP 2021 - 2021 IEEE International Conference on acoustics, speech and signal processing (ICASSP), IEEE, June 6, 2021, pages 920-924, XP033955490, DOI: 10.1109 / ICASSP39728.2021.941339 reveals a seen and unseen emotional style transfer for speech conversion with a new emotional speech dataset.
[0007] One object of the present invention is to provide a speech signal processing device, a speech signal playback system and a method with which a de-emotionalized speech signal of spoken language can be output in real time.
[0008] This problem is solved by the subject matter of the independent patent claims.
[0009] The core idea of the present invention is to provide a transformation of a speech signal containing social media features (SM) and potentially in a particularly complex form into a speech signal wholly or partially freed from SM features and, if necessary, simplified in its formulation, to support speech-based communication for specific groups of listeners, individual listeners, or particular listening situations. The proposed solution comprises a speech signal processing device, a speech signal playback system, and a method that, offline or in real time, removes one or more SM features from a speech signal and presents this freed signal to the listener or stores it in a suitable manner for later playback. A key aspect of this process can be the elimination of emotions.
[0010] A speech signal processing device for outputting a de-emotionalized speech signal in real time or after a time interval is proposed. The speech signal processing device comprises a speech signal acquisition device for acquiring a speech signal. The speech signal comprises at least one piece of emotion information and at least one piece of word information. Furthermore, the speech signal processing device comprises an analysis device, which includes a neural network or artificial intelligence trained to analyze the speech signal with respect to the at least one piece of emotion information and the at least one piece of word information; and a processing device, which includes a neural network or artificial intelligence trained to separate the speech signal into the at least one piece of word information and the at least one piece of emotion information and to process the speech signal.wherein the at least one emotion information is transcribed into further word information either based on training data or on rule-based transcription of recognized emotions, which in turn have been trained inter-individually or intra-individually on the analysis device; and a coupling device and / or a playback device which is / are designed to reproduce the speech signal as a de-emotionalized speech signal, which converts the at least one emotion information into further word information and which comprises at least one word information. The at least one word information can also be understood as at least one first word information. The further word information can be understood as second word information. In the present case, the emotion information is transcribed into the second word information.provided the emotional information is transcribed into word information. The term "information" is used synonymously with "signal" here. The emotional information comprises a suprasegmental feature. Preferably, the at least one piece of emotional information comprises one or more suprasegmental features. As suggested, the captured emotional information is either not reproduced at all, or it is reproduced together with the original word information as a first and second word. This allows a listener to understand the emotional information without difficulty, provided it is reproduced as additional word information. However, it is also conceivable that if the emotional information does not contribute significantly to the overall information,The process involves subtracting the emotional information from the speech signal and reproducing only the original (first) word information. The analysis device could also be described as a recognition system, since it is designed to identify what portion of the captured speech signal represents word information and what portion represents emotional information. Furthermore, the analysis device can be designed to identify different speakers. In this context, a de-emotionalized speech signal refers to one that is completely or partially free of emotion. A de-emotionalized speech signal thus contains only, in particular, first and / or second word information.where one or more word pieces of information can be based on emotional information. For example, a complete removal of emotion can occur in speech synthesis with a robot voice. It is also conceivable, for instance, to generate an angry robot voice. A partial reduction of emotions in the speech signal could be achieved through direct manipulation of the speech audio material, e.g., by reducing the dynamic range, reducing or limiting the fundamental frequency, changing the speech rate, altering the spectral content of the speech, and / or changing the prosody of the speech signal, etc.
[0011] The speech signal can also originate from an audio stream (audio data stream), e.g., television, radio, podcast, audiobook. A speech signal capture device in the narrower sense could be understood as a "microphone." Furthermore, the speech signal capture device could be understood as a device that enables the use of a general speech signal from, for example, the sources just mentioned.
[0012] A technical implementation of the proposed speech signal processing device is based on an analysis of the input language (speech signal) by an analysis device, such as a recognition system (e.g., neural network, artificial intelligence, etc.), which has either learned the conversion to the target signal based on training data (end-to-end transcription) or rule-based transcription based on recognized emotions, which in turn may have been trained inter-individually or intra-individually by a recognition system.
[0013] Two or more speech signal processing devices form a speech signal playback system. With a speech signal playback system, for example, two or more listeners can be provided with individually tailored, de-emotionalized speech signals in real time from a speaker who is emitting speech signals. An example of this is a lesson in a school or a guided tour of a museum, etc.
[0014] Another aspect of the present invention relates to a method for outputting a de-emotionalized speech signal in real time or after a period of time. The method comprises capturing a speech signal that includes at least one word piece of information and at least one emotion piece of information. A speech signal could, for example, be given by a speaker to a group of listeners in real time. The method further comprises analyzing the speech signal with respect to the at least one word piece of information and the at least one emotion piece of information. The speech signal is to be recognized with respect to its word piece and its emotion piece. The at least one emotion piece of information comprises at least one suprasegmental feature, which is to be transcribed into a further, in particular a second, word piece of information.Consequently, the method comprises separating the speech signal into at least one word information and at least one emotion information, and processing the speech signal, wherein the at least one emotion information is transcribed into another word information either on the basis of training data or on the basis of rule-based transcription of recognized emotions, which in turn have been trained inter-individually or intra-individually on the analysis device, and reproducing the speech signal as a de-emotionalized speech signal, which includes the at least one emotion information converted into another word information and / or which includes the at least one word information.
[0015] For reasons of redundancy, the definitions of terms used in connection with the speech signal processing device are not repeated. However, it is understood that these definitions also apply analogously to the procedure and vice versa.
[0016] The core of the technical approach described here is the recognition of information contained in the social media (including emotions) and its insertion into the output signal in a spoken, written, or pictorial form. For example, a speaker who says, very excitedly, "It's outrageous that they're denying me entry here," could be transcribed as "I'm very excited because it's outrageous that..."
[0017] One advantage of the technical teaching disclosed herein is that the speech signal processing device / method is individually tailored to a user by identifying particularly disruptive elements of speech processing for that user. This can be especially important for people with autism, as the individual manifestation and sensitivity to speech processing can vary greatly. Determining individual sensitivity can be achieved, for example, via a user interface through direct feedback or input from close relatives (e.g., parents) or through neurophysiological measurements such as heart rate variability (HRV) or EEG. Neurophysiological measurements have been identified in scientific studies as markers for the perception of stress, exertion, or positive / negative emotions induced by acoustic signals and can thus be used in conjunction with the aforementionedRecognition systems are fundamentally used to determine correlations between speech-related noise (SM) and individual impairment. Once such a correlation has been established, the speech signal processing device / method can reduce or suppress the corresponding, particularly disruptive SM components, while other SM components are not processed or are processed differently.
[0018] Provided that the language is not directly manipulated, but rather generated "artificially" (i.e., in an end-to-end process) without SM, the tolerated SM components can be added to this SM-free signal based on the same information, and / or special SM components can be generated that can support understanding.
[0019] A further advantage of the technical teaching disclosed herein is that, in addition to modifying the SM components, the playback of the de-emotionalized signal can be adapted to the listener's auditory needs. For example, it is known that people with autism have particular requirements for good speech intelligibility and are especially easily distracted from the speech information by background noise in the recording. This can be mitigated by noise reduction, which may be individualized in its degree. Similarly, individual hearing impairments can be compensated for during the processing of speech signals (e.g., by non-linear, frequency-dependent amplification, as used in hearing aids), or the speech signal, reduced by SM components, can be further processed using a generalized, non-individualized method that, for example, increases voice clarity or suppresses background noise.
[0020] A particular potential of this technical teaching lies in its application to communication with autistic individuals and non-native speakers. The approach of automated transcription, as described here, of a speech signal into a new signal—one free of or modified in its SM components, and / or one that accurately reflects the information contained within the SM components—and which is also capable of real-time communication, facilitates and improves communication with autistic individuals and / or non-native speakers, or with individuals in emotionally charged communication scenarios (fire department, military, alarm activation, etc.), or with individuals with cognitive impairments.
[0021] Advantageous embodiments of the present invention are the subject of the dependent claims. Preferred embodiments of the present invention are explained in detail below with reference to the accompanying drawings. These show: Fig. 1 a schematic representation of a speech signal processing device; Fig. 2 a schematic representation of a speech signal playback system; and Fig. 3 a flowchart of a proposed method.
[0022] Individual aspects of the invention described herein are set forth below. Figures 1 to 3 described. In summary, the Figures 1 to 3 The principle of the present invention is clarified. In the present application, identical reference numerals refer to identical or equivalent elements, and not all reference numerals are repeated in all drawings.
[0023] All definitions provided in this application are applicable to the proposed speech signal processing device, the speech signal reproduction system, and the proposed method. The definitions are not repeated repeatedly to avoid redundancy.
[0024] Fig. 1Figure 1 shows a speech signal processing device 100 for outputting a de-emotionalized speech signal 120 in real time or after a period of time. The speech signal processing device 100 comprises a speech signal acquisition device 10 for acquiring a speech signal 110, which includes at least one emotion information 12 and at least one word information 14. The speech signal processing device 100 also comprises an analysis device 20 for analyzing the speech signal 110 with respect to the at least one emotion information 12 and the at least one word information 14. The analysis device 20 could also be referred to as a recognition system, since the analysis device is configured to recognize at least one emotion information 12 and at least one word information 14 of a speech signal 110, provided that the speech signal 110 includes at least one emotion information 12 and at least one word information 14.The speech signal processing device 100 further comprises a processing device 30 for separating the speech signal 110 into at least one word information 14 and at least one emotion information 12, and for processing the speech signal 110. During processing of the speech signal 110, the emotion information 12 is transcribed into a further, in particular a second, word information 14'. An emotion information 12 is, for example, a suprasegmental feature. The speech signal processing device 100 also comprises a coupling device 40 and / or a playback device 50 for reproducing the speech signal 110 as a de-emotionalized speech signal 120, which has converted the at least one emotion information 12 into a further word information 14' and / or includes the at least one word information 14. The emotion information 12 can thus be reproduced as a further word information 14' in real time for a user.This can compensate for, and especially prevent, comprehension problems on the part of the user.
[0025] The proposed speech signal processing device 100 can advantageously facilitate communication in learning a foreign language, understanding a dialect, or for cognitively impaired people.
[0026] Preferably, the speech signal processing device 100 comprises a storage device 60, which stores the de-emotionalized speech signal 120 and / or the captured speech signal 110 in order to reproduce the de-emotionalized speech signal 120 at any given time, in particular to reproduce the stored speech signal 110 as the de-emotionalized speech signal 120 more than once at any given time. The storage device 60 is optional. Both the original speech signal 110, which was captured, and the already de-emotionalized speech signal 120 can be stored in the storage device 60. This allows the de-emotionalized speech signal 120 to be reproduced repeatedly, in particular played back. A user can therefore first have the de-emotionalized speech signal 120 played back to them in real time and then have the de-emotionalized speech signal 120 played back again at a later time.For example, a user could be a student in school who listens to the de-emotionalized speech signal 120 in situ. When reviewing the learning material outside of school, i.e., at a later time, the student could listen to the de-emotionalized speech signal 120 again if needed. In this way, the speech signal processing device 100 can support the user's learning success.
[0027] Saving the speech signal 110 is equivalent to saving a captured original signal. The saved original signal can then be played back later in an unemotionalized state. The unemotionalization can be performed at a later time and then played back in real time. This allows for the analysis and editing of speech signal 110 at a later time.
[0028] Furthermore, it is conceivable to de-emotionalize the speech signal 110 and store it as a de-emotionalized signal 120. The stored de-emotionalized signal 120 can then be played back at a later time, especially repeatedly.
[0029] Depending on the storage capacity of the storage device 60, it is also conceivable to store the speech signal 110 and the corresponding de-emotionalized signal 120 in order to play both signals back at a later time. This can be useful, for example, if individual user settings need to be changed after the speech signal has been recorded and subsequently stored as de-emotionalized signal 120. It is conceivable that a user might not be satisfied with the real-time de-emotionalized speech signal 120, so that post-processing of the de-emotionalized signal 120 by the user or another person would be advisable, allowing future recorded speech signals to be de-emotionalized taking the post-processing of the de-emotionalized signal 120 into account. In this way, the de-emotionalization of a speech signal 120 can be subsequently adapted to the individual needs of a user.
[0030] Preferably, the processing device 30 is configured to recognize speech information 14, which is included in the emotion information 12, and to translate it into a de-emotionalized speech signal 120 and to pass it on to the playback device 50 for playback or to the coupling device 40, which is configured to connect to an external playback device (not shown), in particular a smartphone or tablet, in order to transmit the de-emotionalized signal 120 for playback.It is therefore conceivable that one and the same speech signal playback device 100 reproduces the de-emotionalized signal 120 by means of an integrated playback device 50 of the speech signal playback device 100, or that the speech signal playback device 100 sends the de-emotionalized signal 120 to an external playback device 50 by means of a coupling device 40 in order to reproduce the de-emotionalized signal 120 at the external playback device 50. When transmitting the de-emotionalized signal 120 to an external playback device 50, it is possible to send the de-emotionalized signal 120 to a plurality of external playback devices 50.
[0031] Furthermore, it is conceivable that the speech signal 110 is sent via the coupling device to a number of external speech signal processing devices 100, with each speech signal processing device 100 then defecating the received speech signal 100 according to the individual needs of the respective user of the speech signal processing device 100 and reproducing it as a defecated speech signal 120 to the corresponding user. In this way, for example, a defecated signal 120 can be reproduced for each student in a school class, tailored to their individual needs. This can improve the learning success of a class of students by addressing their individual needs.
[0032] Preferably, the analysis device 20 is configured to analyze background noise and / or emotional information 12 in the speech signal 110, and the processing device 30 is configured to remove the analyzed background noise and / or emotional information 12 from the speech signal 110. As, for example, in Fig. 1The analysis device 20 and the processing device 30 can be two different devices. However, it is also conceivable that the analysis device 20 and the processing device 30 are represented by a single device. The user, who uses the speech signal processing device 100 to have a de-emotionalized speech signal 120 played back, can mark a noise as background noise according to their individual needs, which can then be automatically removed by the speech signal processing device 100. Furthermore, the processing device can remove emotional information 12, which does not make a significant contribution to word information 14, from the speech signal 110. The user can mark emotional information 12, which does not make a significant contribution to word information 14, as such according to their needs.
[0033] As in Fig. 1As shown, the various devices 10, 20, 30, 40, 50, 60 can be in communicative exchange (see dashed arrows). Any other meaningful communicative exchange between the various devices 10, 20, 30, 40, 50, 60 is also conceivable.
[0034] Preferably, the playback device 50 is configured to reproduce the de-emotionalized speech signal 120 without the emotion information 12, or with the emotion information 12 transcribed into further word information 14', and / or with newly imprinted emotion information 12'. The user can decide or mark, according to their individual needs, which type of imprinted emotion information 12' can promote their understanding of the de-emotionalized signal 120 when it is reproduced. Furthermore, the user can decide or mark which type of emotion information 12 is to be removed from the de-emotionalized signal 120. This can also promote the user's understanding of the de-emotionalized signal 120.Furthermore, the user can decide or mark which type of emotional information 12 should be transcribed as further, in particular second, word information 14' in order to include it in the deemotionalized signal 120. The user can thus influence the deemotionalized signal 120 according to their individual needs so that the deemotionalized signal 120 is maximally understandable for the user.
[0035] Preferably, the playback device 50 comprises a loudspeaker and / or a screen to reproduce the de-emotionalized speech signal 120, particularly in simplified language, by means of an artificial voice and / or by displaying computer-generated text and / or by generating and displaying picture card symbols and / or by animating sign language. The playback device 50 can be configured in any way preferred by the user. The de-emotionalized speech signal 120 can be reproduced on the playback device 50 in such a way that the user understands the de-emotionalized speech signal 120 to the best of their ability. For example, it would also be conceivable to translate the de-emotionalized speech signal 120 into a foreign language that is the user's native language. Furthermore, the de-emotionalized speech signal can be reproduced in simplified language, which can improve the user's understanding of the speech signal 120.
[0036] An example of how simplified language is generated from emotionally charged speech material, utilizing emotions as well as intonations (or stresses in pronunciation) to represent them in simplified language, is the following: If someone speaks very angrily and utters the speech signal 110 "You mustn't say that to me," the speech signal 110 would be replaced, for example, by the following de-emotionalized speech signal 120: "I am very angry, because you mustn't say something like that to me." In this case, the processing device 30 would transcribe the emotional information 12, indicating that the speaker is "very angry," into the further word information 14, "I am angry."
[0037] Preferably, the processing device 30 comprises a neural network trained to transcribe the emotion information 12 into further word information 14' based on training data or rule-based transcription. One possible application of a neural network would be end-to-end transcription. In rule-based transcription, for example, the contents of a lexicon can be accessed. When using artificial intelligence, this base of training data, which is provided by the user, can learn the user's needs.
[0038] Preferably, the speech signal processing device 100 is configured to use first and / or second context information to determine a current location coordinate of the speech signal processing device 100 based on the first context information and / or to set associated transcription presets on the speech signal processing device 100 based on the second context information. The speech signal processing device 100 can include a GPS unit (not shown in the figures) and / or a speaker recognition system, which are configured to determine a current location coordinate of the speech signal processing device 100 and / or to recognize the speaker uttering the speech signal 110 and to set associated transcription presets on the speech signal processing device 100 based on the determined current location coordinate and / or speaker information.The first contextual information can include capturing the current location coordinates of the speech signal processing device 100. The second contextual information can include identifying a speaker. This second contextual information can be captured using the speaker recognition system. After identifying a speaker, the processing of the speech signal 110 can be adapted to the identified speaker; in particular, presets associated with the identified speaker can be set to process the speech signal 110. These presets can include, for example, assigning different voices to different speakers in the case of speech synthesis, or very strong de-emotionalization of the speech signal 110 at school but less strong de-emotionalization at home.The speech signal processing device 100 can thus utilize additional, in particular primary and / or secondary, contextual information when processing speech signals 110. This includes, for example, positional data such as GPS, which indicates the current location, or a speaker recognition system that identifies a speaker and adapts the processing depending on the speech. It is conceivable that, upon identification of different speakers, the speech signal processing device 100 could assign different voices to them. This can be advantageous in the case of speech synthesis or when a high degree of emotional de-emotion is required in schools, especially due to the prevailing background noise from other students. In a home environment, however, less emotional de-emotion of the speech signal 110 may be necessary.
[0039] In particular, the speech signal processing device 100 includes a signal exchange device (only indicated in Fig. 2 (indicated by the dashed arrows), which is configured to transmit a signal from a captured speech signal 110 to one or more other speech signal processing devices 100-1 to 100-6, in particular by means of radio, or Bluetooth, or LiFi (Light Fidelity). The signal transmission can be traced from point to multipoint (see Fig. 2 Each of the speech signal processing devices 100-1 to 100-6 can then reproduce a de-emotionalized signal 120-1, 120-2, 120-3, 120-4, 120-5, 120-6, adapted in particular to the needs of the respective user. In other words, one and the same captured speech signal 110 can be transcribed into a different de-emotionalized signal 120-1 to 120-6 by each of the speech signal processing devices 100-1 to 100-6. Fig. 2The transmission of the speech signal 110 is shown unidirectionally. Such unidirectional transmission of the speech signal 110 is suitable, for example, in schools. It is also conceivable that speech signals 110 can be transmitted bidirectionally between several speech signal processing devices 100-1 to 100-6. This can, for example, facilitate communication between users of the speech signal processing devices 100-1 to 100-6.
[0040] Preferably, the speech signal processing device 100 has a user interface 70 configured to categorize the at least one emotion information 12 into undesirable emotion information, neutral emotion information, and / or positive emotion information, according to user-defined preferences. The user interface is preferably communicatively connected to each of the devices 10, 20, 30, 40, 50, and 60. This allows the user to control each of the devices 10, 20, 30, 40, 50, and 60 via the user interface 70 and, if necessary, to input user input.
[0041] For exampleThe speech signal processing device 100 can be configured to categorize the at least one detected emotional information 12 into classes of different interference quality, in particular which have, for example, the following assignment: Class 1 "very disturbing", Class 2 "disturbing", Class 3 "less disturbing" and Class 4 "not disturbing at all"; and to reduce or suppress the at least one detected emotional information 12 that has been categorized into one of Class 1 "very disturbing" or Class 2 "disturbing" and / or to add the at least one detected emotional information 12 that has been categorized into one of Class 3 "less disturbing" or Class 4 "not disturbing at all" to the de-emotionalized speech signal 120 and / or to add a generated emotional information 12' to the de-emotionalized signal 120 in order to support a user's understanding of the de-emotionalized speech signal 120.Other forms of data collection are also conceivable. The example here is merely intended to indicate one way in which emotion information 12 could be classified. Furthermore, it should be noted that in this case, a generated emotion information 12' corresponds to an imprinted emotion information 12'. It is also conceivable to categorize the collected emotion information 12 into more or fewer than four classes.
[0042] Preferably, the speech signal processing device 100 includes a sensor 80 which, upon contact with a user, is configured to identify unwanted and / or neutral and / or positive emotional signals for the user. In particular, the sensor 80 is configured to measure biosignals, such as perform a neurophysiological measurement, or to capture and evaluate an image of a user. The sensor can be a camera or video system used to capture the user in order to analyze their facial expressions in relation to a speech signal 110 perceived by the user. The sensor can be understood as a neuro-interface. In particular, the sensor 80 is configured to measure blood pressure, skin conductance, or the like.In particular, the user can actively mark unwanted emotional information 12, for example, if the sensor 80 detects elevated blood pressure in the user in response to unwanted emotional information 12. The sensor 12 can also detect positive emotional information 12 for the user, specifically if the blood pressure measured by the sensor 80 does not change in response to the emotional information 12. Information about positive or neutral emotional information 12 may be important input parameters for processing the speech signal 110, for training the analysis device 20, for synthesizing the de-emotionalized speech signal 120, etc.
[0043] Preferably, the speech signal processing device 100 comprises a compensation device 90, which is designed to compensate for an individual hearing impairment associated with the user by, in particular, non-linear and / or frequency-dependent amplification of the de-emotionalized speech signal 120. By means of such, in particular, non-linear and / or frequency-dependent amplification of the de-emotionalized speech signal 120, the de-emotionalized speech signal 120 can be acoustically reproduced for the user despite an individual hearing impairment.
[0044] Fig. 2Figure 1 shows a speech signal reproduction system 200, which comprises two or more speech signal processing devices 100-1 to 100-6, as described herein. Such a speech signal reproduction system 200 could, for example, be used in a school lesson. For example, a teacher could speak into the speech signal processing device 100-1, which captures the speech signal 110. Via the coupling device 40 (see Figure 100-1), the speech signal 110 is transmitted to the processing device 100-1. Fig. 1The speech signal processing devices 100-1 could then establish a connection to, in particular to, the respective coupling device, the speech signal processing devices 100-2 to 100-6, which simultaneously transmits the captured speech signal(s) 110 to the speech signal processing devices 100-2 to 100-6. Each of the speech signal processing devices 100-2 to 100-6 can then analyze the received speech signal 110, as already described, and transcribe it into a user-specific, de-emotionalized signal 120 and play it back to the users. Speech signal transmission from one speech signal processing device 100-1 to another speech signal processing device 100-2 to 100-6 can take place via radio, Bluetooth, LiFi, etc.
[0045] Fig. 3A method 300 for outputting a de-emotionalized speech signal 120 in real time or after a time interval is shown. The method 300 initially comprises a step 310 of capturing a speech signal 110, which includes at least one word information 14 and at least one emotion information 12. The emotion information 12 includes at least one suprasegmental feature, which can be transcribed into another word information 14' or which can be subtracted from the speech signal 110. In each case, a de-emotionalized signal 120 results. The speech signal to be captured can be spoken language of a person in situ or can be generated by a media file, by radio, or by a video being played.
[0046] In the following step 320, the speech signal 110 is analyzed with regard to the at least one word information 14 and with regard to the at least one emotion information 12. For this purpose, an analysis device 30 is designed to recognize which speech signal component of the recorded speech signal 110 is to be assigned to a word information 14 and which speech signal component of the recorded speech signal 110 is to be assigned to an emotion information 12, in particular an emotion.
[0047] Following step 320 of the analysis, step 330 then takes place. Step 330 comprises separating the speech signal 110 into at least one word information 14 and at least one emotion information 14, and processing the speech signal 110. A processing device 40 may be provided for this purpose. The processing device may be integrated into the analysis device 30 or may be a separate device independent of the analysis device 30. In any case, the analysis device 30 and the processing device are coupled together so that, after the analysis of the word information 14 and the emotion information 12, these two pieces of information 12 and 14 are separated into two signals. The processing device is further configured to translate or transcribe the emotion information 12, i.e., the emotion signal, into another word information 14.Furthermore, the processing device 40 is configured to alternatively remove the emotion information 12 from the speech signal 110. In any case, the processing device is configured to create a de-emotionalized speech signal 120 from the speech signal 110, which is a sum or superposition of the word information 14 and the emotion information 12. The de-emotionalized speech signal 120 preferably comprises only, in particular the first and second, word information 14, 14' or word information 14, 14' and one or more emotion information pieces 12, which have been classified by a user as permissible, in particular acceptable or non-disruptive.
[0048] Finally, step 340 takes place, after which the speech signal 110 is reproduced as a deemotionalized speech signal 120, which includes at least one emotion information 12 converted into another word information 14' and which includes at least one word information 14.
[0049] With the proposed method 300 or with the proposed speech signal playback device 100, an in situ captured speech signal 110 can be reproduced to the user as a de-emotionalized speech signal 120 in real time or after a time interval, i.e. at a later time, which has the effect that the user, who might otherwise have problems understanding the speech signal 110, can understand the de-emotionalized speech signal 120 essentially without problems.
[0050] Preferably, the method 300 comprises storing the de-emotionalized speech signal 120 and / or the captured speech signal 110; and replaying the de-emotionalized speech signal 120 and / or the captured speech signal 120 at any given time. For example, when storing the speech signal 110, the word information 14 and the emotion information 12 are stored, while when storing the de-emotionalized speech signal 120, for example, the word information 14 and the emotion signal 12 transcribed into another word information 14' are stored. For example, a user or another person can have the speech signal 110 played back, in particular listen to it, and compare it with the de-emotionalized speech signal 120.In the event that the emotion information has not been transcribed quite accurately into further word information 14', the user or another person can change the transcribed further word information 14', in particular correct it. When using artificial intelligence (AI), the correct transcribing of emotion information 12 into further word information 14' for a user can be learned. For example, an AI can also learn which emotion information 12 does not bother the user, or even affects them positively, or appears neutral to a user.
[0051] Preferably, the method 300 comprises recognizing the at least one emotion information 12 in the speech signal 110; and analyzing the at least one emotion information 12 with respect to possible transcriptions of the at least one emotion signal 12 into n different further, in particular second, word information 14', where n is a natural number greater than or equal to 1 and n represents the number of ways to correctly transcribe the at least one emotion information 12 into the at least one further word information 14'; and transcribing the at least one emotion information 12 into the n different further word information 14'. For example, a content-changing SM in a speech signal 110 can be transcribed into n differently modified contents. Example: The sentence "Are you going to Oldenburg today?" can be understood differently depending on the intonation.If the emphasis is on "Fährst" (you are going), an answer like "no, I'm flying to Oldenburg" would be expected; if, on the other hand, the emphasis is on "Du" (you), an answer like "no, I'm not going to Oldenburg, a colleague is" would be expected. A transcription in the first case could be "You will be in Oldenburg today, will you be going there?". Depending on the emphasis, different second-word information can thus result from a single piece of emotional information.
[0052] Preferably, the method 300 comprises the identification of unwanted and / or neutral and / or positive emotional information by a user via an operator interface 70. The user of the speech signal processing device 100 can, for example, define via the operator interface 70 which emotional information 12 he perceives as disturbing, neutral, or positive. For example, emotional information 12 perceived as disturbing can then be treated as emotional information that needs to be transcribed, while emotional information 12 perceived as positive or neutral may remain unchanged in the de-emotionalized speech signal 120.
[0053] Preferably, the method 300 further or alternatively comprises identifying unwanted and / or neutral and / or positive emotional information 12 by means of a sensor 80, which is configured to perform a neurophysiological measurement. The sensor 80 can thus be a neurointerface. A neurointerface is mentioned only as an example. It is also conceivable to provide other sensors. For example, one or more sensors 80 could be provided, which are configured to record different measured variables, in particular the blood pressure, heart rate and / or skin conductance of the user.
[0054] The procedure 300 may include categorizing the at least one captured emotional information 12 into classes of varying levels of disturbance, in particular where the classes may have, for example, the following assignments: Class 1 "very disturbing", Class 2 "disturbing", Class 3 "less disturbing", and Class 4 "not disturbing at all". Furthermore, the procedure may include reducing or suppressing the at least one captured emotional information 12 that has been categorized into Class 1 "very disturbing" or Class 2 "disturbing", and / or adding the at least one captured emotional information 12 that has been categorized into Class 3 "less disturbing" or Class 4 "not disturbing at all" to the de-emotionalized speech signal, and / or adding generated emotional information 12' to support a user's understanding of the de-emotionalized speech signal 120.A user can therefore tailor the 300 procedure to their individual needs.
[0055] Preferably, the method 300 comprises reproducing the de-emotionalized speech signal 120, particularly in simplified language, by an artificial voice and / or by displaying computer-generated text and / or by generating and displaying picture card symbols and / or by animating sign language. In the present case, the de-emotionalized speech signal 120 can be reproduced for the user in a way that is adapted to their individual needs. The above list is not exhaustive. Rather, other methods of reproduction are also possible. This allows the speech signal 110 to be reproduced in simplified language, particularly in real time. The captured, in particular recorded, speech signals 100 can, after transcription into a de-emotionalized signal 120, be replaced, for example, by an artificial voice that contains no or reduced SM components.which no longer contains the SM elements identified as particularly disruptive for the individual. For example, the same voice could always be used to communicate with a person with autism, even if it comes from different conversation partners (e.g., different teachers), if this meets the individual communication needs.
[0056] For example, the sentence "In an ever-increasing commotion, people held up signs saying 'No violence,' but the police officers beat them with batons." could be transcribed into simplified language as follows: "The commotion grew ever larger. People held up signs saying 'No violence.' Police officers beat them with batons."
[0057] Preferably, the method 300 comprises compensating for an individual hearing impairment associated with the user, in particular by non-linear and / or frequency-dependent amplification of the de-emotionalized speech signal 120. This allows the user to be offered a listening experience, provided that the de-emotionalized speech signal 120 is acoustically reproduced to the user, which is similar to the listening experience of a user who does not have a hearing impairment.
[0058] Preferably, the method 300 includes analyzing whether a background noise, particularly one individually defined by a user, is detected in the captured speech signal 110, and, if necessary, subsequently removing the detected background noise. Background noise can be, for example, noise from a barking dog, other people, or traffic noise, etc. If the speech or background noise is not directly manipulated, but rather the speech is generated "artificially" (i.e., in an end-to-end process), then interfering noises are automatically removed in this process. Aspects such as an individual or subjective improvement in clarity, pleasantness, or familiarity can be considered during training or subsequently added.
[0059] Preferably, the method 300 comprises detecting a current location coordinate using GPS, followed by setting presets associated with the detected location coordinate for the transcription of speech signals 110 recorded at the current location coordinate. Because a current location can be detected, such as the school, one's own home, or a supermarket, presets associated with the respective location and relating to the transcription of recorded speech signals 110 at the respective location can be automatically changed or adjusted.
[0060] Preferably, the method 300 comprises transmitting a captured speech signal 110 from one speech signal processing device 100, 100-1 to another speech signal processing device 100 or to several speech signal processing devices 100, 100-2 to 100-6 by means of radio, Bluetooth, or LiFi (Light Fidelity). When using LiFi, signals could be transmitted in a direct or indirect line of sight. For example, a speech signal processing device 100 could send a speech signal 110, particularly optically, to a control interface, to which the speech signal 110 is routed to various outputs and distributed to the various speech signal processing devices 100, 100-2 to 100-6. Each of the various outputs can be communicatively coupled to a speech signal processing device 100.
[0061] Another aspect of the present application concerns a computer-readable storage medium comprising instructions which, when executed by a computer, in particular a speech signal processing device 100, cause it to execute the method as described herein. In particular, the computer, in particular the speech signal processing device 100, may be a smartphone, a tablet, a smartwatch, etc.
[0062] Although some aspects relating to a device have been described, it is understood that these aspects also constitute a description of a corresponding method, such that a block or component of a device can also be understood as a corresponding method step or as a feature of a method step. For reasons of redundancy, a description of the present invention in the form of method steps is omitted here. Some or all of the method steps could be carried out by (or using) a hardware apparatus, such as a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, some or more of the main method steps can be carried out by such an apparatus.
[0063] In the preceding detailed description, various features were sometimes grouped together in examples to streamline the disclosure. This type of disclosure should not be interpreted as indicating that the claimed examples have more features than are expressly stated in each claim. Rather, as the following claims reflect, the subject matter may consist of fewer than all the features of a single disclosed example. Consequently, the following claims are hereby incorporated into the detailed description, with each claim potentially representing a separate example.While each claim can stand as a separate example, it should be noted that, although dependent claims refer back to a specific combination with one or more other claims, other examples also include a combination of dependent claims with the subject matter of any other dependent claim, or a combination of any feature with other dependent or independent claims. Such combinations are included unless it is stated that a specific combination is not intended. Furthermore, it is intended that a combination of features of a claim with any other independent claim is also included, even if that claim is not directly dependent on the independent claim.
[0064] Depending on specific implementation requirements, embodiments of the invention can be implemented in hardware or in software, or at least partially in hardware or at least partially in software. The implementation can be carried out using a digital storage medium, for example, a floppy disk, a DVD, a Blu-ray disc, a CD, a ROM, a PROM, an EPROM, an EEPROM, a FLASH memory, a hard disk, or another magnetic or optical storage medium, on which electronically readable control signals are stored. These control signals can interact with, or interact with, a programmable computer system in such a way as to execute the respective method. Therefore, the digital storage medium enabling the proposed teaching can be computer-readable.
[0065] Some embodiments according to the teaching described herein therefore include a data carrier which has electronically readable control signals which are able to interact with a programmable computer system in such a way that one of the features described herein is carried out as a method.
[0066] In general, embodiments of the teaching described herein can be implemented as a computer program product with a program code, wherein the program code is effective in carrying out one of the methods when the computer program product runs on a computer.
[0067] The program code can also be stored on a machine-readable medium, for example.
[0068] Other embodiments include the computer program for performing one of the features described herein as a method, wherein the computer program is stored on a machine-readable medium. In other words, an embodiment of the method according to the invention is thus a computer program that includes program code for performing one of the methods described herein when the computer program is executed on a computer.
[0069] Another embodiment of a proposed method is thus a data carrier (or a digital storage medium or a computer-readable medium) on which the computer program for performing one of the features described herein is recorded as a method. The data carrier, digital storage medium, or computer-readable medium is typically tangible and / or non-volatile.
[0070] Another embodiment of the proposed method is thus a data stream or a sequence of signals that represents the computer program for carrying out one of the methods described herein. The data stream or sequence of signals can, for example, be configured to be transferred via a data communication connection, such as the Internet.
[0071] Another embodiment comprises a processing device, for example a computer or a programmable logic device, configured or adapted to perform a method for the system described herein.
[0072] Another embodiment includes a computer on which the computer program for carrying out the method for the system described herein is installed.
[0073] Another embodiment of the invention comprises a device or system designed to transmit a computer program for performing at least one of the features described herein, in the form of a method, to a receiver. The transmission can be, for example, electronic or optical. The receiver can be, for example, a computer, a mobile device, a storage device, or a similar device. The device or system can, for example, include a file server for transmitting the computer program to the receiver.
[0074] In some embodiments, a programmable logic device (for example, a field-programmable gate array, an FPGA) can be used to perform some or all of the functionalities of the methods and devices described herein. In some embodiments, a field-programmable gate array can interact with a microprocessor to perform the method described herein. Generally, in some embodiments, the method is performed by any hardware device. This can be general-purpose hardware such as a computer processor (CPU) or method-specific hardware such as an ASIC.
[0075] The embodiments described above merely illustrate the principles of the present invention. It is understood that modifications and variations of the arrangements and details described herein will be obvious to other people skilled in the art. Therefore, it is intended that the invention be limited only by the scope of protection set forth in the following claims and not by the specific details presented herein by way of description and explanation of the embodiments.
Claims
1. Speech signal processing apparatus (100) for outputting a deemotionalized speech signal (120), the speech signal processing apparatus (100) including: - a speech signal acquisition apparatus (10) configured to acquire a speech signal (110) including at least one piece of emotion information (12) and at least one piece of word information (14); - an analysis apparatus (20), including a neural network or an artificial intelligence, configured to analyze the speech signal (110) with respect to the at least one piece of emotion information (12) and the at least one piece of word information (14), - a processing apparatus (30), including a neural network or an artificial intelligence, configured to separate the speech signal (110) into the at least one piece of word information (14) and into the at least one piece of emotion information (12) and to process the speech signal (110), wherein the at least one piece of emotion information (12) is transcribed into a further piece of word information either on the basis of training data or on the basis of rule-based transcription of recognized emotions that in turn have been trained inter-individually or intra-individually by the analysis apparatus (20); and - a coupling apparatus (40) and / or a reproduction apparatus (50) configured to reproduce the speech signal (110) as a deemotionalized speech signal (120) that includes the at least one piece of emotion information (12) converted into the further piece of word information (14') and that includes the at least one piece of word information (14).
2. Speech signal processing apparatus (100) according to claim 1, including a storage apparatus (60) that stores the deemotionalized speech signal (120) and / or the acquired speech signal (110) in order to reproduce the deemotionalized speech signal (120) at an arbitrary point in time, in particular to reproduce the stored speech signal (110) as a deemotionalized speech signal (120) at more than a single arbitrary point in time.
3. Speech signal processing apparatus (100) according to one of the preceding claims, wherein the processing apparatus (30) is configured to recognize the further piece of word information (14') included in the piece of emotion information (12) and to translate it into a deemotionalized speech signal (120) and to forward it to the reproduction apparatus (50) for reproduction by the reproduction apparatus (50) or to the coupling apparatus (40) configured to connect to an external reproduction apparatus (50), in particular a smartphone or tablet, to transmit the deemotionalized signal (120) for reproduction thereof.
4. Speech signal processing apparatus (100) according to one of the preceding claims, wherein the analysis apparatus (30) is configured to analyze interference noise and / or a piece of emotion information (12) in the speech signal (110), and the processing apparatus (40) is configured to remove the analyzed interference noise and / or the piece of emotion information (12) from the speech signal (110), or wherein the analysis apparatus (50) includes a neural network configured to transcribe the piece of emotion information (12) into the further piece of word information (14') on the basis of training data or on the basis of a rule-based transcription.
5. Speech signal processing apparatus (100) according to one of the preceding claims, wherein the reproduction apparatus (50) is configured to reproduce the deemotionalized speech signal (120) without the piece of emotion information (12) or with the piece of emotion information (12) transcribed into the further piece of word information (14') and / or with a newly applied piece of emotion information (12'), or, wherein the reproduction apparatus (50) includes a loudspeaker and / or a screen to reproduce the deemotionalized speech signal (120), in particular in simplified language, by an artificial voice and / or by displaying a computer-written text and / or by generating and displaying image map symbols and / or by animation of sign language.
6. Speech signal processing apparatus (100) according to one of the preceding claims, wherein the speech signal processing apparatus (100) includes a GPS unit and / or a speaker recognition system configured to detect a current location coordinate of the speech signal processing apparatus (100) and / or to recognize the speaker expressing the speech signal (110) and to set associated presettings for transcription at the speech signal processing apparatus (100) on the basis of the detected current location coordinate and / or speaker information, or includes signal exchange means configured to perform signal transmission of a detected speech signal with one or more other speech signal processing apparatuses (100, 100-1 to 100-6), in particular by means of radio, or Bluetooth, or LiFi (Light Fidelity).
7. Speech signal processing apparatus (100) according to one of the preceding claims, comprising an operating interface (70) configured to subdivide the at least one piece of emotion information (12) according to preferences set by a user into an undesired piece of emotion information and / or into a neutral piece of emotion information and / or into a positive piece of emotion information, or configured to categorize the at least one piece of detected emotion information (12) into classes of different interference quality, in particular having the following assignment, for example: class 1 "very interfering", class 2 "interfering", class 3 "less interfering" and class 4 "not interfering at all"; and to reduce or suppress the at least one piece of detected emotion information (12) categorized into one of classes 1 "very interfering" or class 2 "interfering" and / or to add the at least one piece of detected emotion information (12) categorized into one of classes 3 "less interfering" or class 4 "not interfering at all" to the deemotionalized speech signal (120) and / or to add a generated piece of emotion information (12') to the deemotionalized signal to support a user's understanding of the deemotionalized speech signal (120).
8. Speech signal processing apparatus (100) according to one of the preceding claims, comprising a sensor (80) configured, upon contact with a user, to identify undesired and / or neutral and / or positive pieces of emotion information (12) for the user either on the basis of training data or on the basis of rule-based transcription of recognized emotions that in turn have been trained inter-individually or intra-individually by the analysis apparatus (20), wherein the sensor (80) is configured to measure biosignals, such as to perform a neurophysiological measurement, or to acquire and evaluate an image of a user.
9. Speech signal processing apparatus (100) according to one of the preceding claims, comprising a compensation apparatus (90) configured to compensate for an individual hearing impairment associated with a user by non-linear and / or frequency-dependent amplification of the deemotionalized speech signal (120).
10. Speech signal reproduction system (200) comprising two or more speech signal processing apparatuses (100, 100-1 to 100-6) according to one of the preceding claims.
11. Method (300) for outputting a deemotionalized speech signal (120) in real time or after expiry of a time period, the method (300) comprising: acquiring (310) a speech signal (110) including at least one piece of word information (14) and at least one piece of emotion information (12); analyzing (320) the speech signal (110) with respect to the at least one piece of word information (14) and with respect to the at least one piece of emotion information (12); separating (330) the speech signal into the at least one piece of word information (14) and into the at least one piece of emotion information (12) and processing the speech signal (110), wherein the at least one piece of emotion information (12) is transcribed into a further piece of word information either on the basis of training data or on the basis of rule-based transcription of recognized emotions that in turn have been trained inter-individually or intra-individually by the analysis apparatus (20); reproducing (340) the speech signal (110) as a deemotionalized speech signal (120) that includes the at least one piece of emotion information (12) converted into a further piece of word information (14') and that includes the at least one piece of word information (14).
12. Method (300) according to claim 11, including: storing the deemotionalized speech signal (120) and / or the acquired speech signal (110); reproducing the deemotionalized speech signal (120) and / or the acquired speech signal (110) at an arbitrary point in time. or recognizing the at least one piece of emotion information (12) in the speech signal (110); analyzing the at least one piece of emotion information (12) with respect to possible transcriptions of the at least one piece of emotion information (12) into n different further pieces of word information (14'), wherein n is a natural number greater than or equal to 1 and n indicates a number of possibilities for transcribing the at least one piece of emotion information (12) correctly into the at least one further piece of word information (14'); transcribing the at least one piece of emotion information (12) into the n different further pieces of word information (14').
13. Method (300) according to one of claims 11 or 12, including: identifying undesired and / or neutral and / or positive pieces of emotion information (12) by a user by means of an operating interface (70), or identifying undesired and / or neutral and / or positive pieces of emotion information (12) by means of a sensor (80), in particular configured to measure biosignals, such as to perform a neurophysiological measurement, or to acquire and evaluate an image of a user, or including: categorizing the at least one piece of detected emotion information (12) into classes of different interference quality, in particular having at least one of the following assignments, for example: class 1 "very interfering", class 2 "interfering", class 3 "less interfering" and class 4 "not interfering at all"; reducing or suppressing the at least one piece of detected emotion information (12) categorized into one of classes 1 "very interfering" or class 2 "interfering" and / or adding the at least one piece of detected emotion information (12) categorized into one of classes 3 "less interfering" or class 4 "not interfering at all" to the deemotionalized speech signal and / or adding a generated piece of emotion information (12') to support a user's understanding of the deemotionalized speech signal (120).
14. Method (300) according to one of the preceding claims 11 to 13, including: reproducing the deemotionalized speech signal (120), in particular in simplified language, - by an artificial voice and / or - by displaying a computer-written text and / or - by generating and displaying image map symbols and / or - by animation of sign language, or compensating for an individual hearing impairment associated with a user by, in particular non-linear and / or frequency-dependent, amplification of the deemotionalized speech signal (120), or analyzing whether an interference noise, in particular individually defined by a user, is detected in the acquired speech signal (110), removing the detected interference noise, or detecting a current location coordinate by means of GPS, setting presettings associated with the detected location coordinate for transcription of speech signals (110) acquired at the current location coordinate, or transmitting an acquired speech signal (110) from one speech signal processing apparatus (100) to another speech signal processing apparatus (100) or to a plurality of speech signal processing apparatuses (100, 100-1 to 100-6) by means of GPS, or radio, or Bluetooth, or LiFi (Light Fidelity).
15. Computer-readable storage medium comprising instructions which, when executed by a computer, cause the latter to carry out the method (300) according to one of claims 11 to 14.