Device and method for automated utterance, sentence boundary and phrase recognition

WO2026201944A1PCT designated stage Publication Date: 2026-10-01FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2026/058207
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-05-09
Filing Date
2026-03-23
Publication Date
2026-10-01

Smart Images

  • Figure EP2026058207_01102026_PF_FP_ABST
    Figure EP2026058207_01102026_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a device for the speech analysis of a recorded first audio signal according to one embodiment, in which spoken language of a first person from a population is recorded, a population comprising a plurality of persons having the same or at least similar speech behaviour properties. The device comprises a speech recognition module (110) for speech recognition detection, which divides the recorded first audio signal into a plurality of speech segments that have speech and subdivide it into a plurality of speech pause segments that do not have speech. The device also comprises a speech analysis module (120), which is designed to determine a temporal pause length of one or more speech pause segments of the plurality of speech pause segments. The speech analysis module (120) is designed to determine utterance boundaries or sentence boundaries of one or more sentences or utterances according to the temporal pause length of the one or more speech pause segments in the recorded first audio signal. Furthermore, the speech analysis module (120) is designed to determine the utterance boundaries or the sentence boundaries of the one or more sentences or utterances according to a temporal duration of utterance boundaries or sentence boundaries which is typical of the population.
Need to check novelty before this filing date? Find Prior Art

Description

Device and method for automated Utterance, sentence boundary and phrase recognition Description The application relates to a device and a method for automated utterance, sentence boundary and phrase recognition. Speech recognition systems and context models, or models based on semantic rule systems such as NLP models (Natural Language Processing models) or Natural Language Understanding (NLU models), are trained on adult language and are generally unsuitable for child language, such as that spoken by children under five years old. Models for adult language (for natural language recognition or the recognition / analysis of text generated from this language) are based on probabilities in the use of grammatical relationships between words within a sentence, i.e., the relationship between word classes, the so-called parsing (in NLP), e.g., through...

[0023] However, developing children's language does not follow these probabilities from the outset, or rather, the rule system is not yet fully developed, so that relationships between sentence elements or word classes do not function reliably in children's language. These relationships are, however, necessary in existing models to determine sentence boundaries in adults and utterance boundaries in children based on syntax using NLP methods. With existing models for adult language, their application to toddlers fails because, according to the rules of language acquisition conventions, children, whose language does not yet have a main clause structure, only produce utterances and not sentences. Identifying the beginning and end of sentences or utterances in a continuous, recorded digital speech stream that is not limited to one or a few utterances remains an unresolved or insufficiently solved problem for automated speech analysis in children under five years of age. With current technology, the recognition of utterance boundaries based on recorded speech data cannot be automated. Furthermore, the division into utterances or sentences is important for many advanced language analyses, such as phrase recognition and the determination of means. FH250504PDE-2026097150.DOCXLength of Utterance (MLU; German: medium utterance length) etc. is important for a wide range of applications and is therefore helpful for further automated language analyses, e.g. with the help of NLP, NLU in the form of an attached word class recognition, e.g. by

[0023] , and subsequently verb position recognition, or forms the prerequisite for it. Phrases are propositions that can contain a maximum of one finite verb. If there are additional finite verbs, the utterance must be divided into several phrases. Phrases such as "I don't know." or "Nothing." (e.g., in response to the question "What is happening here?") are not considered phrases, even if they contain a finite verb. Furthermore, word repetitions, single words such as yes and no answers, self-corrections, interruptions, and incomprehensible utterances are also considered non-phrases. Certain parts of the phrases are also excluded, with the remainder of the phrase being evaluated.

[0025] The MLU can be calculated, for example, by distinguishing between phrases and non-phrases. Each utterance can be assigned a corresponding annotation with one of two labels (e.g., P for "phrase" or NP for "non-phrase"). The calculation of the Mean Length of Utterance refers to the phrases; non-phrases are not counted. Prior art includes Automatic Speech Recognition (ASR) systems, Speech-to-Text (S2Text) conversion, Voice Activity Detection (VAD), and Speech Activity Detection (SAD) (see, for example,

[0022] : International Telecommunication Union, Recommendation P.56: Objective measurement of active speech level, Dec. 2011, https: / / www.itu.int / rec / T-REC-P.56-201112-l / en, or other VAD and SAD implementations). Speech-to-Text or pure text analysis often utilizes NLP analysis systems, which, for example, operate on a punctuation-based basis, for further analysis.Here, text is divided into sentences or words at various processing stages

[0023] , words are assigned to word classes, and relations between the word classes are recognized using model-based parsing in order to draw conclusions about the grammatical relationship of the words within a sentence. Thus, it is possible, for example, to recognize nominal or verbal phrases.

[0024] This represents a subcategory of NLP, Natural Language Understanding (NLU). So far, NLP and NLU models have not been successfully applied to understanding child language. FH250504PDE-2026097150.DOCX. Furthermore, regarding the state of the art, Large Language Models (LLM) should be mentioned. Currently, sentence and utterance boundaries in children (over 5 years old) are identified by manually limiting the time frame of the data beforehand. This time frame is defined, for example, by the length of a sentence or by segmenting the utterance, or by using Speech-to-Text algorithms and corresponding punctuation marks or other (manual) pre-annotations or labels. Even within automatic transcription for adults using automatic speech recognition systems (LM, LLM, Kl), reliably matching the transcript to the acoustic speech output and subsequently subdividing the continuous speech stream into utterances or sentences is not currently possible.Problematic examples include single-word recognition based on large LLMs, the delimitation and identification of word repetitions, hesitations, sentence fragments based on the aforementioned systems, or the recognition of phrases and non-phrases in this speech stream to identify the actual syntactically evaluable sentence or utterance using such systems. Due to the lack of automatic sentence / utterance boundary detection, phrases within these cannot be automatically identified, which makes automatic determination of the MLU impossible. However, since determining the MLU is desirable as an important linguistic indicator of language proficiency, the MLU is currently determined manually through sentence / utterance and phrase transcription, as well as counting the word-to-phrase ratio (MLU = number of words per phrase / number of phrases). The object of the invention is achieved by the subject matter of the independent claims. Preferred embodiments are provided in the dependent claims. This makes it possible to perform further syntactic or semantic analyses within the recognized boundaries (utterance, sentence, phrase) using NLP / NLU, or, in combination with word boundary recognition, to determine the MLU fully automatically. A device for speech analysis of a recorded first audio signal according to an embodiment is provided in which spoken language of a first person from a population is recorded, wherein a population is a plurality of persons with the same or at least similar characteristics in speech behavior. FH250504PDE-2026097150.DOCX comprises a speech recognition module for speech detection, which divides the recorded initial audio signal into a plurality of speech segments containing speech and a plurality of speech pause segments containing no speech. Furthermore, the device comprises a speech analysis module configured to determine the temporal pause length of one or more speech pause segments within the plurality of speech pause segments. Depending on the temporal pause length of the one or more speech pause segments in the recorded initial audio signal, the speech analysis module is configured to determine utterance boundaries or sentence boundaries of one or more sentences or utterances.Furthermore, the language analysis module is designed to determine the utterance boundaries or sentence boundaries of one or more sentences or utterances, depending on a temporal duration of utterance boundaries or sentence boundaries that is typical for the population. Furthermore, a method for speech analysis of a recorded first audio signal, in which the spoken language of a first person from a population is recorded, is provided according to one embodiment, wherein a population comprises a plurality of persons with the same or at least similar characteristics in speech behavior. The method comprises: Dividing the recorded first audio signal into a plurality of speech segments that contain speech and into a plurality of speech pause segments that do not contain speech. Determining the length of a pause between one or more language pause segments or the majority of language pause segments. Depending on the temporal length of the one or more speech pause segments in the recorded first audio signal, utterance boundaries or sentence boundaries of one or more sentences or utterances are determined, whereby the utterance boundaries or sentence boundaries of the one or more sentences or utterances are determined depending on a temporal duration of utterance boundaries or sentence boundaries that is typical for the population. Furthermore, a computer program with program code for carrying out the procedure described above is provided when the computer program is executed on a computer or signal processor. FH250504PDE-2026097150.DOCX As described above, the application of existing models for adult language to toddlers fails because, according to the rules of language acquisition conventions, children, whose language does not yet exhibit a main clause structure, only produce utterances and not sentences. This is remedied by reliably enabling the recognition of utterances, sentence boundaries, and phrases in children by utilizing and evaluating the pause structure of child language. This makes it possible to conduct further syntactic or semantic analyses within the recognized boundaries (utterance, sentence, phrase) using NLP / NLU. Implementations combine digital speech processing and clinical language acquisition research. In specific implementations, the MLU can be determined automatically and efficiently within a spontaneous speech stream by using automatic sentence / utterance boundary detection combined with phrase recognition, and by determining the word count based on the word boundary-to-phrase ratio. (MLU = number of words (based on word boundaries, content recognition not necessary) per phrase / number of phrases). The use of NLP / NLU to determine word content and interpretation is also possible. For this extension, an ASR with VAD / SAD boundaries and word boundaries is used in the speech analysis module. Furthermore, an NLP / NLU model with word class recognition is required. Preferred embodiments of the invention are described below with reference to the drawings. The drawings illustrate: Fig. 1 shows a device for speech analysis of a recorded audio signal according to one embodiment. Fig. 2 shows a device for speech analysis of a recorded audio signal according to a further embodiment, whose speech analysis module has a sentence boundary detector and a phrase detector. Fig. 3 shows a device for speech analysis of a recorded audio signal according to a further embodiment, the FH250504PDE-2026097150.DOCX The language analysis module has an extended analysis module in addition to the sentence boundary recognizer and the phrase recognizer. Fig. 4 shows an exemplary representation of the analysis of different sentence parts and their word annotations (to clarify the reference to the utterance / sentence structure), depending on the strength of the filtering (thresholds), and their phrase (P) and non-phrase (NP) annotation according to one embodiment. Fig. 5 shows an exemplary representation of the pause duration based on a given population of detected word repetition, detected hesitation and so-called syntactic guardrails, or meaningful / nonsensical utterances according to one embodiment. Fig. 6 shows an exemplary representation of an automated recognition of different sentences or sentence parts depending on the strength of the filtering according to one embodiment. Fig. 7 shows an automation of utterance / sentence boundaries or phrase recognition with subsequent NLP word class recognition and further analyses according to one embodiment. Fig. 1 shows a device for speech analysis of a recorded audio signal according to one embodiment. The audio signal contains spoken language of a first person from a population, wherein a population comprises a plurality of persons with the same or at least similar characteristics in speech behavior. The device comprises a speech recognition module 110 for speech recognition detection, which divides the recorded first audio signal into a plurality of speech segments that contain speech and into a plurality of speech pause segments that do not contain speech. Furthermore, the device includes a speech analysis module 120, which is configured to determine the temporal pause length of one or more speech pause segments of the majority of speech pause segments. FH250504PDE-2026097150.DOCXThe speech analysis module 120 is designed to determine utterance boundaries or sentence boundaries of one or more sentences or utterances, depending on the temporal pause length of one or more speech pause segments in the recorded first audio signal. Furthermore, the language analysis module 120 is equipped to determine the utterance boundaries or sentence boundaries of one or more sentences or utterances, depending on a temporal duration of utterance boundaries or sentence boundaries that is typical for the population. Regarding the difference between the terms "utterance" and "sentence": Adults speak of sentences and sentence boundaries. Young children speak of utterances and utterance boundaries, since they usually do not yet express themselves in sentences. According to one embodiment, the typical temporal duration of utterance or sentence boundaries for the population can depend on the temporal duration of utterance or sentence boundaries in a plurality of recorded audio signals from a plurality of persons in the population. In one embodiment, the device can, for example, be designed to determine the typical temporal duration of the utterance boundaries or sentence boundaries for the population by analyzing the temporal duration of the utterance boundaries or sentence boundaries in the majority of the recorded audio signals of a majority of persons in the population. According to one embodiment, the device can, for example, be configured to determine which population the first person belongs to from a plurality of two or more populations. In such an embodiment, the device is therefore capable of performing speech analysis for different populations. In one embodiment, the speech analysis module 120 can be configured, for example, to determine the utterance boundaries or sentence boundaries of one or more sentences or utterances in such a way that, depending on which population from the majority of the two or more populations the first person belongs to, a typical temporal duration of utterance boundaries or sentence boundaries is selected for that population, and the utterance boundaries or sentence boundaries are adjusted according to this typical The temporal duration is determined using FH250504PDE-2026097150.DOCX. This results in a population-specific analysis of the utterance or sentence boundaries. According to one embodiment, if the first person belongs to a first population from the majority of two or more populations, the typical duration of utterance boundaries or sentence boundaries for that population can, for example, have a first duration. Conversely, if the first person belongs to a second population from the majority of two or more populations that differs from the first population, the typical duration of utterance boundaries or sentence boundaries for that population can, for example, have a second duration that differs from the first. Thus, utterance boundaries and sentence boundaries can differ depending on the population. In one embodiment, the device may, for example, be configured to provide an input interface through which a user can specify which population the first person belongs to from a plurality of two or more populations. For instance, in a particular embodiment, it is possible that a user specifies, via input devices such as a keyboard or mouse, or via voice input using a microphone, which population the aforementioned first person belongs to. According to one embodiment, the device can, for example, be configured to determine, by analyzing the recorded first audio signal, which population the first person belongs to from a plurality of two or more populations. In such an embodiment, the population to which the person belongs (for example, the population "children" or "adolescents" or "middle-aged adults" or "seniors") would be determined by analyzing the recorded audio signal. According to one embodiment, the speech analysis module 120 can be configured, for example, to determine that an utterance boundary or a sentence boundary exists in the recorded first audio signal if the temporal pause length of the speech pause segment is greater than or equal to a first temporal limit and less than or equal to a second temporal limit, wherein the first temporal limit and the second temporal limit depend on the temporal duration of the utterance boundaries or sentence boundaries that is typical for the population. FH250504PDE-2026097150.DOCXI In one embodiment, the speech analysis module 120 can be configured, for example, to determine a typical individual temporal duration of the utterance boundaries or sentence boundaries of the first person by individually adjusting the temporal duration of the utterance boundaries or sentence boundaries typical for the population for the first person, depending on the temporal pause length of the one or more speech pause segments in the recorded first audio signal; and / or wherein the speech analysis module 120 can be configured, for example, to adjust the first temporal limit and / or the second temporal limit depending on the temporal pause length of the one or more speech pause segments in the recorded first audio signal. According to one embodiment, the speech analysis module 120 can be configured, for example, to disregard, when determining the typical individual duration of the utterance boundaries or sentence boundaries of the first person or when adjusting the first limit and / or the second limit, such speech pause segments of the one or more speech pause segments whose duration falls short of the typical duration of the utterance boundaries or sentence boundaries for the population by more than one first deviation value or whose duration exceeds the typical duration of the utterance boundaries or sentence boundaries for the population by more than one second deviation value. In one embodiment, the speech analysis module 120 can, for example, be configured to identify an utterance in the recorded first audio signal by defining the beginning and end of the utterance through two utterance boundaries in the audio signal, between which there is no further utterance boundary. Alternatively, the speech analysis module 120 can, for example, be configured to identify a sentence in the recorded first audio signal by defining the beginning and end of the sentence through two sentence boundaries in the audio signal, between which there is no further sentence boundary. According to one embodiment, the speech analysis module 120 can be configured, for example, to determine phrase boundaries of a plurality of phrases depending on the temporal pause length of the one or more speech pause segments in the recorded first audio signal, wherein the speech analysis module 120 can be configured, for example, to determine the phrase boundaries of the one or more phrases depending on a temporal duration of phrase boundaries that is typical for the population. Regarding the concept of a phrase: A sentence or utterance can, for example, contain one or more phrases, such as a main clause and a subordinate clause, or FH250504PDE-2026097150.DOCX for example a phrase in which a meaningful statement is made, such as "I want to play" and a phrase without meaningful content, e.g. with "uh" sounds: e.g. "uh baggi". In one embodiment, the temporal duration of phrase boundaries typical for the population can depend on the temporal duration of phrase boundaries in a plurality of recorded audio signals from a plurality of persons in the population. According to one embodiment, the device can, for example, be designed to determine the typical temporal duration of phrase boundaries for the population by analyzing the temporal duration of phrase boundaries in the majority of recorded audio signals from a majority of persons in the population. In one embodiment, the device can, for example, be configured to determine a typical duration of utterance or sentence boundaries for each population of the majority of the two or more populations. In this way, utterance or sentence boundaries can be used, for example, by means of population-specific audio samples, to train population-specific AI subsystems (e.g., a first AI (sub)model for the children population, a second AI (sub)model for the adults population, each trained with audio recordings of the respective population). According to one embodiment, the device can, for example, be configured to determine the typical temporal duration of the phrase boundaries for each population of the plurality of the two or more populations by analyzing the temporal duration of the phrase boundaries in the plurality of the recorded audio signals of a plurality of persons in the population. In one embodiment, the typical temporal duration of utterance boundaries or sentence boundaries for a first population from the majority of the two or more populations can, for example, have a first temporal duration, and the typical temporal duration of utterance boundaries or sentence boundaries for a second population from the majority of the two or more populations can, for example, have a second temporal duration that differs from the first temporal duration. In one embodiment, the typical duration of phrase boundaries for the population can be shorter than the typical duration of utterance or sentence boundaries for the population. FH250504PDE-2026097150.DOCX According to one embodiment, the speech analysis module 120 can be configured, for example, to determine that a phrase boundary is present in the recorded first audio signal if the temporal pause length of the speech pause segment is greater than or equal to a third temporal limit and less than or equal to a fourth temporal limit, wherein the third temporal limit and the fourth temporal limit depend on the temporal duration of phrase boundaries typical for the population. In one embodiment, the speech analysis module 120 can, for example, be configured to determine a typical individual time duration of the phrase boundaries for the first person by individually adjusting the typical time duration of the phrase boundaries for the first person, depending on the time pause length of the one or more speech pause segments in the recorded first audio signal. And / or, the speech analysis module 120 can, for example, be configured to adjust the third time boundary and / or the fourth time boundary depending on the time pause length of the one or more speech pause segments in the recorded first audio signal. According to one embodiment, the speech analysis module 120 can, for example, be configured to identify a phrase in the recorded first audio signal by defining the beginning and end of the phrase through a first phrase boundary or a first utterance boundary and a second phrase boundary or a second utterance boundary in the audio signal, between which no further phrase boundary or utterance boundary lies. Alternatively, the speech analysis module 120 can, for example, be configured to identify a phrase in the recorded first audio signal by defining the beginning and end of the phrase through a first phrase boundary or a first sentence boundary and a second phrase boundary or a second sentence boundary in the audio signal, between which no further phrase boundary or sentence boundary lies. In one embodiment, the speech analysis module 120 can be configured to automatically determine and output a Mean Length of Utterance (MLU) in the audio signal, approximating the linguistic measure, using the temporal phrase length of the majority of phrases in the audio signal. FH250504PDE-2026097150.DOCX According to one embodiment, the device may further include a dialogue agent that adapts the pause lengths of its speech output to the speech pause segments of the majority of speech pause segments of a user in the audio signal. In one embodiment, the device can, for example, be configured to determine and output information about the quality of the first person's speech production based on the determination of the utterance limits or sentence boundaries. In one embodiment, the device can be configured, for example, to perform word class recognition using NLP, where the word class recognition can be downstream of the speech analysis module 120. According to one embodiment, for example, the combination of utterance / sentence boundary and word class recognition can be used for sentence element recognition. According to one embodiment, the device can, for example, be configured to perform a verb position analysis that is downstream of word class recognition using NLP. For example, word class recognition can be performed downstream of the speech analysis module (120). Similarly, verb placement analysis can be performed downstream of word class recognition. Alternatively, verb placement analysis and word class recognition can also be integrated into, or performed by, the speech analysis module (120). In one embodiment, the device can, for example, be configured to determine and output information about the quality of first-person speech production based on verb placement analysis. According to one embodiment, the device can, for example, be configured to determine and output the number of occurrences of words of the specific word class in the first recorded audio signal. According to one embodiment, the device can, for example, be configured to perform a verb placement / position analysis based on the sequence of word classes determined by NL1 / P that occurs within the sentence or utterance boundaries. This analysis allows conclusions to be drawn about the main clause or subordinate clause and its complexity within the utterance, sentence, or phrase based on the verb position, as well as providing information about the syntactic abilities of the individual. FH250504PDE-2026097150.DOCX To conduct a speech recording from the population. Because verbs have specific places within an utterance or sentence structure where they correctly stand. In one embodiment, the speech analysis module 120 can, for example, be configured to determine the sentence element boundaries of a plurality of sentence elements depending on the temporal length of the one or more speech pause segments in the recorded first audio signal. The speech analysis module 120 can, for example, be configured to determine the sentence element boundaries of one or more phrases depending on a temporal duration of sentence elements that is typical for the population. For example, a sentence element can comprise one or more words. According to one embodiment, the typical temporal duration of sentence element boundaries for the population can depend on the temporal duration of sentence element boundaries in a plurality of recorded audio signals from a plurality of persons in the population. In one embodiment, the device can, for example, be designed to determine the typical temporal duration of sentence element boundaries for the population by analyzing the temporal duration of the sentence element boundaries in the majority of the recorded audio signals of a majority of persons in the population. According to one embodiment, the typical duration of sentence element boundaries for the population can be shorter than the typical duration of utterance boundaries or sentence boundaries for the population. In one embodiment, the typical duration of sentence element boundaries for the population can be shorter than the typical duration of phrase boundaries for the population. According to one embodiment, the speech analysis module 120 can be configured to determine that a sentence element boundary is present in the recorded first audio signal if the temporal pause length of the speech pause segment is greater than or equal to a fifth temporal limit and less than or equal to a sixth temporal limit, wherein the fifth temporal limit and the sixth temporal limit depend on the temporal duration of the sentence element boundaries that is typical for the population. In one embodiment, the speech analysis module 120 can, for example, be configured to determine a typical individual temporal duration of the sentence element boundaries for the first person by individually adjusting the temporal duration of the sentence element boundaries typical for the population for the first person, depending on the temporal pause length of the one or more speech pause segments in the recorded first audio signal. And / or, the speech analysis module 120 can, for example, be configured to adjust the fifth temporal boundary and / or the sixth temporal boundary depending on the temporal pause length of the one or more speech pause segments in the recorded first audio signal. In exemplary embodiments, the device is configured to store and / or output one, several or all of the information specified above. Preferred embodiments are described below: Preferred embodiments enable automated recognition of utterance boundaries and sentence boundaries within a spontaneously produced and non-time-limited speech stream, such as recorded speech data, of a speaker, based on the language-specific prosody of the produced speech streams and sensitive detection of the pause structure using a VAD. For the development of this system, for example, individual statistics of the pause structure of the digital speech stream to be analyzed from a speaker of a population can be combined with the descriptive statistics of the pause analysis including threshold determination of an entire larger speech corpus of the corresponding population, consisting of several speakers of this population (e.g. children's speech corpus, here children under 5 years; other populations possible). This recognition of prosodic pause elements and their patterns, and thus the recognition of sentence / utterance boundaries, forms the basis for detecting linguistic phrases or semantically more complex sentence structures such as main clauses, subordinate clauses, and sentence elements, using further statistics and appropriate filtering of this output. In the development process, the data can be, for example, orthographically annotated and labelled according to utterance, sentence structure and phrases, and analyzed sensitively, i.e. very accurately in terms of timing, with regard to the pause structure by means of a SAD (see for example

[0022] ). FH250504PDE-2026097150.DOCX These automatically extracted pause times can, for example, be related to the manual labels of the utterances, sentence parts, semantics (optional phrases / non-phrases), whereupon new patterns and repetitions, e.g., in the pause structure, can be discovered, which can, for example, provide information about the presence of such sentence parts. Subsequently, further statistical analyses of the pause structure can be carried out in combination with a comparison of different annotation levels, such as utterances, sentence parts, semantics, and phrases / non-phrases, for example, based on a given corpus in the field of child language. From this, for instance, conclusions can be drawn about the acoustic analysis of deep structural patterns in the language in order to determine sentence part boundaries and sentence boundaries, or, in the case of children, utterance boundaries. Based on the pause structure and statistics, automated detection of utterance boundaries can be performed for children's language (e.g., the language of children under five years old), for example, without prior time limits, pre-annotation, punctuation labels, etc., including semantic assessment. Further filtering and the use of thresholds can also enable automated phrase recognition. In some implementations, the system can be re-adapted for a new population / corpus (development process). In specific implementations, this can be fully automated (automated process). For example, with known populations, manual adjustments may no longer be necessary. As a logic system, this can be adapted for any language and age group. One example of implementation (e.g., a development process) can be described as follows: First, the pause structure for a population can be determined. Pause detection can be performed using a fine SAD (see e.g.

[0022] ) and time limit output. Then, an evaluation of the SAD result can be performed in comparison to manually annotated utterances, sentence elements, and semantics. For example, histograms can be used for pattern recognition. The lower and upper limits of the population for the pause length that corresponds to the FH250504PDE-2026097150.DOCX Differentiation between single-word utterances (possibly meaningless utterances) and more complex sentences, including pauses in thought, word repetitions, and sentence fragments. This can result in pattern recognition / identification. Furthermore, statistics can be compiled on a children's corpus (+). This can involve using the upper and lower bounds to determine segments (boundiower x boundupper) and merged segments (x < bounder), i.e., an analysis of different boundaries based on the pause structure to filter individual words, complex sentence structures (main clause, subordinate clause), etc. Furthermore, statistics on pause structure and length can be generated. This can be achieved by individually analyzing the pause structure within these elements to identify distinctive pause types. Pattern recognition of prosodic pauses can also be used to delineate sentence elements and sentence boundaries. Prosodic pauses can be identified to classify repetitions, delays, and sentence boundaries. Additionally, rules for individually delineating segments can be established. In this context, the individual mean of the minimum and maximum values ​​in the histogram can be determined. Finally, a comparison with manual annotations, particularly regarding sentence elements and content, can be performed. Optionally, automated phrase recognition can be performed. Furthermore, optionally, a further subdivision into non-phrases (one-word / short utterances, repetitions, ...), into main clause, subordinate clause, i.e. into complex sentence structures (e.g. separate recognition of sentence boundaries and phrases) can be made. Furthermore, optional filtering can be used, especially with regard to overly short dialogic one-word utterances; short utterances; filler words and with regard to complex sentence structures based on prosodic pause structures. In some embodiments, only some of the above points may be implemented. Fig. 2 shows a device for speech analysis of a recorded audio signal according to a further embodiment, whose speech analysis module 120 has a sentence boundary detector 121 and an (optional) phrase detector 122. In a FH250504PDE-2026097150.DOCXExecution form: Based on the output of phrase recognition 122, an extended phrase recognition is performed. Fig. 3 shows a device for speech analysis of a recorded audio signal according to a further embodiment, whose speech analysis module has, in addition to the sentence boundary detector 121 and the (optional) phrase detector 122, an (optional) extended analysis module 123. In one embodiment, the extended analysis module 123 serves to evaluate the production of utterances and / or sentences and / or phrases. In another embodiment, the extended analysis module 123 serves for word class recognition. Further analyses can follow the extended analysis module 123 or within the framework of the extended analysis module, such as a determination of the word class position and / or the verb position and / or the sentence type in utterance / sentence or phrase. Figure 4 shows an exemplary representation of the analysis of different sentence parts and their word classes, depending on the strength of the filtering (thresholds), and their phrase (P) and non-phrase (NP) annotation according to one embodiment. Fig. 5 shows an exemplary representation of the pause duration of detected word repetition, detected hesitation and so-called syntactic guardrails according to one embodiment. Fig. 6 shows an exemplary representation of an automated recognition of different sentences or sentence parts depending on the strength of the filtering (thresholds) according to one embodiment. Another example of implementation (e.g., an automated process) can be described as follows: First, the automatically detected word boundaries (start, stop times) can be read using a SAD, integrated into (e.g. adapted) ASR. FH250504PDE-2026097150.DOCXDes Furthermore, elements can be merged based on specific thresholds. Fixed filtering limits can be determined using (+). Pause limits below the threshold in seconds can be determined. This allows for the separation of speech segments and a controlled, automated merging of automatically recognized segments (such as utterances, sentence fragments, main clauses / subordinate clauses) into utterances and sentence structures. In this context, filtering of excessively short, dialogic one-word utterances, short utterances, and filler words can be performed based on the length of the recognized speech segment. Furthermore, automated detection of sentence / utterance beginnings and ends (e.g., segment identification) is possible. This can be done using automatically generated individual statistics of the audio file (speaker) being analyzed. Individual threshold values ​​can be determined (via the histogram): If possible, additional automated detection of individual prosodic pauses and thus sentence structure, and segmentation of individual speech output into segments (utterances / sentence parts / main clause / subordinate clause) according to individual statistics based on SAD analysis in real time can be performed; if statistics cannot be calculated individually, an application to the threshold values, e.g., based on the statistics of the entire corpus, can be made. Furthermore, the individual elements can be categorized based on their respective threshold values. For example, the distinction between word repetition, hesitation, and meaningful utterances can be determined based on prosodic pause structure, using either individually or statistically determined values, and enabling the recognition of more complex sentence structures. Optionally, automated phrase recognition can be performed: This can involve further subdivision into non-phrases (e.g., one-word / short utterances, repetitions, etc.), such as main clauses and subordinate clauses. Filtering can be applied to complex sentence structures, for example, regarding excessively short dialogic one-word utterances; short utterances; filler words; or regarding complex sentence structures based on prosodic pause patterns. In some embodiments, only some of the above points may be implemented. FH250504PDE-2026097150.DOCXFig. 7 shows an automation of utterance / sentence boundaries or phrase recognition with subsequent NLP word class recognition and further analyses according to one embodiment. In specific implementations, a system for the fully automated detection of sentence and utterance boundaries is provided. This system is based on the recognition of prosodic pause elements and patterns, independent of the semantic content of the ASR (Action Resonance Structure), and the subsequent identification and automatic recognition of word classes and semantic levels (NLP model), which in turn allow inferences to be made about the syntactic guardrails. Here, the prosodic pause structure serves as a means of providing syntactic guardrails. This serves as a basis for subsequent syntactic and semantic analyses, such as the recognition and filtering of linguistic phrases or semantically more complex sentence structures like main clauses, subordinate clauses, or the verb position within these, or sentence elements. In particular, this system is based on the corpus analysis of child language, which does not yet fully follow the rule systems of adult language.It is particularly worth emphasizing here that the analysis works for spontaneous speech utterances without time limitations. These methods are applicable to any population and adult model, as well as any language situation. Even with adult language, recognizing sentence boundaries and distinguishing between meaningful utterances or phrases—and thus differentiating them from word repetitions, hesitations, etc.—can be useful for conducting further analyses, particularly regarding the assessment of language proficiency. Here, the MLU score and verb position are especially important markers. With other languages, training on language-specific sentence intonation, for example, can be implemented. Based on the pause structure and statistics, implementations thus enable the recognition of sentence boundaries for spontaneous speech utterances (or other situations) for any population without prior time limitations, pre-annotation, punctuation labels, etc., enabling sentence beginning and end recognition and semantic evaluation. In one embodiment, a speech recognition module, e.g. a VAD or a SAD, is designed to detect speech ranges and pauses, silence between the end of the smallest utterance unit / word end and the following utterance beginning / word beginning. FH250504PDE-2026097150.DOCXA basic analysis module can be designed for utterance or sentence boundary detection. Optionally, an extended analysis can be performed, for example, filtering the utterance or sentence boundary detection to identify phrases. Phrases will generally be smaller (shorter) than or equal to the sentence or utterance boundaries. In some implementations, word class recognition can be performed. Furthermore, other or additional analyses can optionally be carried out, such as word class position, verb position, and sentence type (main clause, subordinate clause) in an utterance / sentence or phrase. In some embodiments, a different approach to the pause areas in the smallest pause areas can be taken, such as the recognition of sentence elements. A specific implementation example is described in detail below. In this example, sentence boundaries are determined within an existing population and setting (context). A setting could be, for example, a game situation, a dialogue situation, a reading situation, spontaneous speech, etc. For example, it can be assumed that the lower and upper limits are known. For instance, it may be known that the most frequent lower and upper pauses are in the range of approximately 1 to 1.5 seconds. In one embodiment, an ASR system (which may include, for example, a SAD) may be integrated. If the embodiment is provided as software, for example, the provision of the executable file (which may include, for example, program code with population-specific limits) may be provided. The following is a specific example of this implementation. In a first step, an audio recording of the existing population can be loaded into the system (step A). Then, using VAD / SAD (speech range detection) integrated into an ASR, the system can automatically detect word boundaries (start and stop times) of the entire recording (step B). FH250504PDE-2026097150.DOCXDes Furthermore, the lower limit associated with the population can be used to merge elements: Words whose word boundaries (from the stop of the first word to the start of the next word) are separated by a pause (difference between stop and start) smaller than the lower limit can be combined. This creates speech segments whose duration is greater than that of the individual words (Step C). The lower and upper limits associated with the population (determined, for example, using descriptive statistics and a frequency analysis of a first and second local maximum at a resolution of 0.1 seconds of the histogram of the existing population within a population-appropriate pause range for utterance / sentence boundaries, e.g., between 1 and 1.5 seconds) are used to group the speech segments: If, for example, there is a pause between two segments that lies within the range between a lower and upper limit (lower limit <= pause <= upper limit), these segments are grouped together. This allows, for example, the filtering out of sentence fragments. Conversely, using the upper limit filters out even larger segment gaps, which can correspond to sentence fragments, hesitations, etc. See also FIG. 5 (Step D). The segments created in this way (e.g., in steps C and D) are filtered according to their duration. For example, a minimum duration (e.g., 0.5 s) can be defined as the lower limit for the length of a segment. If this segment duration is not met, the segment can be filtered out. In this way, for example, single words are filtered out. Thus, one-word utterances that are too short, unnecessary, or unclear are removed (step E). The segments identified above (e.g., in step C) can then be considered individually: Within each segment, the pause structure (which, for example, contains only pauses less than or equal to the lower limit) is examined again. To determine an individual limit, the frequency of the pauses can be determined, for example, using descriptive statistics (such as a histogram). The individual limit can be defined as the mean between the local minimum and the local maximum. If the individual limit is less than a limit determined by the population, the mean between the lower and upper limits of the population can be used as the limit. If a pause greater than the limit determined in this way exists within a segment, the segment can be split into two segments at this pause (step F). FH250504PDE-2026097150.DOCX The segments determined above (e.g., in step D) can now be considered individually: Within the segments, the pause structure (which, for example, contains only pauses less than or equal to the upper limit) is examined again. To determine an individual limit, the frequency of pauses can be determined using descriptive statistics (e.g., a histogram). The individual limit can be defined as the mean between the local minimum and the local maximum. If the individual limit is less than a limit determined by the population, the mean between the lower and upper limits of the population is used as the limit. If a pause greater than the limit determined in this way is present within a segment, the segment is split into two segments at this pause. This serves to check for a complex sentence structure or...the division into main clause and subordinate clause (step G). Furthermore, the segments determined above (e.g., in steps F and G) can be combined to form an output of the system: This output then contains, for example, the start and stop times of the automatically detected sentence boundaries. In the case of the population of children under six years of age, these are referred to as utterance boundaries instead of sentence boundaries (step H). In embodiments, a system according to the invention can, for example, be constructed as shown below: A first step could be, for example, the inclusion or use of a defined population and a (defined) setting (context). A further step could involve the integration of a fine SAD or the integration of an ASR system with fine SAD / VAD with word-level boundaries. Furthermore, a determination of the break structure of the population may be provided. Furthermore, it may be necessary to identify significant pauses. For example, peaks in the frequency analysis could be determined using descriptive statistics in the pause range, typical for the specific setting. In spontaneous speech, for instance, the lower limit might be around 1 second. FH250504PDE-2026097150.DOCX The concepts and procedures already described above can be used to implement the details. In various embodiments, extensions to the system may be provided. For example, sentence element recognition, as described above, can be implemented with individual pauses. Furthermore, phrase recognition can be implemented, for example, by means of further filtering. Furthermore, a determination of the MLU can be implemented by, for example, extending the code to include a calculation and a counting of phrases. Furthermore, a subsequent implementation of an NLP module, e.g., by

[0023] , can determine the word classes within the output utterance or sentence boundaries or phrases. Likewise, by extending the code to determine the verb position within these utterance or sentence boundaries or phrases, a statement can be made about the presence or ability to apply or produce certain syntactic rules. For this purpose, the word class "verb" is compared to the preceding and following word classes such as nouns, pronouns, conjunctions, etc., within the utterance or sentence boundaries or phrases with the applicable syntactic rule values ​​for main clauses, subordinate clauses, and specific verb positions within them. Other additions are conceivable. For example, in another embodiment, sentence elements can be determined based on the lower and upper limits associated with the population. These limits can be determined, for example, using descriptive statistics and a frequency analysis, whereby a local maximum can be determined at a resolution of 0.1 seconds of the histogram of the existing population within a smaller pause range for sentence elements appropriate to the population, e.g., between less than 1 second and 1.5 seconds pause length, or a local maximum can be determined within the pause ranges that lie within the utterance and sentence-phrase limits associated with the population. FH250504PDE-2026097150.DOCX According to one embodiment, the development of an exemplary system can proceed as follows, for example. As a first step, for example, a determination of the speech ranges of a given population (here spontaneous speech of children) can be carried out using a fine SAD (step A). Furthermore, a determination of pauses between the speaking sections can be implemented (step B). Furthermore, the break structure can be analyzed using descriptive statistics, for example using histograms (step C). Furthermore, speech segments can be grouped together if they are separated by a pause smaller than a first threshold, e.g. 1 s, since only pauses larger than such a first threshold (e.g. 1 s) are considered relevant for utterance / phrase recognition (step D). Furthermore, the speech ranges from step D can be combined, separated by a pause of a length between the first limit and a second limit (e.g., from the range 1 s <= x <= 1.5 s). The second limit (e.g., 1.5 s) can be determined, for example, as a rough guideline by examining the first histograms (step E). Furthermore, the exclusion of remaining segments whose duration is less than a third threshold (e.g., < 0.5 s) can be provided. This serves to eliminate background noise and single-word utterances (step F). Furthermore, a switch to using a SAD integrated into an ASR system may be implemented. This is more robust against non-speech background noise. In most cases, an ASR result (word start and word end) is also used for the analyses following utterance recognition (e.g., MLU) (step G). Furthermore, the first and second limits can be adjusted. For example, the limits defining the range from 1 s to 1.5 s can be changed to limits defining the range from 0.8 s to 1.3 s. This can be achieved, for example, by defining a first and second maximum of the histogram at a certain resolution. FH250504PDE-2026097150.DOCX is used from 0.1 seconds, which involves considering a larger population and using the ASR-internal SAD (step H). Furthermore, an analysis of the pause structure within the summarized speech areas determined in steps D and E may be considered. This could involve, for example, a manual, subjective identification of the different types of pauses, such as 'hesitation', 'word repetition', and 'demarcation' (step I). In a further step, a statistical analysis can be carried out of the types of breaks defined in step I (e.g. three), which differ statistically significantly from each other in their break duration (step J). Furthermore, a fixed threshold can be determined using descriptive statistics to distinguish between different types of pauses (step K). For example, the pause types 'hesitation' and 'word repetition' can be distinguished from the pause type 'demarcation'. For example, in a specific population, a value of 0.4 s for the speech ranges determined in step D results in a discrimination accuracy of 73%. For the speech ranges determined in step E, a value of 0.5 s results in a discrimination accuracy of 75%. However, this value does not determine the individuality of the pause structure for each member of the population (here, each child). Not all children are optimally represented by the thresholds. Other populations exhibit different numerical values. Furthermore, an individual threshold value can be determined for each child, whereby a separate analysis of the pause structure within the speech ranges determined in steps D and E can be performed. Descriptive statistics are determined using a histogram, and various methods for determining the individual threshold value are tested. In one embodiment, the mean between the first local minimum and the first local maximum can be determined as the most suitable (step L). Furthermore, the use of the individual limit values ​​determined in step L for further segmentation of the speech areas determined in steps D and E, if necessary, may be provided for if a pause greater than the determined limit value exists within an area (step M). FH250504PDE-2026097150.DOCX The inventive concept allows the language of a subject from the population (e.g., a child) to be analyzed. The following statements or feedback can be generated (e.g., by the system): A feedback (output) about the number of phrases or (complex) sentences that the subject of the population (e.g. the child) spoke during the recording. A feedback (output) that outputs the MLU of the subject to the population (e.g., the child). A statement about the level of a dropout rate and / or repetition rate and / or hesitancy of the subject in the population (e.g., the child). A statement about the verb position within an utterance / sentence / phrase and thus correctness and ability in the application of syntactic rules of the subject of the population (e.g. the child). In one embodiment, direct feedback and a prompt to repeat would be conceivable to achieve fluent speech. For example, direct feedback to the speaker, such as a child, would indicate whether the last spoken sentence was grammatically correct or incorrect, for instance, regarding sentence fragments or verb placement. If a sentence is grammatically incorrect, the system could, for example, prompt the speaker to repeat the speech input. A statement about the rate of single-word utterances. These statements allow for objective conclusions regarding the quality of speech production and the assessment of speech style. This is difficult and prone to error to manually verify in real time, even for speech therapists or coaches, or by recording and evaluating the speech. This requires a significant investment of time, resources, and expertise. Even speech therapists without experience or further expertise cannot always accurately categorize speech into phrases, or it represents an enormous time commitment. A manual transcription of the speech production (Speech-to-Text) is necessary, followed by a subsequent categorization of the transcript into phrases and a count of the words per phrase in relation to the proportion of phrases. Therefore, determining the MLU (Maximum Language Level) is only possible based on the manual categorization of the recorded and transcribed speech into phrases and non-phrases. FH250504PDE-2026097150.DOCX subsequent counting is possible. The concepts according to the invention thus solved a high time-consuming and error-prone effort. The following are application areas of the invention: Thus, such implementations can be generally used in speech analysis systems. For example, implementations can be applicable to any population and language, as well as any communication situation and setting, with the appropriate training. To determine the MLU fully automatically, sentence delimitation, in particular phrase recognition in combination with single-word recognition based on word boundaries from an ASR system, can be used. (For example, the word boundaries can be determined based on the ASR, and the pauses are then available as a pause structure based on the word boundaries.) As the most common form of assessment of progress in grammatical development worldwide, the MLU can be used to determine the language development level of children, but also to identify communication deficits such as in autism spectrum disorder (ASD) or other pathologies and age-related neurological diseases. Various embodiments of the invention can be used in these applications. Implementations can be used as part of a more comprehensive syntax analysis for various applications, such as (foreign) language acquisition. Application areas also include the identification and combined use of individual and population-specific break structures and their thresholds. These devices can also be used in the early detection / diagnosis of speech and language disorders. This includes the early detection / diagnosis of genetic syndromes and complex developmental disorders (such as autism spectrum disorder), as well as the diagnosis of neurological disorders (especially differential diagnosis). Another area of ​​application lies in measuring disease progression, such as dementia. These devices can also be used to measure the effectiveness of interventions or therapies. For example, FH250504PDE-2026097150.DOCX These methods are repeatedly applied to determine disease progression or treatment success. In some implementations, the individual pause structure can be used to automatically detect speaker changes, for example, in a dialogue system. This can also be used to generate pauses appropriate to the dialogue and adapted to the individual speaker. This solves the problem of how speech output from conversational agents can be perceived as natural. It thus becomes possible to conduct a natural dialogue with the conversational agent based on the pause structure of the dialogue partner. In some implementations, the speaker's native language or the currently spoken language can also be identified based on the pause structure. Application areas include, in particular, the language analysis of children, especially around the time of starting school and in daycare settings, in education, in healthcare, in early intervention diagnostics, such as interdisciplinary early intervention diagnostics, early intervention for children with special needs, and in examinations by doctors, such as pediatricians during routine check-ups. Further areas of application include subsequent language diagnostics for therapists, speech therapists, special education teachers, early intervention for children with special needs, and in hospitals, e.g., university hospitals, such as stroke centers / units, social pediatric centers, centers for rare diseases, for school psychological services, and in early intervention centers. Furthermore, application areas include automated communication applications and dialogue systems of all kinds. Additional application areas include forensic language analysis (e.g., by police, military, etc.). Advantages arise from the general applicability of the designs, freedom of application in different situations, where even short recordings are sufficient, and from the automation of result generation and result interpretation. Implementation models realize automatic phrase recognition and sentence boundary recognition. FH250504PDE-2026097150.DOCX In principle, application areas of embodiments lie in the field of recognition and analysis of speech development disorders, in the field of determining speech fluency and speaker posture, in the field of stress / accent recognition, for foreign language acquisition, in the field of monotony analysis, for speech optimization, etc. The use of embodiments in specific application areas is described in more detail below. One application of these devices lies in the early detection of specific language impairment (SLI). Approximately 7-8% of all preschool children are affected by a SLI. Possible symptoms include, for example, omitting words in sentences or unusual verb placement in main clauses. Following diagnosis, individually tailored therapy is provided (see [1]). "The relevance of early detection of primary SLI has been reflected in recent years, among other things, in the high number of speech therapies, which are usually prescribed after the third birthday. According to the medical reports of the Scientific Institute of the AOK (German statutory health insurance), SLI (ICD-10-GM-22, code F80) was the most frequent diagnosis among speech therapy patients in childhood over the past three years. SLI diagnoses dominated among AOK-insured speech therapy patients in 2018 (55.2%), in 2019 it was 56.1%, and in 2020 it had already reached 57.0%." (see [2]).For relevance, see also [4], [5], [6] and [7]. A large number of children require speech therapy (see also [7]). Both the early detection and the therapy of language development disorders can be supported by specific implementations. Sentence structure is of particular importance in this context, making automatic recognition of this structure necessary. A specific embodiment lies in stress / accent / intonation recognition: Only those who master the intonation of a language can be considered competent speakers (see

[0012] ). Intonation is a subfield of phonetics and phonology. Intonation is also important for the grammar of German, both in teaching German as a foreign language (GFL) and as a native language (see

[0013] ). Intonation must be learned just like morphology and sentence structure. Intonation errors are grammatical errors. They can significantly impair communication (see

[0013] ). The distinctions between accented and unaccented parts of the utterance, between rising and falling accents, and between high and low accent tones fulfill important functions in German for organizing the flow of information between communication partners and thus for integrating utterances into the communication and interaction context.FH250504PDE-2026097150.DOCX (see

[0013] ). Intonation systematically and regularly encodes information that is essential for successful communication (see

[0013] ). In a foreign language, however, communication problems are almost inevitable when the intonation system and phonetic realization are so different (see

[0013] ). Intonation occurs not only at the word level but also at the sentence level. Therefore, phrase recognition is necessary for successful automatic assessment. Automating this assessment can be of interest in various areas due to increasing digitalization. These include (foreign) language acquisition, speech therapy, self-improvement, public speaking (radio, theater, politics, etc.), professional contexts, teaching, customer service and sales, etc. The use of embodiments in this field of application can provide assistance. Another example concerns foreign language acquisition, specifically learning German: Millions of people worldwide are learning German as a foreign language (see [8] and

[0011] ). Increasing digitalization is also impacting how people learn languages. The use of online courses for personal development has been steadily increasing for years (see [8], especially the table on p. 48, and also [9]). Furthermore, the use of NLP is becoming increasingly widespread and is growing rapidly, and this growth is expected to continue in the future (see

[0010] ). In online foreign language acquisition, automatic phrase recognition of expressions can support the assessment of learning success, for example, by evaluating sentence intonation and accent. Another area of ​​application lies in the detection of monotonous speech disorders or pathologies: In Germany, over 100,000 people live with aphasia – a language disorder that frequently occurs after a stroke. Approximately 25,000 people are newly diagnosed with it each year. Strokes are the most common cause of aphasia, accounting for about 80 percent of cases, affecting around 270,000 people annually in Germany. Aphasia develops in about one-third of all first-time strokes, often resolving within the first four weeks. In about 20 percent of those affected, the language disorder remains permanent after a stroke.

[0014] Regular speech training can help reduce these problems. E-learning resources are used in this context (see

[0015] ). Objective feedback is necessary to complete effective self-training.The training material typically consists not only of individual words, but also of sentences or entire texts. This leads to increased speaking effort, but also to more natural intonation and more. FH250504PDE-2026097150.DOCX an increased naturalness of the situation. Especially at the text level, automatic recognition of phrases or sentence boundaries can support the automated assessment of sentence melody, intonation, and stress, according to various implementations. Another embodiment lies in speech optimization, which is also highly relevant (

[0016] ,

[0017] ). In presentations, not only the content is important, but also the speaking itself, whereby vocal characteristics and communication skills can reinforce a positive impression (see

[0018] ), and are often trained in soft skills training (see

[0019] ). A digital platform with integrated feedback can enable independent, time-flexible training. Automatic sentence boundary detection, as described in the embodiments, can help to specifically examine sentences and refine the training. Not every sentence in a speech has the same importance and requires pronounced emphasis. Automatic detection of particularly emphasized sentences can help to identify the focus of the speech and determine which sentences may require more intensive training. Another embodiment involves the recognition of the speaker's attitude, which can occur on two levels: firstly, on the perceptual level, where the speaker is judged by others in terms of competence, trustworthiness, etc., based on their manner of speaking; and secondly, on the speaker's subjective level, where their manner of speaking expresses their own attitude or self-confidence. There is a prosodic difference between perceived and one's own sense of security / confidence (see

[0020] ). Automatic recognition of both the speaker's subjective and perceptual attitudes according to embodiments can be of interest in various fields, and their analysis can provide insights into personal views (see

[0018] ). Here, the use of automatic phrase recognition according to embodiments can be useful for analyzing attitudes in the context of specific sentences and, consequently, their content.Perceptual attitudes can also be relevant in the context of learning. For example, it has been shown that learners were better able to use utterances perceived as confident when learning concepts (see

[0021] ). Accordingly, a perceptually perceived confident attitude can help to convey content or to persuade others. Training this linguistic self-confidence can be supported by an objective tool that uses phrase / sentence recognition according to embodiments to segment the signal. FH250504PDE-2026097150.DOCX Although some aspects have been described in connection with an apparatus, it is understood that these aspects also constitute a description of the corresponding method, such that a block or component of an apparatus is also to be understood as a corresponding method step or as a feature of a method step. Similarly, aspects described in connection with or as a method step also constitute a description of a corresponding block, detail, or feature of a corresponding apparatus. Some or all of the method steps may be performed by (or using) a hardware apparatus, such as a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, some or more of the main method steps may be performed by such an apparatus. Depending on specific implementation requirements, embodiments of the invention can be implemented in hardware or in software, or at least partially in hardware or at least partially in software. The implementation can be carried out using a digital storage medium, for example, a floppy disk, a DVD, a Blu-ray disc, a CD, a ROM, a PROM, an EPROM, an EEPROM, a FLASH memory, a hard disk, or another magnetic or optical storage medium, on which electronically readable control signals are stored. These control signals can interact with a programmable computer system in such a way that the respective method is carried out. Therefore, the digital storage medium can be computer-readable. Some embodiments according to the invention therefore include a data carrier which has electronically readable control signals which are able to interact with a programmable computer system in such a way that one of the methods described herein is carried out. In general, embodiments of the present invention can be implemented as a computer program product with a program code, wherein the program code is effective in carrying out one of the methods when the computer program product runs on a computer. The program code can also be stored on a machine-readable medium, for example. FH250504PDE-2026097150.DOCX Other embodiments include the computer program for carrying out one of the methods described herein, wherein the computer program is stored on a machine-readable medium. In other words, an embodiment of the method according to the invention is thus a computer program that includes program code for carrying out one of the methods described herein when the computer program runs on a computer. Another embodiment of the methods according to the invention is thus a data carrier (or a digital storage medium or a computer-readable medium) on which the computer program for carrying out one of the methods described herein is recorded. The data carrier or the digital storage medium or the computer-readable medium is typically tangible and / or non-volatile. Another embodiment of the method according to the invention is thus a data stream or a sequence of signals that represents the computer program for carrying out one of the methods described herein. The data stream or sequence of signals can be configured, for example, to be transferred via a data communication connection, such as the Internet. Another embodiment comprises a processing device, for example a computer or a programmable logic device, which is configured or adapted to perform one of the methods described herein. Another embodiment comprises a computer on which the computer program for performing one of the procedures described herein is installed. Another embodiment of the invention comprises a device or system designed to transmit a computer program for carrying out at least one of the methods described herein to a receiver. The transmission can be, for example, electronic or optical. The receiver can be, for example, a computer, a mobile device, a storage device, or a similar device. The device or system can, for example, include a file server for transmitting the computer program to the receiver. FH250504PDE-2026097150.DOCX In some embodiments, a programmable logic device (for example, a field-programmable gate array, an FPGA) can be used to perform some or all of the functionalities of the methods described herein. In some embodiments, a field-programmable gate array can interact with a microprocessor to perform one of the methods described herein. Generally, in some embodiments, the methods are performed by any hardware device. This can be general-purpose hardware such as a computer processor (CPU) or method-specific hardware such as an ASIC. The embodiments described above merely illustrate the principles of the present invention. It is understood that modifications and variations of the arrangements and details described herein will be obvious to other people skilled in the art. Therefore, it is intended that the invention be limited only by the scope of protection set forth in the following claims and not by the specific details presented herein by way of description and explanation of the embodiments. FH250504PDE-2026097150.DOCX References: [1] dbl: language development disorder; https: / / www.dbl-ev.de / kinder-und-jugendliche / sprachentwicklungsstoerung Last accessed: March 4, 2025 [2] Kiese-Himmel, C. (2022). Early detection of primary language development disorders - increasing relevance due to changes in diagnostic criteria? Bundesgesundheitsblatt-Gesundheitsforschung-Gesundheitsschutz, 65(9) , 909-916. [3] Market size and market share of speech and language disorders until 2030, https: / / www.theinsightpartners.com / de / reports / speech-and-language-disorder-market last accessed: 04.03.2025 [4] US market size for speech therapy, share | Analysis

[2030] , https: / / www.fortunebusinessinsights.com / de / us-logop-diemarkt-105574 Last accessed: March 4, 2025 [5] Children's Atlas - Outpatient Diagnoses ICD-3 Codes - bifg, https: / / www.bifg.de / publikationen / reporte / arztreport / arztreport-kinderatlas-ambulante-diagnosen-nach-top-icd-10-dreistellern last accessed: 04.03.2025 [6] How many children are there in Germany?, https: / / alleantworten.de / wie-viel-kinder-gibt-es-in-deutschland#:~:text=ln%20Deutschland%20leben%20derzeit% 2010%2C65%20Millionen%20Kinder%20im,13%2C75%20Millionen%20Kinder%20und%20Jugendliche%20unter%2018%20Jahren. Last accessed: March 4, 2025 [7] 2.7 million children need speech therapy, but only 70,000 receive it! - VDLS ■ Association of German Speech Therapists and Speech-Language Therapy Professionals, https: / / www.vdls-ev.de / millionen-kinder-benoetigen-logopaedische-therapie (last accessed: March 4, 2025) [8] German as a Foreign Language Worldwide - DAAD, https: / / www.daad.de / de / der-daad / kommunikation-publikationen / veroeffentlichungen-publikationen / deutsch-als-fremdsprache-weltweit / Last accessed: March 5, 2025 [9] Online language courses - worldwide revenue until 2033 | Statista, FH250504PDE-2026097150.DOCX https: / / de.statista.com / statistik / daten / studie / 1402768 / umfrage / umsatz-markt-online-sprachkurse-weltweit / Last accessed: March 5, 2025

[0010] Natural Language Processing - Worldwide | Market Forecast, https: / / de.statista.com / outlook / tmo / kuenstliche-intelligenz / natural-language-processing / weltweit Last accessed: March 5, 2025

[0011] German-speaking people worldwide | Statista, https: / / de.statista.com / statistik / daten / studie / 1119851 / umfrage / deutschsprachige-menschen-weltweit / last accessed: 05.03.2025

[0012] Intonation, https: / / grammis.ids-mannheim.de / systematische-grammatik / 2336 last accessed: 05.03.2025

[0013] Blühdorn, H. (2013). Intonation in German just a question of beautiful sound? Pandaemonium Germanicum, 16, 242-278.

[0014] Aphasia: Definition, Symptoms and Forms | German Brain Foundation, https: / / hirnstiftung.org / 2023 / 12 / aphasie / last accessed: March 5, 2025

[0015] Ganzeboom, M. et al. (2022). 'A serious game for speech training in dysarthric speakers with Parkinson's disease: Exploring therapeutic efficacy and patient satisfaction'. In: International Journal of Language & Communication Disorders.

[0016] Events Market: Events and Participants | Statista, https: / / de.statista.com / statistik / daten / studie / 233136 / umfrage / veranstaltungen-und-teilnehmer-auf-dem-veranstaltungsmarkt-in-deutschland / #:~:text= The number of events at conferences and congresses increased to around 311 million per year. Last accessed: March 14, 2025

[0017] Events industry: Distribution by type 2023 | Statista, https: / / de.statista.com / statistik / daten / studie / 159424 / umfrage / verteilung-der-veranstaltungen-nach-veranstaltungsart / last accessed: 14.03.2025

[0018] Dietrich, Bryce J., Matthew Hayes, and Diana Z. O'brien. "Pitch perfect: Vocal pitch and the emotional intensity of congressional speech." American Political Science Review 113.4 (2019): 941-962. FH250504PDE-2026097150.DOCX

[0019] Online Soft Skills Training Market Size | Forecast - 2032, https: / / www.alliedmarketresearch.com / online-soft-skills-training-market-A295265 letzter Zugriff: 14.03.2025

[0020] Pon-Barry, H., & Shieber, S. (2010, May). Assessing self-awareness and transparency when classifying a speaker’s level of certainty. In Speech Prosody.

[0021] Barr, D. J. (2003). Paralinguistic correlates of conceptual structure. Psychonomic Bulletin & Review, 10, 462-467.

[0022] International Telecommunication Union, Recommendation P.56: Objective measurement of active speech level, Dec. 2011, https: / / www.itu.int / rec / T-REC- P.56-201112-l / en.

[0023] https: / / stanfordnlp.github.io / stanza / index.html Letzter Zugriff: 23.03.2025.

[0024] https: / / intrafind.com / de / blog / natural-language-processing-best-practice Letzter Zugriff: 23.03.2025.

[0025] Kauschke, Dörfler, Siegmüller (2022): Patholinguistic diagnostics in speech and language disorders (PDSS) FH250504PDE-2026097150.DOCX

Claims

38 Patent claims 1. Device for speech analysis of a recorded first audio signal in which spoken language of a first person from a population is recorded, wherein a population comprises a plurality of persons with the same or at least similar characteristics in speech behavior, wherein the device comprises: A speech recognition module (110) for speech detection, which divides the recorded initial audio signal into a plurality of speech segments that contain speech and into a plurality of speech pause segments that do not contain speech, a speech analysis module (120) that is trained to determine the temporal pause length of one or more speech pause segments of the majority of speech pause segments, wherein the speech analysis module (120) is designed to determine utterance boundaries or sentence boundaries of one or more sentences or utterances, depending on the temporal pause length of the one or more speech pause segments in the recorded first audio signal, wherein the language analysis module (120) is designed to determine the utterance boundaries or the sentence boundaries of one or more sentences or utterances, depending on a temporal duration of utterance boundaries or sentence boundaries that is typical for the population.

2. Device according to claim 1, where the typical temporal duration of utterance boundaries or sentence boundaries for the population depends on the temporal duration of utterance boundaries or sentence boundaries in a plurality of recorded audio signals from a plurality of persons in the population.

3. Device according to claim 2, wherein the device is designed to determine the typical temporal duration of the utterance boundaries or sentence boundaries for the population by analyzing the temporal FH250504PDE-2026097150.DOCX39 To determine the duration of utterance boundaries or sentence boundaries in the majority of recorded audio signals from a majority of people in the population.

4. Device according to one of the preceding claims, wherein the device is designed to determine which population the first person from a plurality of two or more populations belongs to.

5. Device according to claim 4, wherein the language analysis module (120) is designed to determine the utterance boundaries or sentence boundaries of the one or more sentences or utterances in such a way that, depending on which population from the majority of the two or more populations the first person belongs to, a temporal duration of utterance boundaries or sentence boundaries typical for that population is selected, and the utterance boundaries or sentence boundaries of the first person are determined depending on this typical temporal duration.

6. Device according to claim 5, where, if the first person belongs to a first population from the majority of the two or more populations, the temporal duration of utterance boundaries or sentence boundaries typical for this population has a first temporal duration, and where if the first person belongs to a second population from the majority of the two or more populations that is different from the first population, the temporal duration of utterance boundaries or sentence boundaries typical for this population has a second temporal duration that is different from the first temporal duration.

7. Device according to one of claims 4 to 6, wherein the device is configured to provide an input interface through which a user can input which population the first person from the plurality of two or more populations belongs to. FH250504PDE-2026097150.DOCX40 8. Device according to any one of claims 4 to 7, wherein the device is configured to determine, by analyzing the recorded first audio signal, which population the first person belongs to from the plurality of two or more populations.

9. Device according to one of the preceding claims, wherein the speech analysis module (120) is configured to determine that an utterance boundary or a sentence boundary is present in the recorded first audio signal if the temporal pause length of the speech pause segment is greater than or equal to or greater than a first temporal threshold and less than or equal to or less than a second temporal threshold, wherein the first temporal threshold and the second temporal threshold depend on the temporal duration of the utterance boundaries or sentence boundaries that is typical for the population.

10. Device according to claim 9, wherein the speech analysis module (120) is trained to determine a typical individual temporal duration of the utterance boundaries or sentence boundaries of the first person by individually adjusting the temporal duration of the utterance boundaries or sentence boundaries typical for the population for the first person, depending on the temporal pause length of the one or more speech pause segments in the recorded first audio signal, and / or wherein the speech analysis module (120) is configured to adjust the first temporal limit and / or the second temporal limit depending on the temporal pause length of the one or more speech pause segments in the recorded first audio signal.

11. Device according to claim 10, wherein the speech analysis module (120) is designed to disregard, when determining the typical individual temporal duration of the first-person utterance or sentence boundaries or when adjusting the first limit and / or the second limit, those speech pause segments of the one or more speech pause segments whose duration exceeds the typical temporal duration of the utterance or sentence boundaries for the population by FH250504PDE-2026097150.DOCX falls below a first deviation value or its duration exceeds the typical temporal duration of the utterance boundaries or sentence boundaries for the population by more than a second deviation value.

12. Device according to one of the preceding claims, wherein the speech analysis module (120) is configured to identify an utterance in the recorded first audio signal by determining the beginning and end of the utterance through two utterance boundaries in the audio signal, between which there is no further utterance boundary; or wherein the speech analysis module (120) is configured to identify a sentence in the recorded first audio signal by determining the beginning and end of the sentence by two sentence boundaries in the audio signal, between which there is no further sentence boundary.

13. Device according to one of the preceding claims, wherein the speech analysis module (120) is designed to determine phrase boundaries of a plurality of phrases, depending on the temporal pause length of the one or more speech pause segments in the recorded first audio signal, wherein the language analysis module (120) is designed to determine the phrase boundaries of one or more phrases depending on a temporal duration of phrase boundaries typical for the population.

14. Device according to claim 13, where the typical temporal duration of phrase boundaries for the population depends on the temporal duration of phrase boundaries in a plurality of recorded audio signals from a plurality of persons in the population.

15. Device according to claim 14, wherein the device is designed to determine the typical temporal duration of the phrase boundaries for the population by analyzing the temporal duration of the FH250504PDE-2026097150.DOCX To determine phrase boundaries in the majority of recorded audio signals from a majority of people in the population.

16. Device according to one of claims 13 to 15, further dependent on claim 5, wherein the device is designed to determine for each population of the majority of the two or more populations a typical duration of utterance boundaries or sentence boundaries for that population.

17. Device according to claim 16, wherein the device is configured to determine, for each population of the plurality of the two or more populations, the typical temporal duration of the phrase boundaries for the population by analyzing the temporal duration of the phrase boundaries in the plurality of the recorded audio signals of a plurality of persons of the population.

18. Device according to claim 17, where the typical temporal duration of utterance boundaries or sentence boundaries for a first population from the majority of the two or more populations has a first temporal duration, and where the typical temporal duration of utterance boundaries or sentence boundaries for a second population from the majority of the two or more populations has a second temporal duration that differs from the first temporal duration.

19. Device according to claims 13 to 18, where the typical duration of phrase boundaries for the population is shorter than the typical duration of utterance or sentence boundaries for the population.

20. Device according to one of claims 13 to 19, FH250504PDE-2026097150.DOCX43 wherein the speech analysis module (120) is configured to determine that a phrase boundary is present in the recorded first audio signal if the temporal pause length of the speech pause segment is greater than or equal to or greater than a third temporal limit and less than or equal to or less than a fourth temporal limit, wherein the third temporal limit and the fourth temporal limit depend on the temporal duration of phrase boundaries typical for the population.

21. Device according to claim 20, wherein the speech analysis module (120) is trained to determine a typical individual temporal duration of the phrase boundaries of the first person by individually adjusting the temporal duration of the phrase boundaries typical for the population for the first person, depending on the temporal pause length of the one or more speech pause segments in the recorded first audio signal, and / or wherein the speech analysis module (120) is configured to adjust the third temporal limit and / or the fourth temporal limit depending on the temporal pause length of the one or more speech pause segments in the recorded first audio signal.

22. Device according to one of claims 13 to 21, further dependent on claim 12, wherein the speech analysis module (120) is configured to identify a phrase in the recorded first audio signal by defining the beginning and end of the phrase through a first phrase boundary or a first utterance boundary and through a second phrase boundary or a second utterance boundary in the audio signal, between which there is no further phrase boundary or utterance boundary; or wherein the speech analysis module (120) is configured to identify a phrase in the recorded first audio signal by determining the beginning and end of the phrase by a first phrase boundary or a first sentence boundary and by a second phrase boundary or a second sentence boundary in the audio signal, between which there is no further phrase boundary or sentence boundary. FH250504PDE-2026097150.DOCX44 23. Device according to one of claims 13 to 22, wherein the speech analysis module (120) is designed to determine and output an approximation of a linguistic measure of a Mean Length of Utterance (MLU) in the audio signal using the temporal phrase length of the majority of phrases in the audio signal.

24. Device according to one of the preceding claims, wherein the device further comprises a dialogue agent which adapts the pause lengths of its speech output to the speech pause segments of the majority of speech pause segments of a user in the audio signal.

25. Device according to one of the preceding claims, wherein the device is designed to determine and output information about the quality of first-person speech production based on the determination of utterance or sentence boundaries.

26. Device according to one of the preceding claims, wherein the device is configured to perform word class recognition using NLP, e.g. wherein the word class recognition is downstream of the speech analysis module (120).

27. Device according to one of the preceding claims, wherein the device is designed to perform a verb position analysis, e.g. which is downstream of word class recognition using NLP.

28. Device according to claim 27, the device is designed to determine and output information about the quality of first-person speech production based on verb placement analysis.

29. Device according to one of claims 26 to 28, FH250504PDE-2026097150.DOCX45 wherein the device is configured to determine and output a number of occurrences of words of the specified word class in the first recorded audio signal.

30. Device according to one of the preceding claims, wherein the speech analysis module (120) is designed to determine sentence element boundaries of a plurality of sentence elements, depending on the temporal pause length of the one or more speech pause segments in the recorded first audio signal, wherein the language analysis module (120) is designed to determine the sentence element boundaries of one or more phrases depending on a temporal duration of sentence elements that is typical for the population.

31. Device according to claim 30, where the typical temporal duration of sentence element boundaries for the population depends on the temporal duration of sentence element boundaries in a plurality of recorded audio signals from a plurality of persons in the population.

32. Device according to claim 31 , wherein the device is designed to determine the temporal duration of the sentence element boundaries typical for the population by analyzing the temporal duration of the sentence element boundaries in the majority of the recorded audio signals of a majority of persons in the population.

33. Device according to one of claims 30 to 32, where the typical duration of sentence element boundaries for the population is shorter than the typical duration of utterance boundaries or sentence boundaries for the population.

34. Device according to claim 33, further dependent on any one of claims 13 to 23, FH250504PDE-2026097150.DOCX46 where the typical duration of sentence element boundaries for the population is shorter than the typical duration of phrase boundaries for the population.

35. Device according to one of claims 30 to 34, wherein the speech analysis module (120) is configured to determine that a sentence element boundary is present in the recorded first audio signal if the temporal pause length of the speech pause segment is greater than or equal to or greater than a fifth temporal limit and less than or equal to or less than a sixth temporal limit, wherein the fifth temporal limit and the sixth temporal limit depend on the temporal duration of the sentence element boundaries typical for the population.

36. Device according to claim 35, wherein the speech analysis module (120) is trained to determine a typical individual temporal duration of the sentence element boundaries of the first person by individually adjusting the temporal duration of the sentence element boundaries typical for the population for the first person, depending on the temporal pause length of the one or more speech pause segments in the recorded first audio signal, and / or wherein the speech analysis module (120) is configured to adjust the fifth time limit and / or the sixth time limit depending on the time interval of the one or more speech pause segments in the recorded first audio signal.

37. Method for language analysis of a recorded first audio signal in which spoken language of a first person from a population is recorded, wherein a population comprises a plurality of persons with the same or at least similar characteristics in language behavior, wherein the method comprises: Dividing the recorded first audio signal into a plurality of speech segments that contain speech and into a plurality of speech pause segments that do not contain speech. FH250504PDE-2026097150.DOCX Determining the temporal pause length of one or more language pause segments of the plurality of language pause segments, where, depending on the length of the pauses in one or more speech pauses in the recorded first audio signal, utterance boundaries or sentence boundaries of one or more sentences or utterances are determined, wherein the utterance boundaries or the sentence boundaries of one or more sentences or utterances are determined depending on a temporal duration of utterance boundaries or sentence boundaries that is typical for the population.

38. Computer program comprising program code for carrying out the method according to claim 37, when the computer program is executed on a computer or signal processor. FH250504PDE-2026097150.DOCX