A computer-implemented system and method for deriving and generating a voice response to a user input

GB2638529APending Publication Date: 2025-08-27EVANS CELEVATORON
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
GB2024017309
Authority / Receiving Office
GB · GB
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-27
Filing Date
2024-11-26
Publication Date
2025-08-27

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A computer-implemented system for providing text to speech (TTS) comprises a user communications device 104 having an input device 134 to receive text data from a text data source; and an audio playback device 134 for outputting an audio response. The user communications device 104 being configured to receive a text data stream from a text data source. In real time, as soon as said text data stream has started to be received, a speech emphasis handler module is used to segment the incoming text data into text data segments based on natural speech patterns and / or emphasis. The text data segments are passed, as they are created and in order, to a speech recognition module to transform each said text data segment into representative digital audio data. In real time, as the representative digital audio data is generated, digital audio data is fed to an audio buffer and stored. In real time and in order, the audio segments are delivered to the audio playback device 134. The text data source may be a Large Language Model (LLM) or Large Multimodal Model (LMM) API. The application may produce natural-sounding speech in real-time, without waiting for the entire text to be downloaded.
Need to check novelty before this filing date? Find Prior Art

Description

Field of the Invention This invention relates generally to a computer-implemented system and method for deriving and generating a voice response to user inputs to a Large Language Model (LLM) or Large Multimodal Model (LMM) and, more particularly, but not necessarily, to a computer-implemented voice chat system and method utilizing an LLM or LMM. Background of the Invention Text to speech conversion engines exist. However, they usually require the entire piece of text to be input and recorded before conversion and speech synthesis can begin. There is a need for an improved text to speech synthesis apparatus and aspects of the present invention seek to address this issue, amongst others. Statements of Invention In accordance with an aspect of the invention, there is provided a computer-implemented system for facilitating a text to speech application, comprising a user communications device configured to communicate with a Large Language Model (LLM) or Large Multimodal Model (LMM) API, the user communications device comprising: a processor; a memory having instructions stored therein to be executed under control of the processor; an input device for receiving text data from a text data source; and an audio playback device for outputting an audio response; the user communications device being configured, under control of the processor, to execute instructions stored in the memory to: receive a text data stream from a text data source; in real time, as soon as said text data stream has started to be received, use a speech emphasis handler module to segment said incoming text data into text data segments based on natural speech patterns and / or emphasis, and pass the text data segments, in order, as they are created, to a speech recognition module to transform each said text data segment into representative digital audio data; in real time, as said representative digital audio data is generated, start feeding said digital audio data to an audio buffer and store it therein; and deliver, in real time and in order, said audio segments to said audio playback device. By segmenting the incoming text data as it is being received, and passing each segment, as soon as it has been created, to the speech recognition module, the text-to-audio data transformation process can be effected as soon as the text data from the text data source starts to be received, feeding the audio data to the audio buffer as it is generated; and then, as soon as the first audio segment is defined, starting to continuously deliver the audio segments, in order, to the audio playback device, the claimed system enables a synthesized speech response to a text input in natural speech patterns to the user, in real time, without interruption or quality degradation. Other aspects an optional features of the invention are as set out in the appended claims. These and other aspects of the invention will be apparent from the following detailed description. Brief Description of the Drawings Embodiments of the invention will now be described, by way of examples only, and with reference to the accompanying drawings, in which: Figure 1 is a schematic block diagram illustrating an exemplary voice chat system and communications server apparatus for driving and generating voice responses to user inputs; Figure 2 is a schematic flow diagram illustrating architecture components of a text to speech system for deriving and generating voice responses to either user or other text inputs, in terms of an exemplary iterative flow between the architecture components thereof; Figure 3 is a schematic flow diagram illustrating a real-time speech user flow of an exemplary speech to text to speech system for deriving and generating voice responses to user or other text inputs; Figure 4 is a schematic block diagram illustrating a Speech Recognition module of an exemplary speech to text to speech system for deriving and generating voice responses to user or other text inputs; Figure 5 is a schematic block diagram illustrating the Communication Settings of an exemplary text to speech system for deriving and generating voice responses to user or other text inputs; Figure 6A is a schematic flow diagram illustrating the functions of a real-time speech synthesis flow of an exemplary text to speech system for deriving and generating voice responses to user or other text inputs; Figure 6B is a schematic flow diagram illustrating the functions of an audio management module of an exemplary speech to text to speech system for deriving and generating voice responses to user or other text inputs; Figure 7 is a schematic block diagram illustrating the user communications device inputs and outputs in an exemplary text to speech system for deriving and generating voice responses to user or other text inputs; Figure 8 is a schematic block diagram illustrating some of the aspects of the Generative Model APIs utilised in a text to speech system according to an exemplary embodiment of the invention; Figure 9 is a schematic block diagram illustrating further Communication Settings of an text to speech system according to an exemplary embodiment of the invention; and Figure 10 is a schematic high level flow diagram illustrating the data flow in a text to speech system according to an exemplary embodiment of the invention. Detailed Description In the following description of various exemplary embodiments of the invention, reference is made to the accompanying drawings that form a part hereof, and in which are shown, by way of examples and illustration, specific embodiments in which the invention may be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice the invention, and it is to be understood that other embodiments may be utilized and that variations and modifications may be made, without departing from the scope of the invention as defined in the appended claims. The following detailed description is therefore not to be taken in a limited sense. The specification may refer to “an”, “one” or “some” embodiment(s) in some parts. This does not necessarily imply that each such reference is to the same embodiment(s), or that the feature applies to only one embodiment. Single features of different embodiments may also be combined to make another embodiment. As used herein, the singular forms “a”, “an” and “the” are intended to include the plural forms as well, unless expressly stated otherwise. It will be further understood that the terms “includes”, “comprises”, “including” and / or “comprising”, when used in this specification, specify the presence of stated features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, operations, steps, elements, components and / or groups thereof As used herein, the term “and / or” includes any and all combinations and arrangements of one or more of the associated listed items. Example embodiments of the invention utilize a unique combination of novel speech recognition and speech Recognition modules, together with innovative Communications Settings, communicating, via a backend API, with an LLM or LMM platform, for example, or with a text source such as those used to store pdfs, books, etc, to provide an improved text to speech system facilitating real time voice response functionality. Referring to Figure 1 of the drawings, a computer-implemented voice chat system 100 is illustrated. The voice chat system 100 comprises a communications server apparatus 102 and a user communications device 104. These devices 102, 104 are connected in the communication network 108 (for example, the internet) via communication links 110, 112 implementing, for example, internet communication protocols. The user communications device 104 may be able to communicate through other communications networks, such as public switched telephone networks (PSTN networks), including mobile cellular communication networks, but these are omitted from Figure 1 for the sake of clarity. Communications server apparatus 102 may be a single server, as illustrated schematically in Figure 1, or have the functionality performed by the server apparatus 102 distributed across multiple server components. Indeed, in some example embodiments, the server functionality may be incorporated in the user communications device 104 itself, eliminating the need for the network 108 and communications links 110, 112. In the example of Figure 1, communications server apparatus 102 may comprise a number of individual components including, but not limited to, one or more microprocessors 116, a memory 118 (e.g. volatile memory such as RAM for the loading of executable instructions 120, the executable instructions 120 defining the functionality of the communications server apparatus 102, carried out under control of the processor 116. Communications server apparatus 102 also comprises an input / output module 122 (which may include a transmitter / receiver module) allowing the server to communicate over the communication network 108 (where applicable), receive data representative inputs (from the user communications device 104) and output a response. Communications server apparatus 102 also comprises a database (and other data stores) 126, as well as document libraries 124. In this embodiment, the database (and data stores) 126 and document libraries 124 are part of the communications server apparatus 102, however, it should be appreciated that the database (and data stores) 126 and document libraries 124 can be separated from the communications server apparatus 102, and be connected thereto via the communication network 108 or via another communication link (not shown). Indeed, the entire communications server apparatus 102 could be implemented in the Cloud, rather than on a specific device or set of devices, and the invention is not intended to be limited in this regard. User communications device 104 may comprise a number of individual components including, but not limited to, one or more microprocessors 128, a memory 130 (e.g. a volatile memory such as RAM) for the loading of executable instructions 132, the executable instructions (e.g. in the form of an app) defining the functionality of the user communications device 104, under control of the processor 128. User communications device 104 also comprises an input / output module 134 (which may or may not include a transmitter / receiver module), allowing the user communications device 104 to communicate with the communications server apparatus 102 over the network 108 (if applicable and required). User interface 136 is provided for user control. If the user communications device 104 is, say, a smart phone or tablet device, the user interface 136 may comprise a touch screen panel display, as is prevalent in many smart phone and other hand held computing devices. Alternatively, if the user communications device 104 is, say, a desktop or laptop computer, the user interface may have, for example, computing peripheral devices, such as display monitors, computer keyboards, and the like. User communications device 104 may also comprise a sound module 138, incorporating a microphone and speaker. These may be integrated in the user communications device 104, and / or this functionality may be provided by one or more connected (wired or wirelessly) peripheral devices such as a headset, earphones or headphones. Referring additionally to Figure 7 of the drawings, the user input devices that may be utilized with an example system include a microphone 906 (integrated into the user communications device 104 or as a connected peripheral device), a keyboard or touch pad 907 (either integrated into the user communications device 104 or as a connected peripheral device), a headphone or headset microphone 908, a camera 910 (particularly if a Large Multimodal Model is being utilized), and speakers 909 (either integrated into the user communications device 104 or as connected peripheral devices). The inputs may also include command buttons / functions such as play / pause buttons and remote control command buttons / functions, and a media file input 911. The user output devices that may be utilized with an example system include speakers (either integrated (912a) into the user communications device 104 or as a connected peripheral device 912b), headphones / earphones / telephone headset 913, a screen 914 (either integrated into the user communications device 104 or as a connected peripheral device), a text file output 915, and a media file output 916. Referring now additionally to Figure 2 of the drawings, in an example text to speech system such as a voice chat system 100, the communications server apparatus 102 may be a Large Language Model (LLM) or Large Multimodal Model (LMM) platform, or it may be another source of text such as a pdf or other type of written prose, having an architecture and / or format that will be known to a person skilled in the art, and will therefore not be discussed in great detail herein. It will be appreciated that any desired text source could be used in a system of the invention, depending on the application to which it is required to be adapted, and the present invention is not intended to be limited in this regard. The user communications device 104 may be configured as a ‘chat’ app having a user interface 136 and a backend API 150 for communicating with the text source 102. A user 152 can input a query (in the form of text data, voice data and / or even image data in the case where an LMM server is being utilized) via the user interface 136, or via a remote control function from a peripheral device, such as two-way earphones or a telephone headset, for example. The query input device(s) are denoted generally, in Figure 2, at 154. If the input query is in the form of voice data, it passes first to a Speech Recognition module 200 for generating text data representative of the input query, as described in more detail hereinafter. Data representative of Communication Settings 207, derived from Chat Settings 508, Character Settings 526 and Conversation Settings 548, as will be described in detail hereinafter, is added as a preamble to the text data and the resultant data is transmitted as a query to the LLM platform 102, via the backend API 150. A text response is generated by the LLM platform 102 and delivered back to the user communication device 104, via the backend API 150, and fed to a Speech Recognition module 202, which generates a respective audio (voice) response representative of the text response (again utilizing the Communication Settings 207), as will be described in more detail hereinafter. The audio response is delivered to the user 152 via, for example, the device speaker or headphones / earphones, denoted generally, in Figure 2, at 156. Referring additionally to Figure 10 of the drawings, a high level data flow diagram is illustrated, defining the data flow within an example system, and the interaction between the various modules and functions thereof, as will be described in further detail hereinafter. Referring now to Figure 3 of the drawings, in a voice chat session using the voice chat system illustrated schematically, and by way of example only, in Figure 2 of the drawings, the user can start a ‘conversation’, at step 10, either simply by speaking their query by entering text via a keypad or keyboard. If a session is to be initialized by the user’s voice, there may be a requirement for them to speak a certain word or command before stating their query, or, alternatively, the system could be configured to be always ‘on’ when the microphone is switched on. Once a chat session has been in initialized and the device microphone is active, a Speech Recognition module (to be described hereinafter) performs a speech recognition process (at step 12) and, at step 14, transcribes the speech input, in substantially real time, to produce a corresponding text message output. Alternatively, the user may enter a query simply by entering their text message. In either case, the text message data is sent (at step 16, via a backend API, to the associated LLM platform, together with data representative of pre-selected or pre-configured Communication Settings. The text input or query received by the LLM platform causes the LLM to generate (at step 18) a customized text response (based on the query itself together with the user’s Communication Settings, which specify certain user-specific parameters and preferences, as will be described in more detail hereinafter). The text response is displayed on the user’s device (step 20) and also fed to a real-time Speech Synthesis module (to be described hereinafter) to generate (at step 22), in real-time, a spoken response to the user’s original input query. It is known, in relation to LLMs, to input a text query and obtain a text response. Speech Synthesizers are also known, but these tend to wait for a complete text input before commencing speech synthesis, and then only commence playback of the synthesized speech once the complete text response has been synthesized. This results in unnatural delays. Similarly, chat apps that will respond to a speech query tend to wait until the speech query has been completed before transcribing the speech into a text message for input to the LLM via the backend API, resulting in delays and unnatural ‘conversational’ flow. Furthermore, known systems are unable to personalise, to any great extent, either the LLM response to a query, or the manner in which the query response is presented back to the user. In contrast, an innovative chat application according to an example embodiment of the invention utilizes Large Language Models (LLMs) and / or Large Multimodal Models (LMMs) to provide real-time, dynamic and personalised voice chat conversations. The system employs a unique, multi-layered approach involving realtime speech recognition and transcription, real-time speech synthesis, dynamic conversational frameworks, and character-based personalisation, to significantly enhance user engagement and experience. However, it will be appreciated that the text input may comprise a different text input, such as a media file in the form of a book or other document. Example embodiments may be configured to work for any type of media that can be converted into text for handling by the text-to-speech and audio playback and management functions to be described hereinafter. As illustrated in Figure 2 of the drawings, and as described above, an example voice chat system comprises a Speech Recognition module 200, a Speech Synthesis Module 202, and Communication Settings 206, all interacting with a backend API 208 that provides the user device interface to a selected LLM (or LMM) 210. In addition, the system includes an Audio Management module 204, to ensure that synthesized speech is delivered to the user without interruption or quality degradation, as will be described in more detail hereinafter. A user can initiate an interaction with the system by means of a ‘Start conversation’ function. This could be achieved by pressing a button on their user interface, or on their device, or it could be done by way of a voice command, and the system could incorporate any appropriate means, or even multiple means, to allow a user to initiate a ‘conversation’, and the present invention is not necessarily intended to be limited in this regard. Once the ‘Start Conversation’ function has been activated, it triggers both the user device microphone and the Speech Recognition module concurrently, to both capture (and record) the user’s voice in real time and to perform a substantially real time speech recognition and transcription process to generate text data representative of the user’s speech input. Referring to Figure 4 of the drawings, the Speech Recognition module 22 / 200 has, as its input 100, user speech received at the user device microphone 101. This speech is fed to an Audio Recording Handler function 102, which is configured to manage capture of the user’s speech and convert this raw input into a format suitable for further processing. As soon as the ‘Start Conversation’ function has been activated, the Audio Recording Handler 102 acts to initialize audio recording (block 104), by setting up the necessary parameters and devices for capturing audio data such that, as soon as speech is received from the microphone input 101, the Speech Recognition module 200 can start recording and processing it. Accordingly, the Start Recording function 106 initiates a recording process, capturing the user’s speech in real time. The Audio Recording Handler 102 also incorporates a ‘Pause and Resume Recording’ function 108 and a ‘Stop Recording’ function 110, which may, for example, be activated by selected user inputs or actions, or (in the case of the ‘Stop Recording’ function 110) when the speech input has reached an end. Importantly, a Segmentation function 112 is configured to receive a continuous audio stream representative of the user’s speech input and, in real time, segment it into logical units that mirror natural speech patterns (e.g. phrases, sentences, etc.). This segmented audio data is more manageable than a continuous audio stream, and can therefore be processed more efficiently by the speech recognition function, as it is being received at the microphone input. This segmentation also means that there is no need for the speech recognition function to wait until the entire audio message has been received to process it into logical language patterns - it can be done as the audio message is still being received. Thus, as soon as the user starts to speak, the resultant audio stream at the microphone input 101 is segmented (in real time, as it comes into the Audio Recording Handler function 102) and the segmented audio data is output (at block 11) to a Speech Recognition Handler 116. When the ‘Start Conversation’ function is activated, speech recognition is initialized (block 118) such that, as soon as the Speech Recognition Handler 116 starts to receive segmented audio data from the Audio Recording Handler function 102, it starts speech recognition (at block 120). Speech recognition functions perse (i.e. speech-to-text conversion engines) will be known to a person skilled in the art, and will not be described in detail herein. However, by segmenting the audio data before passing it to the Speech Recognition Handler 116, the speech recognition function 120 can process the audio data in segments in real time, as it comes in from the Audio Recording Handler 102. A ’Handle Recognition Results’ function 128 processes the results from the speech recognition engine and converts the segmented audio data into coherent text, and a ‘Real-time Transcription’ function 130 is configured to display, in real time the resultant text data as it is being generated, providing the user with real time visual feedback, confirming what they have said. The Speech Recognition Handler 116 includes a ‘Pause and Resume’ function 122 and a ‘Stop Recognition’ function 124, which can both be activated by user commands 126, and the ‘Stop Recognition’ function 124 would also be triggered when all of the segmented audio data has been received and processed, and acts to finalize the conversion process and deliver the transcribed text data to the backend API (150, Figure 2). The Speech Recognition Handler 116 may also provide feedback on the status of the recognition process (at block 132). Of course, in some cases, the user may have typed their query / input 111 to the user interface, in which case, the text data so input would be delivered directly to the backend API (150, Figure 2), bypassing the Speech synthesis module altogether. However, in a system configured to cause the user query to be “read back” to them, then it, too, would be passed through the speech synthesis module. In yet other case, the system could be configured to ‘read’, for example, a text process such as a book. Either way, the text data is supplemented by a preamble, comprising data representative of Conversation Model settings, before (in the case of a voice chat system) being passed, by the backend API (150, Figure 2) to the LLM (or LMM). Referring to additionally to Figure 5 of the_drawings, the Communications Settings of an example system are illustrated schematically, and described in more detail below. Communication Settings allow users to customize various parameters, to influence the ‘conversation’ and personalize their use experience of a voice chat session. The nature and variety of the parameters that can be so customized is considered to be unique, and the degree to which the ‘conversation’ can be personalized is considered to be novel and highly innovative relative to known voice chat and LLM / LMM query systems. Furthermore, the Conversation Model within the Communication Settings are adaptive, and a feedback loop (24, Figure 3), incorporated into the user interaction flow, allows various communication settings to be adapted over time, according to a particular user’s interaction with the system. The preamble added to the text message data at the backend API (150, Figure 2) before it is transmitted to the LLM / LMM, is derived from he values of the various parameters and characteristics within a so-called Conversation Model, and acts to enable tailoring of the LLM / LMM responses, and the manner in which they are expressed to the user, according to various user preferences (and past interactions with the system). Accordingly, the Communication Settings 207 comprise a Conversation Model 506, a Chat Model 502, and a Character Model 504. Chat Model (502): Allows a user to customize various parameters or Chat Settings 508, which contribute to a personal user experience by influencing the ‘conversation’. A settings panel, provided at the user interface of the system, allows users to: • Adjust voice parameters: Using, for example, sliders, knobs or dials, users can modify voice attributes like pitch, speed, accent, and dialect. These changes directly influence the Speech Synthesizer Handler’s output (to be described hereinafter). • Select voice options: image selectors or drop-downs, for example, could allow users to choose different voice types, ensuring a more personal and engaging experience. • Date and Time dials: these may be critical if the application has functionalities related to scheduling or reminders. It may not directly influence speech Recognition, but could be part of user preferences. • Save and load preferences: using a UserPreferences structure, the application could store user-defined settings, ensuring a consistent experience across chat sessions. Chat Settings (508): This comprises of a number of ‘modules’ to allow the user to customize parameters such as User Personalization 510, Communication Style 512, Topic Preferences 514, Tone &Format 516, Response Length 518, Response Variation 520, Scheduled Usage 522 and Dominant Hand 524, as described in more detail below, as well User Personalization 525a, Response Accuracy 525b and Closure Sensitivity 525c. User Personalization (525a): The application may allow customization of parameters such as Name, Age and Language through a View Model and stored settings, thereby allowing the user experience to be tailored: • Name: Enables user to set their preferred UserName • Age: The user’s age for tailoring the ‘conversation’ to the right audience • Language: Language of interaction, any accent or dialect parameters Communication Style (512): This module governs the complexity of language and structure of sentences in the ‘conversation’. The user can choose their own conversation options, which could be anything ranging from ‘simple’ to ‘advanced’, again allowing the user experience to be tailored to their own preferences. The example system utilises a sophisticated algorithm to determine the complexity and structure of the language used during ‘conversations’. Users can choose any type of ‘converation’ imaginable, or select from a range of pre-set options, from ‘poetry’ to ‘report’, allowing for a highly customizable linguistic experience. This feature may have particular utility in educational or academic settings, for example, where the app can be customized to adapt to different knowledge levels. Topic Preference (514): A Topic Preference function may enable the user to prioritize certain topics of interest. This affects the direction of the conversation, ensuring that the dialogue remains engaging and relevant to the user. The system employs machine learning to prioritize conversation topics based on user interactions and declared interests. This feature ensures that ‘conversations’ naturally drift toward subjects that are both relevant and engaging for the user, thus improving user retention and satisfaction, Tone &Format (516): The example app also allows users to set the emotional and psychological tone of the ‘conversations’, offering editable options such as ‘friendly’, ‘formal’, or ‘sarcastic’. Users can set this tone of the ‘conversation’ through Character Settings 526 (to be described hereinafter), allowing or a more emotionally resonant and context-appropriate interaction (ranging from, for example, a customer service app to a personal ‘chat’ app), which may be critical, at least for some applications. Response Accuracy (525b): Users can enhance the accuracy of the app’s responses by enabling the system to resend queries to the Language model multiple times. This setting allows for more accurate and well-thought-out responses by leveraging techniques such as Chain of Thought reasoning or Mixture of Experts processing scenarios, improving the overall quality of thee interaction and ensuring that the app meets their informational needs. Response Length - Tokens (518):_Users can adjust the length of the app’s responses, choosing from, for example, a ‘fast’ or ‘thoughtful’ pace. This feature enables the app to ‘fit into’ various ‘conversation’ dynamics, from quick exchanges to more drawn-out, thoughtful responses. Response Variation - Temperature (520): The example system algorithm can either prioritize consistency or introduce variety in its responses, depending on user settings. This feature can either reduce or increase predictability of the ‘conversation’, making it adaptable to different user preferences for conversational flow. Closure Sensitivity (525c): Users can customize how conversations with the Language Model conclude by adjusting Closure Sensitivity settings. Options include formal, personal, or impersonal endings to conversations, among others. This feature ensures that interactions reach a natural and satisfactory conclusion that aligns with the user’s communication style. Scheduled Usage (522): The example app may include built-in timing and scheduling features, such as: • Alarm Clock: Users can set alarms within the Chat Settings interface 508. • Time Limits: Allows the user to set time limitations on "conversations’. • Reminders: Enables a user to set reminders for various tasks. • Adjustable Message Alerts: allows the user to customize message alert timings. Dominant Hand (524): The feature enables the user to set left-handed or right-handed user interface adjustments. Feedback Prompt Settings (590) Referring to Figure 9, the Feedback Prompt Settings (590) framework enhances the interaction between the user and the Language Model (LLM) by serving as an intermediary that refines and extends user inputs to generate better or more complex responses. When enabled, it analyzes the user's input and provides suggestions or enhancements before the input is sent to the LLM. The Feedback Prompt Settings (590) contribute to a more engaging and productive conversational experience by improving the overall quality of interactions with the LLM. Closure Settings (591) ensure that conversations with the Language Model reach a natural and satisfactory conclusion. This module refines feedback prompts to provide complete and conclusive statements, giving users a sense of closure. It works with Context Understanding (592) to maintain conversational coherence while adapting responses based on user interaction patterns. Context Understanding (592) automatically recognizes the context of the user's input to generate relevant and precise feedback prompts through continuous analysis of conversation flow. It identifies key themes or subjects within the user's request to provide focused feedback, working with the Chat Model (502) to ensure appropriate responses. Conversational Brevity (594) creates concise feedback prompts that are easy to understand and respond to, ideal for spoken interactions. This maintains conversation flow while adjusting in real-time through the feedback loop (24) to match user interaction styles. Adaptive Language (596) adjusts the language and tone of feedback prompts to match the user's style—formal, casual, technical, or creative. This personalization makes interactions more natural and engaging, improving user comfort and confidence in the system. Sensitivity and Privacy (597) include filters to avoid repeating sensitive or private topics in feedback prompts, respecting user confidentiality. This module works with security measures to ensure appropriate content delivery, maintaining compliance with age or privacy requirements established in Communication Settings (207). Error Detection and Correction (598) identifies potential errors or ambiguities in the user's input and suggests corrections to spelling or repetition. It implements realtime corrections while maintaining communication through both the Language Model API (150) and Speech Synthesis module, enhancing the effectiveness of feedback prompts. Visual and Auditory Output (599) offers both visual (on-screen text) and auditory (spoken) output for feedback prompts. It integrates with the Speech Synthesis module to ensure synchronized delivery of visual and spoken content, maintaining accessibility standards for all users. Character Model (504): allows a user to set detailed characteristics or ‘Character Settings’ 526 to define the ‘Character’ of the chatbot that is seen to respond to their inputs. Such parameters may include the Unique Identifier 528 of a chatbot, their Name 530, a Description 532 and Biography 534, their Role &Activities 536, their Background 538, a parameter defining their Rapport and Tonality 540, their Character Traits 542, their Voice Traits 544 and the Character Media 546. Some of these Character Settings 526 are explained in more detail below. User ID (528): Each chatbot character has a unique identifier managed by the Character View Model, serving backend functions like tracking interactions and storing character-specific data, and enabling advanced analytics. This ensures that each interaction is correctly attributed to the corresponding character, providing seamless user experience. This User ID 528 serves as a cornerstone for backend operations, mapping interactions and behavioural traits to each character. The Unique ID 528 ensures data integrity and aids in the advanced analytics of usr-character interactions. (Name 530): The character’s display name serves as a user-friendly identifier and appears prominently during user interactions. This feature enhances the personal connection between the character and the user, enriching the conversational experience. The display name serves as the front-end identifier for each chatbot character. This name is what users see during interactions, making it useful in establishing a persona for the chatbot. The name can, in some example systems, be changed dynamically, to give the user a more personalised experience. Description (532): A brief description is provided for each character, offering users an immediate understanding of the character’s personality or role. This description may include ‘hints’ at the characters backstory, interests or expertise, serving as an engaging introduction. Biography (534): Each character may feature a brief backstory that serves, not only as an introductory narrative, but also as an immersive tool for deeper conversations. This backstory can be invoked contextually during ‘conversations’ to create multidimensional and engaging interaction. Role &Activities (536): The role variable defines the character’s function within the conversational framework, for example ‘advisor friend’ or ‘storyteller’. This variable guides the type of responses generated by the example system, ensuring that they are aligned with the character’s predefined role. Character Traits (538) are descriptive attributes managed by the Character View Model, adding layers of ‘personality’ to each character. The character traits are an array of descriptive attributes like ‘friendly’, ‘sassy’, etc., that add nuanced layers of personality to each character. These traits influence the behaviour and responses of the character in the conversation, making them more relatable and engaging. Voice Traits (544) include attributes like pitch, speed, language, and accent, identified by a voiceidentifier variable. This allows each character to have a unique vocal personality, enhancing the auditory experience of the conversation. Communication Model (506): Comprises an adaptive Conversation Settings’ module 548, in which a number of parameters can be set initially, but which can be adapted over time, by the example system, based on the user’s interactions with it. Accordingly, the example system (or app) comprises a dynamic architecture, including a real-time feedback loop (24, Figure 3), which collects and analyses user responses to adjust conversational parameters ‘on the fly’. The example system utilize advanced machine learning algorithms to adapt ‘conversations’ based on user feedback and ongoing context, recognising conversational cues and user sentiment. This feature allows the app to continually refine and optimize the conversational experience, to ensure that the dialogue remains engaging, relevant and contextually appropriate. Thus, utilizing the above-mentioned advanced machine learning algorithms, the example system has the ability to adapt conversations based on ongoing context. This includes recognizing conversational cues, user sentiment, and other contextual factors to steer the conversation in a direction that is both meaningful and engaging for the user. Further, the example system employs memory usage patterns, which enable the app to ‘recall’ and reference past interactions. This feature ensures that conversations are coherent and contextually relevant, thereby enhancing user engagement. The adaptive parameters in the Conversation Settings module 548 may include: Perceptiveness 549, Predictability 550, Perspective 552, Posture 554, Distancing 556, Intent 559, and: Imagination (558): The system incorporates algorithms that influence the levels and types of creative responses generated by the app. This allows the example app to craft unique and imaginative replies, offering a rich and varied conversational experience. Resonance (560) gauges the level of agreement, interest, trust and fulfilment in a ‘conversation’. The example system adapts, in real-time, to align with the user’s sentiment, making the ‘conversation’ more engaging and satisfying. Dissonance (562) identifies potential mismatches in conversational mapping, helping the system avoid topics that might lead to conversational breakdowns. This ensures that the ‘conversations’ remain coherent and meaningful. Conceptual Thinking (564) allows the example app to engage in abstract ‘thought’ processes, allowing for deeper and more open-ended ‘conversations’. This capability enables the system to ‘discuss’ more complex topics, theories and ideas, enriching the conversational experience. Pragmatic Thinking (566) guides the example app to offer practical and actionable responses. This is especially useful for solution-oriented conversations, where the user seeks advice, directions or specific information. As a result of these Conversation Settings 548, the system is able to display intelligent expression when responding to user queries, using Adaptive Conversation facilitated by an Iterative User Feedback Loop (24, Figure 3) that generates feedback updates to the User Interaction Flow (Figure 3). One of the principal innovative features of the example system is its ability to adapt, in real time, based on user feedback and interaction patterns. This adaptive model ensures that the chat experience is continually refined, unlike conventional voice chat apps. This feature also contributes to the adaptability of the system to various different applications and settings. By incorporating principles of human cognition, the example system extends the communication logic into familiar real-world scenarios. It takes into account physical posturing, context and human-like cognitive processes to craft ‘conversations’, that are not only intelligent, but also deeply relatable and human-like. Media Analysis Settings (900) Referring to Figure 9, the Media Analysis Settings (900) framework comprises interconnected processing modules that handle various media types and data formats It serves as the primary analysis pipeline, receiving input through Media Input (902) and coordinating with Media Analysis (904) for on-device processing. It specializes in analyzing and interpreting various content types within media, leveraging advanced techniques to extract and format semantic data, text, tables, spatial content, and more for comprehensive media analysis. It maintains communication with Communication Settings (207) while performing real-time analysis of incoming media content. The analyzed data is then provided to the Media Synthesis Settings (922) as Synthesis Prompts (924), Media Input (902), Categories (930, 932), Media Formats (934), and Artistic Styles (936), enabling personalized media generation based on user preferences. Media Analysis Input (902) accepts real-time media captured or uploaded for analysis, including camera footage, photographs, scanned documents, spatial content, and other data formats. This module starts the analysis process by receiving various media formats that need to be processed and interpreted. These input media can also serve as references for the Media Synthesis Model (920), providing source material for generating new content. Media Analysis (904) works with Vision ML frameworks for advanced processing capabilities, supporting multiple input formats like PDF, DOC, and image files, spatial content, and AV content. The module ensures processing efficiency and accuracy through continuous validation, using on-device machine learning to protect user privacy. The extracted features and metadata are provided to the Media Synthesis Settings (922) as Synthesis Prompts (924) and Input Data, facilitating the generation of personalized media content. Text Content Extraction (908) uses Optical Character Recognition (OCR) technology to extract and format text found within the media, turning visual text into editable and searchable text. The module maintains formatting information while coordinating with Table Data Formatting (910) for structured content processing. The formatted text informs the Media Synthesis Model (920) for text generation and document composition. Table Data Formatting (910) identifies and structure tabular data present within the media, enabling analysis of data previously embedded in media formats. The module maintains cell relationships and header hierarchies while coordinating with Text Content Extraction (908) for complete data preservation. The structured data enables the Synthesis Settings (922) to maintain data relationships in generated content. Document Capture (912) processes and formats data from multiple media files into coherent document formats, like combining scanned pages into a single PDF. It coordinates with Text Content Extraction Functions (908) to maintain document structure while preserving textual and formatting elements. These documents provide reference data to the Synthesis Model (920) for maintaining document fidelity. Image Analysis (906) analyzes and formats the content within media to provide structured semantic data. This involves identifying objects, scenes, and other relevant information using advanced machine learning models. It interfaces with the ML Analysis framework to provide structured semantic data for further processing. The semantic data informs the Media Synthesis Model (920) as Synthesis Prompts (924), Categories (930, 932), or Style Parameters (936), in composing contextually appropriate imagery or integration into other media. Audio Video Analysis (907)_ensures accurate media processing while considering both content and metadata during processing to enable rich media experiences with spatial audio and synchronized playback. It supports both analysis and synthesis of media content by handling: • Media Streams and Encoding: Processes and creates audio and video media files, including various codecs, formats, editing and compositing. • Audio Management: Handles and generates spatial audio positioning, acoustic models, environmental sounds and effects for immersive experiences. • Visual Processing: Analyzes visual quality, and manages high-resolution enhanced visual media with applied effects, filters, and color spaces. • Playback Control: Extracts and manages synchronization data, streaming optimization, quality settings, playback settings and synchronization. • Metadata Handling: Extracts and generates content metadata, temporal markers, interactive elements, and positioning to improve organization and retrieval. This comprehensive media data enables the Synthesis Settings (922) as Input Media, Synthesis Prompts (924), Media Formats (934), and Artistic Styles (936), to maintain audiovisual coherence in generated content. Spatial Analysis (909) processes data in spatial and 3d formats, ensuring cross-platform compatibility and interoperability. It manages spatial data for 3D content and augmented reality experiences, supporting both analysis and synthesis of spatial content. It handles: • Depth Maps and Point Clouds: Captures and Creates detailed geometry and depth information of real-world and generated environments. • Surface Geometry and Material Properties: Analyzes and Synthesizes surface textures, material characteristics, surface and structural properties. • Environmental Mapping: Assesses and Generates lighting conditions, reflections, ambient occlusion, and global illumination. • Position Analysis: Processes and places spatial coordinates, objects, orientation tracking, motion data, and scene hierarchies. • Interaction Spaces: Defines and Generates interactive and gesture zones, tracking volumes, collaboration boundaries, and interaction boundaries • Scene Understanding: Understands object relationships, spatial context, activity zones, and scene semantics for immersive experiences. The spatial data and models enable the Synthesis Model (920) as Input Media, Synthesis Prompts (924), or Categories (930, 932), to generate geometrically accurate 3D content. Screen Analysis (911 ^facilitates features like screen recording analysis, context-aware interactions, and supports data related to windows, sessions, viewers, participants, and context parameters. It manages screen content and user interface analysis through systematic state monitoring, supporting both analysis and synthesis of display content. It handles: • State Management: Monitors and manages window contexts, session states, view hierarchies, and interface elements for seamless interactions. • Interaction Analysis: Tracks user engagement metrics, focus areas, event processing, and interaction patterns; adapts interfaces based on user and environmental factors. • Experience Flow Tracking: Monitors progress markers, completion states, participant interactions, conversation dynamics, and supports collaborative activities like screen sharing and meeting recordings. Interface patterns inform the Synthesis Settings (922) as Synthesis Prompts (924) or Input Data, in creating responsive display elements. Formatted Media Data (914) manages and structures additional data extracted from the media, such as metadata or contextual information, enhancing the depth of analysis or method of synthesis. This contextual data guides the Synthesis Settings (922) in adapting generated content with Synthesis Prompts (924), Categories (930, 932), or Style Parameters (936), enabling more personalized and context-aware media generation. Media Presentation (918) provides multiple output formats for the Media Synthesis Model (920) and the Media Analysis API (580) based on Communication Settings (207) preferences. This module works with Adaptive Analysis to adjust output detail and format according to user requirements, ensuring the extracted data is accessible and useful for both chat interactions and APIs. Biometric ID (Secure User Profiles) (916) integrates device-level security features into the Media Analysis Settings to ensure secure access and processing. Using the on-device framework, it supports various biometric authentication methods: • Facial Recognition uses device cameras for secure access via facial biometrics. • Fingerprint Authentication employs device sensors for quick and secure verification. • Iris Scan offers access using iris scanning for high-security or spatial applications. • User ID &Password Integration allows authentication using User ID and password. These biometric security measures collectively safeguard sensitive media data and send validation data to the Media Analysis API (580), ensuring that further media analysis of synthesis is conducted securely. Biometric data and user authentication status can influence the Media Synthesis Settings (922), allowing personalized content generation based on user identity and access levels. Security and Compliance (917) integrates comprehensive security controls into the Media Analysis Settings to ensure secure content processing and validation; using a layered security approach, it supports various protection mechanisms: • Quality Management: Controls audit trails and timestamps by applying digital signatures and secure key storage for content validation. (Data points: audit trails, timestamps, digital signatures, secure keys.) • Data Processing Validation: Ensures content integrity during operations by maintaining attributable, legible, contemporaneous, original, and accurate records. (Data Points: integrity checks, ALCOA records, validation.) • Runtime Operations: Manages authenticated access through multi-factor verification and session controls while monitoring process isolation. (Data points: logs, multi-factor data, timeout settings, isolation metrics.) • Post-Processing Security: Handles secure record retention and backup management in both viewable and electronic formats. (Data points: records, backups, formats, metadata.) • Communication Security: Coordinates protected data transmission between components while binding signatures to records. (Data points: Encrypted communication channels, bound digital signatures, transmission logs, secure API protocols.) These security measures collectively ensure compliant content processing across regulated industries while maintaining necessary audit trails and documentation for international standards. The implementation sends validation data to the Media Analysis API (580), confirming the secure analysis or synthesis of all processed content. Security measures ensure that data provided to the Media Synthesis Settings (922) complies with regulatory standards, maintaining data integrity and confidentiality during synthesis processes. Referring to Figure 9, the Media Synthesis Settings (922) framework includes multiple interconnected modules that control media generation and analysis, offering users extensive control over their media synthesis experiences: Media Synthesis Input (921) accepts processed or direct media for synthesis operations, including analyzed content from Media Analysis Settings (900) and raw input media that bypasses analysis. This module serves as the entry point for the Media Synthesis Model (920), receiving various media types including analyzed images, documents, AV media, biometric data, and other formats. The input content can come either from Media Analysis Settings (900) results or directly from Media Analysis Input (902) through the bypass pathway, providing source material and reference data for media generation processes. Media Synthesis Model (920) allows users to customize various parameters within the Media Synthesis Settings (922), contributing to a personalized user experience by influencing the generated media content. The system features a dynamic architecture with a real-time feedback loop (24, Figure 3), utilizing advanced machine learning algorithms, the Media Synthesis Model (920) adapts media generation based on user feedback and ongoing context, recognizing user preferences and usage patterns. Synthesis Feedback Prompts (924) allow users to adjust feedback prompts for media, streamlining communication. This module interfaces directly with the standard Feedback Prompt Settings (590) in Communication Settings (207) to maintain consistency with user preferences while enhancing information delivery through visual and auditory channels. Synthesis Response Length (926) ensures media responses are engaging and concise when spoken. This setting keeps conversations flowing and avoids overwhelming users with lengthy dialogue. The system automatically adjusts response length based on interaction context, user preferences, and the current conversation state maintained through the feedback loop (24). Spoken Media Descriptions (928) optionally integrate with the Speech Synthesis module to provide audio descriptions and context for media content. This adapts dynamically based on user preferences, enhancing accessibility for visually impaired users and offering detailed descriptions for those seeking deeper understanding. Predefined Categories (930) offer standardized content classification options, including photography, painting, drawing, and charts. These categories interface with the ML Analysis framework to ensure appropriate processing parameters are applied to each content type generated by the Media Synthesis API (586). Custom Categories (932) let users create personalized classification systems for media content, tailoring media suggestions and responses to their unique interests or projects. These categories integrate with the Media Analysis Settings (900) to govern how content is processed, stored, and presented by the Media Synthesis API (586), aligning with user preferences in Communication Settings (207). Media Format (934) provides options for various file types, as well as landscape, portrait, or other media formats based on specific project requirements or personal preferences. This ensures proper format handling throughout the processing pipeline while maintaining data integrity. Artistic Styles (936) offer a range of visual effects and styles for users to apply to generated media, from classic to contemporary, enhancing creative expression. Detail Adjustment (938) gives users granular control over processing depth and output refinement, allowing them to set the level of detail in generated media. This module communicates with both the Media Analysis Settings (900) and the Media Synthesis API (586) to adjust processing parameters in real-time based on preferences and content needs. ML Analysis (940) provides an option to activate further analysis for media before it is uploaded, coordinating with the Media Analysis (904) framework for advanced processing. The module ensures processing efficiency and accuracy through continuous validation. Data Presentation (942) lets users choose how they receive extracted data, catering to different learning styles and accessibility needs. This module works with Adaptive Analysis (944) to adjust output detail and format according to user requirements. Adaptive Analysis (944) implements dynamic adjustments of the media analysis based on user input, offering a tailored approach for both casual and professional interactions, enhancing user satisfaction. Download Data (946) provides immediate access to newly extracted data during analysis, crucial for users needing instant information for decision-making before further processing. Referring additionally to Figure 8 of the drawings, the Text API Settings and Data Language Model (or Backend) API 150 defines the LLM (or LMM) Settings 568, such as User ID 570, URL 572, API Key 574, Message, Model, Temperature, Character Description, History Management, Request Customization (headers), Date Formatting, and error handling. Such LLM or LMM settings may be specific to the LLM or LMM being used, and will be familiar to a person skilled in the relevant art. As such, these settings will not be discussed in further detail herein. Referring to Figure 6A of the drawings, the Speech Synthesis (Realtime Speech Synthesis) module 202 receives, as its input data from the Backend API 150, a text data stream representative of the LLM / LMM response to the user input. The text data stream is prefixed by a data preamble representative of data points from the Chat Model 502, data points from the Character Model 504, data points from the Communication Model 506 and data points from the Backend API 150. This input data is fed into a Speech Synthesizer Handler 602 (or Speech Generation API), as described in more detail below. Speech Synthesizer Handler - Speech Generation API (602): This is tasked with the conversion of text-based responses into audible speech, enriching the dynamic and engaging conversation experience. The ‘Speech Synthesizer Handler’, together with the associated Speech Emphasis Handler 910, functions are at the heart of converting text-based responses into audible speech. Speech Synthesizer Handler 602, together with the Speech Emphasis Handler 910, convert text-based responses from the LLM or LMM (or text received from elsewhere) into lifelike, engaging speech. Generated text from the LLM or LMM, and pre-processed by the Speech Emphasis Handler 910, is fed into this function, which then utilizes various voice settings such as pitch, rate, and timbre for speech Recognition. The real-time speech output is a lifelike, engaging speech that aligns with the specific character's defined voice traits (from the Communication Settings (207, Figure 5). This speech aligns with all of the specific voice traits defined in the Character Settings (526, Figure 5). Speech Emphasis Handler - API Speech Handling (910): Referring to Figure 6A of the drawings, a text data stream 582 is received into a Speech Emphasis Handler 910. The text data stream 582 may be derived from, for example, the Language Model APIs (150, Figures 2, 8) (in response to a user input or query, or the output of the Speech Recognition function), but alternatively could comprise text derived from a media file, such as a book. The Speech Emphasis Handler 910 is configured to pre-process the incoming text data stream 582 by breaking it into logical units that mirror natural speech patterns and include, at 615, pre-processing data to take account of issues such as phrasal handling 622, detection of filler words 617, detection of intentional duplicate words 619, phrase detection 621, phenome handling 614, pronunciation correction 616 and semantic analysis 623. The pre-processed text is then passed in segments or “chunks” to the Speech Synthesizer Handler 602. Key Unique Functionalities: Detects Filler Words (617): Identifies common filler words (like "urn", "uh", "like", "you know") and decide how you want to handle them. Inserts an adjustable and dynamic pause after a filler word to mimic natural speech patterns. Also has an option to skip filler word entirely. Detects Intentional Duplicates (619): Looks for repeated words or phrases that are next to each other and considers them as intentional for emphasis. Treats these duplicates differently, by adding a pause between them to emphasize the repetition, or by altering the prosody to convey the intended emphasis. Phrase Detection (621): Extends the tokenization to detect noun phrases, verb phrases, and such, using the NLTagger to detect parts of speech and then apply simple heuristic rules to group words into phrases. Semantic Analysis (623): Use the NLLanguageRecognizer to determine the dominant language of the text and potentially aid in understanding the context for the chatbot character customization. For deeper semantic analysis, this can connect to a machine learning model or an NLP service that provides semantic analysis. Pronunciation Correction (616): Beyond basic text-to-speech Recognition, the handler contains algorithms to correct and refine pronunciation, ensuring the output sounds natural. This ensures uncommon or complex words are pronounced correctly. The system may accesses a prebuilt phonetic dictionary, and as well uses algorithmic methods to predict the pronunciation of words, with NLP libraries for parsing larger complex combinations of words and phrases. Phoneme Handling (614): For intricate words or names, the handler breaks down the text into individual sound units (phonemes) to ensure accuracy in pronunciation. Data Transformation: Text fragment ("synthesis") -> Phonemes ("s", "i", "n", "0", "e", "s", "i", “s") Inputs: • Text data, typically in a string format, generated by the Large Language Model (LLM), Large Multimodal Model (LMM) or Al Chatbots. • The speech engine can also be used to read the users input messages when a user, recording software, or audience may want both sides of the chat spoken. • Voice settings (e.g., pitch, rate) based on user preferences or character attributes and in the form of parameters or configuration objects. • User commands to control synthesis, usually as events or triggers, such as, start, pause, resume, or stop synthesis. • Apply Pre Processing Delays to Chunks (615): After determining where the emphasis should be, this applies these delays to the processed chunks of text data to be spoken. Utilizing these pre-processing functions on the text data stream as it starts being received into the Realtime Speech Synthesis Module 202, and thereby enabling the text data to be broken up into natural / intuitive “chunks” of text, means that the “chunks” of text can be immediately passed to the Speech Synthesizer Handler 602, to be converted into audio (speech) data, while the rest of the text data stream is still being received and pre-processed. Ultimately, what this means is that, unlike conventional text-to-speech applications, he present application is able to perform the txt to speech functionality ‘on the fly’, as the text is coming in, and without having to wait for it to finish in some predetermined fashioned (e.g. by a full stop or end of paragraph symbol) before commencing the speech conversion. Speech Synthesizer Handler (602): • Initialize Speech Recognition (604): Prepares the necessary libraries and APIs to begin the text-to-speech conversion. The system may load specific voice profiles or configurations based on user preferences or character settings. • Synthesize Speech (606): At this stage, the handler converts the provided “chunks” of text data (received from the Speech Emphasis Handler 910, into raw audio data, sounding like lifelike speech. The system uses phonetic mappings and predefined voice settings to generate this audio. The output is a continuous waveform representing the speech, usually in a digital format such as PCM (Pulse Code Modulation). The input text may be a text data stream 582 generated by an LLM platform in response to a user query, or it may be a text data stream 582 generated by a Media Analysis module 580 (associated with an LMM platform) in response to an input image. If the user wants the system to create an image or other media in response to their input, then the output from the Media Analysis module 580 may comprise multiple (e.g. four) parallel text responses 584 from which the user can select (or the system may be configured to select one automatically based on the Communication Settings 207). In the latter case, the selected data stream is delivered to a media synthesis module 586 to generate the associated media for output on the user communications device 104. Data Transformation: Text ("Hello, World!") -> Digital Audio Data (PCM waveform) Key • Pause and Resume Synthesis (608): Offers the ability to temporarily halt and then continue the speech Recognition process, ensuring the user retains control over the conversation flow. • Stop Synthesis (610): Ends the synthesis process, either upon completion of the text or upon user request. • Voice Selection and Control (612): This feature contains functionality to select and control different voice options, allowing for a customizable user experience. This allows users to customize the voice's attributes, towards such variables as pitch for gender, language for accent, and filters for “roboticness”, etc. The synthesized speech will adjust its characteristics based on adjustments to these settings. Data Input: User voice preferences (e.g., "Female voice, British accent, More robotic”) In a Speech Phrasal Handler receiving synthesized speech from the Speech Synthesizer Handler 602, the following functions may be provided: • Inflection and Modulation (618): The synthesized speech is not monotonous.. The handler uses NLP and ML libraries from Apple to interpret the context and sentiment of the text, adding necessary pitch changes, stress patterns, rhythm, inflections and modulations based on the context and sentiment of the text to make the speech sound more natural and expressive. These are customized approaches not often used to make speech apps. • Emphasis Delays (625): Based on the phrases' positions and semantic values, the program decides on the delays. For instance, it can add some longer pauses after concluding significant phrases or before starting important statements. • Speed and Pitch Control (620): Depending on user settings or character attributes, the handler can adjust the speed (rate) and pitch of the synthesized speech. This function likely considers various voice settings not typically used in speech applications, such as dynamically adjusting the pitch, rate, and timbre, to include more voices and promote more user satisfaction. Data Points and Logic Systems of Realtime Speech Synthesis: • Text Data: The primary input for the Speech Synthesizer Handler is the text data. This could be a user's query, a response from the Large Language Model, or any other textual data. • Voice Settings: Using the settings panels, users can adjust parameters such as pitch, rate, accent, and dialect. This customization ensures that the synthesized speech aligns with a specific character's attributes or user preferences. Settings such as accent and dialect can significantly change the pronunciation of words, ensuring that the output is consistent with regional variations. • Synthesis Engine: At the heart of the handler, the synthesis engine takes in the text and the voice settings to produce audio data. The engine ensures that the speech mimics human nuances, factoring in inflections, modulations, and pronunciation corrections. • Phoneme Handling: For complex words or those not natively recognized by the synthesis engine, the handler breaks down the text into phonemes, ensuring accurate pronunciation. User Interactions: • Users can initiate, pause, resume, or stop the synthesis process. • Through the settings panels, users can adjust voice parameters, enabling a personalized experience. So, these API calls are not just network calls, they also refer to using any internal Frameworks, Libraries and Models for Image, Text, Voice and Lyrics, as they all have their own on device APIs which use similar functionality. When these models become small enough, they may be used locally without network connections, making this architecture network independent, which would include a wider degree of privacy, increased speeds, greater accessibility, and much lower costs to operate. As each “chunk” of text data is converted to speech data, it is passed to the Audio Buffer Player 650. Audio Buffer Player (650): The output of the Speech Phrasal Handler of the Synthesizer Handler 602 is fed to an Audio Buffer Player 650, which also has as inputs User commands for playback control and other additional data streams. The Audio Buffer Player 650 serves as a real-time audio processing unit, ensuring that the synthesized speech is delivered to the user without interruptions or quality degradation. It acts as a bridge between the synthesized audio data and the device's audio output systems. The Audio Buffer Player 650 manages the audio buffer for real-time processing, ensuring smooth and glitch-free audio output. It constantly manages the incoming text stream and the outgoing audio stream, ensuring that audio data is effectively processed and ready for playback. By intelligently managing buffer sizes, it minimizes latency and ensures a smooth audio output, enhancing real-time interactivity. Key Functionalities: • Initialize Audio Buffer (704): Initializes the audio buffer for real-time processing, which sets up the buffer structure to hold real-time audio data. This defines the size of the buffer and prepares it for data streaming. Data Structure: Circular Buffer or FIFO (First In First Out) Queue • Buffer Management (706): This script is responsible for managing the real-time segmented audio data 705 in a buffer. It constantly manages an audio buffer, ensuring that audio data is effectively processed (block 713) and ready for playback. It also intelligently decides the size and management of this buffer to ensure smooth audio output. As audio data arrives, it's stored in the buffer. The system ensures that there's always a steady stream 711 of audio data available for playback, minimizing underflows or overflows (block 712). Data Transformation: Continuous Digital Audio Data -> Segmented Audio Chunks in Buffer • Preprocessing (708): This phase enhances the audio data for playback. Before the audio data is played, it undergoes a preprocessing phase. This ensures that the audio stream is ready for real-time playback, optimized for human hearing and understanding. (It might involve equalization, noise reduction, or dynamic range compression to ensure optimal listening quality.) • Pronunciation and Inflection Handling (710): While this is primarily a function of the synthesizer, the buffer player can apply last-minute corrections or adjustments based on real-time feedback or user inputs. The audio buffer also takes care of some pronunciation nuances. It ensures that the audio, when played back, mimics natural human speech in terms of pronunciation, inflection, and pacing. • Realtime Playback (714): This involves streaming the buffered audio data to the device's speakers. It ensures the audio plays without glitches or delays, and that the audio data in the buffer is played back in real-time, with minimal latency. This also makes sure the playback of the audio buffer is controllable with the device's connected audio systems, like headphone buttons and speaker controls. • Interrupt Handling (716): The buffer can handle interruptions, pausing the audio when needed, and resuming it seamlessly. In case of network interruptions, or incoming calls, mixing with other applications audio or handling user-triggered pauses, the buffer player can temporarily halt playback and later resume from the same point. • Buffer Overflow and Underflow Management (712): This functionality ensures smooth audio playback. If the buffer is too full (overflow), the system might discard or compress some data. If the buffer runs out of data (underflow), it might pause playback until more data arrives. This manages potential issues like buffer overflow (too much data) or underflow (too little data) to maintain a consistent audio experience. Inputs: • Realtime audio data streams, typically as digital audio packets or chunks. • Synthesized speech data from the Speech Synthesizer Handler 602. • User commands for playback control such as play, pause, or stop, usually as events or triggers. Outputs: • Processed audio data 713 ready for playback. • Continuous audio stream 711, usually sent directly to the device's audio hardware. • Feedback signals 716, like buffer status or playback position, typically in the form of flags, counters, or timestamps, often indicating the status of the audio buffer (e.g., buffer full, buffer empty). Data Points and Logic Systems: • Audio Data Stream: The continuous flow of audio data, either from real-time sources or from the Speech Synthesizer Handler. • Buffer: The buffer holds chunks of audio data, ensuring smooth playback. The buffer size can be dynamic, adapting based on the data flow rate and ensuring minimal latency. • Preprocessing: Before playback, the audio data undergoes preprocessing. This step might involve equalizing the audio to optimize it for the device's speakers or enhancing it for better human comprehension. • Playback Management: This system handles the real-time playback of audio data, catering to user commands like play, pause, or stop. It also manages interruptions, ensuring seamless user experience. User Interactions: • Users can control the playback using play, pause, stop, or other commands. • Through the settings panels, users can adjust audio parameters like volume or balance. Speech FX Handler (627): A sophisticated module designed to enhance the user experience by integrating both audio and visual effects that change in color and shape according to the character. This dual feedback mechanism offers users a visual feedback loop from both their own voice and the speech synthesis, reducing cognitive load and creating a more intuitive and satisfactory experience. This feature significantly improves accessibility by offering visual cues for speech, making the app more inclusive for users with hearing impairments. Speech Effects (628): This feature allows selection from a range of effects such as equalization adjustments, echo, reverb, ring modulation, or vocoders. By customizing these audio effects, users can make conversations more engaging and immersive, enhancing the overall audio quality. Visual Speaking Indicators (629): This feature offers a visual representation of speech activity, alerting users to ongoing speech from both themselves and the app. It enhances engagement and accessibility by providing immediate visual cues, especially beneficial for users with hearing impairments. These indicators reduce cognitive load by offering a visual feedback loop, making the interaction more intuitive and satisfying. Feedback on Synthesis Process (630): Users receive immediate visual feedback on the speech synthesis process through Feedback on Synthesis Process (624). It improves the overall interactive experience by keeping users informed in real-time, thereby reducing uncertainty and enhancing trust in the app's responses. Synthesized Speech Data (631): The app efficiently manages and processes synthesized speech data through Synthesized Speech Data (626) to ensure seamless audio playback and visual representation. It ensures that both the auditory and visual elements of speech are synchronized, enhancing the overall quality of the interaction. Audio Manager (718): Referring additionally to Figure 6B of the drawings, he Audio Manager 718 takes, as its inputs, processed audio data from the Audio Buffer Player 650 and any user interactions 720, and is responsible for managing real-time audio output through the device's audio systems. These functions take the processed audio buffer and manage its playback through the device's audio output systems. Through smart buffering and state management, these functions enable glitch-free, real-time audio playback. Key Functionalities: • Initialize Audio (804): Prepares the audio systems for playback, ensuring optimal settings for real-time interactions. • Manage Audio Playback (806): Handles the initiation, pause, resume, and stop of audio playback. • Handle Remote Control Events (808): Processes events from remote controls like play, pause, and skip. • Voice Selection and Control (810): This file holds functionality for selecting different voice options Inputs: • Processed audio data from the audio buffer player 650. • User interactions (like play, pause, or skip commands) 720. Outputs: • Real-time audio output 812 to the device's speakers. • Feedback 814 on the current state of audio playback (e.g., playing, paused). Bluetooth Manager (816): This is related to Bluetooth Low Energy (BLE) management. While this file might not directly contribute to the speech-to-text and text-to-speech functionalities, it could play a role in how the app interacts with external devices or accessories. Key Functionalities: • Initialize BLE (818): Prepares the application for BLE communications, ensuring that the device's Bluetooth capabilities are available and functional. • Scan for Devices: Searches for available BLE devices in proximity. • Connect and Disconnect: Provides the ability to establish or terminate connections with external BLE devices. • Data Transfer (819): Facilitates the transfer of data between the app and connected BLE devices. Inputs: • Commands to start or stop scanning, connect to a device, or send / receive data. • Data received from external BLE devices. Outputs: • Feedback on the status of BLE operations (e.g., device connected, data sent). • Data to be sent to external BLE devices. Remote Control (820): This is dedicated to handling remote control functionalities, which might be crucial for user interactions during real-time audio playback. Key Functionalities (822): • Play Command: Processes the command to initiate audio playback. • Pause Command: Halts the audio playback temporarily. • Resume Command: Resumes paused audio playback. • Skip Command: Allows users to skip to a specific portion of the audio stream. • Interrupt Playback: An important function that enables users to interrupt ongoing audio playback, either by pressing a button or through other interactive controls. This ensures the user can intervene and change the course of the conversation as needed. Inputs: • User interactions, such as pressing play, pause, resume, skip, or interrupt commands. Outputs: • Commands to the audio management system to execute the desired actions. • Feedback on the status of the remote control operations. It will be appreciated, from the foregoing description, that the example system utilizes a unique approach to real-time, dynamic, and personalized ‘conversations’, facilitated by advanced Large Language Models or Large Multimodal Models. The system or app incorporates unique user-centric settings and innovative technologies to offer an engaging and realistic conversational experience. With features like realtime transcription, instant speech feedback and audio buffering, the app is distinguished from, and highly innovative relative to, traditional voice chat applications. The example system utilizes a multi-layered approach that synergistically combines immediate speech and synthesis with dynamic conversational formats and personality-based characters (or ‘robots’). These features enable the system to be adapted for many different applications and settings (depending on the LLM or LMM to be used), and enable it to facilitate engaging and realistic ‘conversations’. The above-described system utilizes a speech emphasis handler 602 for text “chunking, audio buffering, a speech recognition handler, incoming speech visualisation, outgoing listening visualization, and audio management to realise a unique real-time text to speech chat system, that overcomes the drawbacks associated with known text to speech conversion systems. The unique features of various embodiments of the invention features lie in its unique approach to real-time, dynamic, and personalized conversations, facilitated by advanced Large Language Models. The app incorporates unique user-centric settings, and innovative technologies to offer a truly engaging and realistic conversational experience. With features like real-time transcription, instant speech feedback, and audio buffering, embodiments of the invention set themselves apart from traditional chat applications and text to speech applications. Embodiments of the invention employ a Multilayered Approach that synergistically combines immediate speech and synthesis with dynamic conversational formats and personality-based robots. This innovative approach sets the app apart, making it exceptionally effective in facilitating engaging and realistic conversations. It will be apparent to a person skilled in the art, from the foregoing description, that modifications and variations can be made to the described embodiments without departing from the scope of the invention as defined by the appended claims.

Claims

1. A computer-implemented system for facilitating a text to speech application, comprising a user communications device configured to communicate with a Large Language Model (LLM) or Large Multimodal Model (LMM) API, the user communications device comprising:a processor;a memory having instructions stored therein to be executed under control of the processor;an input device for receiving text data from a text data source; andan audio playback device for outputting an audio response;the user communications device being configured, under control of the processor, to execute instructions stored in the memory to:receive a text data stream from a text data source;in real time, as soon as said text data stream has started to be received, use a speech emphasis handler module to segment said incoming text data into text data segments based on natural speech patterns and / or emphasis, and pass the text data segments, in order, as they are created, to a speech recognition module to transform each said text data segment into representative digital audio data;in real time, as said representative digital audio data is generated, start feeding said digital audio data, as a continuous audio data stream, to an audio buffer and store it therein; anddeliver, in real time and in order, said digital audio data to said audio playback device.

2. A system according to claim 1, further comprising an audio buffer player configured to segment the continuous digital audio data stream, as it is being received into the audio buffer, into digital audio segments based on natural speech patterns.

3. A system according to claim 1 or claim 2, wherein the speech emphasis handler module is configured to receive and dissect text data and break down sentences into phrases.

4. A system according to claim 3, wherein said speech emphasis handler module is configured to utilise heuristic rules to group words into phrases.

5. A system according to any of the preceding claims, wherein said speech emphasis handler module is configured to receive text data and identify filler words therein.,6. A system according to any of the preceding claims, wherein said speech emphasis handler module is configured to receive text data and identify duplicated words, phrases or sentences therein.

7. A system according to any of the preceding claims, wherein said speech emphasis handler module is configured to receive text data and identify therein emotion and tone elements.

8. A system according to any of the preceding claims, wherein said speech emphasis handler module configured to receive text data and identify therein natural delays or pauses.

9. A system according to any of the preceding claims, wherein said playback device comprises a speech phrasal handling module configured to receive said audio segments and identify and incorporate emphasis delays in said audio data as it is being played back, in use.

10. A system according to any of the preceding claims, wherein said playback device comprises a speech phrasal handling module configured to receive said audio segments, and apply inflection and modulation methods to add pitch changes and stress patterns to said audio data as it is being played back, in use.11 .A system according to any of the preceding claims, wherein said playback device comprises a speech phrasal handling module configured to control the speed and pitch of the audio data as it is being played back, in use.

12. A system according to any of the preceding claims, further comprising a customisation module configured to, in use, allow a user to customise one or more characteristics of the audio data as it is being played back, in use.

Citation Information

Patent Citations

  • Speech synthesis method and device, computer equipment and storage medium

    CN114765022A

  • Text-to-voice processing method and device, equipment and storage medium

    CN115223541A

  • Method and system for text-to-speech synthesis of streaming text

    US20230335111A1

  • Text-to-speech synthesis method, electronic device, and computer-readable storage medium

    US20230410791A1