Method for generating customized automatic speech recognition transcript and electronic device performing the method

The electronic device personalizes ASR transcripts by extracting audio features and recipient context, addressing the limitations of current voice assistants by providing contextually relevant and user-specific responses.

WO2026105983A1PCT designated stage Publication Date: 2026-05-21SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SAMSUNG ELECTRONICS CO LTD
Filing Date
2025-03-28
Publication Date
2026-05-21

AI Technical Summary

Technical Problem

Current voice assistant systems lack personalization and flexibility to accommodate individual user preferences and unique speech patterns, leading to frustration and inefficiencies, particularly for users with distinct speaking styles or accents, as they do not adapt to changing language patterns or command preferences over time.

Method used

An electronic device generates customized automatic speech recognition (ASR) transcripts by extracting audio features, retrieving recipient context from conversation history, and using a formality vector to adjust interaction flow, capturing both generic formality and user idiosyncrasies through modules like audio feature extraction, recipient-specific context retrieval, and behavioral customization engines.

Benefits of technology

The solution enables personalized and contextually relevant responses, enhancing user experience by adapting to individual user preferences and conversational dynamics, ensuring responses align with the user's intended tone and style.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025004001_21052026_PF_FP_ABST
    Figure KR2025004001_21052026_PF_FP_ABST
Patent Text Reader

Abstract

An electronic device and method for generating customized automatic speech recognition (ASR) transcripts are provided. The method involves extracting audio features received from an audio input and retrieving a recipient-specific context based on conversation history with the recipient. Further, the method involves generating a first customized ASR transcript based on audio features, the recipient context, and a formality vector. A second customized ASR transcript is generated, incorporating both the first customized ASR transcript and the unique behavior of the individual user in relation to the message recipient.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD FOR GENERATING CUSTOMIZED AUTOMATIC SPEECH RECOGNITION TRANSCRIPT AND ELECTRONIC DEVICE PERFORMING THE METHOD

[0001] The present disclosure relates to the field of speech recognition. In particular, the present disclosure relates to a method for generating a customized automatic speech recognition (ASR) transcript and an electronic device performing the method.

[0002] Voice assistant technology has rapidly evolved, allowing users to interact with electronic devices through voice commands to accomplish various tasks. These virtual assistant platforms, commonly referred to as Intelligent Personal Assistants (IPAs) or Automated Virtual Assistants, enable users to retrieve information, provide customer support, and assist in managing daily activities. By facilitating hands-free operation, IPAs enhance accessibility and user convenience in digital environment. Various IPA implementaions offer different capabilities, leveraging advanced speech recognition, maching learning, and artificial intelligence techniques to provide personalized services and inprove user experience.

[0003] The functionality of intelligent systems depends on their ability to accurately recognize and respond to user inputs. These assistants utilize natural language processing (NLP) and machine learning to interpret speech, retrieve information, and execute commands. They can deliver news updates, track stock prices, provide weather forecasts, manage schedules, and assist with retail purchases. Their capabilities can be enhanced by leveraging contextual data, such as user location, past interactions, and external data sources.

[0004] Despite such advancements, current voice assistant systems rely on standardized command structures designed for broad user adoption. While this approach enables seamless interaction for most users, it lacks flexibility to accommodate individual preferences and unique speech patterns. Users must often conform to predefined phrases, limiting natural interaction. This lack of personalization can cause frustration and inefficiencies, especially for those with distinct speaking styles or accents. Repeating or rephrasing commands to obtain desired results undermines the intuitive and convenient nature of voice-enabled technology.

[0005] The rigid structure of these systems does not account for evolving user interactions. For example, voice assistants do not adapt to users' changing language patterns or command preferences over time. This limitation is particularly evident when command phrasing varies due to regional dialects, colloquialisms, or personal speech habits. Current technology also lacks the ability to personalize responses or adjust interaction flow based on a user's conversation history or unique engagement. As a result, voice assistants may generate contextually irrelevant responses, reducing overall user experience.

[0006] Therefore, there is a need for a technology that can address or mitigate one or more of the issues in existing solutions.

[0007] According to an embodiment of the disclosure, there is provided a method performed by an electronic device for generating a customized automatic speech recognition (ASR) transcript. The method may include extracting one or more audio features associated with a received audio input from a user. The method may include retrieving a recipient context from one or more messages of a recipient based on conversation history with the recipient. The method may include generating a first customized ASR transcript based on at least one of the one or more audio features, the recipient context, and a formality vector. The first customized ASR transcript may capture the generic formality of the message. The method may include generating a second customized ASR transcript based on at least one of the first customized ASR transcript and user behavior. The second customized ASR transcript may capture idiosyncrasies of the user with the recipient of the message.

[0008] According to an embodiment, there is provided an electronic device for generating a customized automatic speech recognition (ASR) transcript. The electronic device may include a memory configured to store at least one instruction a computer program. The electronic device may include at least one processor configured to execute the at least one instruction to perform extracting one or more audio features associated with a received audio input from a user. The at least one processor is configured to execute the at least one instruction to perform retrieving a recipient context from one or more messages of a recipient based on conversation history with the recipient. The at least one processor is configured to execute the at least one instruction to perform generating a first customized ASR transcript based on at least one of the one or more audio features, the recipient context, and a formality vector. The first customized ASR transcript captures the generic formality of the message. The at least one processor is configured to execute the at least one instruction to perform generating a second customized ASR transcript based on at least one of the first customized ASR transcript and user behavior. The second customized ASR transcript captures idiosyncrasies of the user with the recipient of the message.

[0009] To further clarify the advantages and features of the present disclosure, a more particular description of the disclosure will be rendered by reference to an embodiment thereof, which is illustrated in the appended drawings. It is appreciated that these drawings depict only typical embodiments of the disclosure and are therefore not to be considered as limiting its scope.

[0010] In order to more clearly explain the technical solutions in an embodiment of the disclosure, the accompanying drawings to be used in the description of the embodiment of the disclosure will be briefly described below.

[0011] FIG. 1illustrates an environment for generating a customized automatic speech recognition transcript, according to an embodiment of the present disclosure;

[0012] FIGs. 2a-2f illustrate a detailed architecture for generating the customized automatic speech recognition transcript, according to an embodiment of the present disclosure;

[0013] FIG. 3illustrates a block diagram of a system for generating a customized automatic speech recognition transcript, according to an embodiment of the present disclosure;

[0014] FIG. 4illustrates a flowchart depicting a method for generating a customized automatic speech recognition transcript, according to an embodiment of the present disclosure;

[0015] FIG. 5 illustrates a scenario for generating a customized automatic speech recognition transcript, according to an embodiment of the present disclosure; and

[0016] FIG. 6 illustrates a scenario for generating a customized automatic speech recognition transcript, according to an embodiment of the present disclosure.

[0017] Reference will be made in detail to example embodiments of the present disclosure, examples of which are illustrated in the drawings, wherein like reference numerals refer to like elements throughout the drawings.

[0018] The terms used in the present disclosure will be briefly described, and then embodiments of the present disclosure will be described in detail.

[0019] The terms used in the present disclosure are general terms as much as possible and have been widely used nowadays in consideration of the functions of the present disclosure, which, however, may be changed according to an intention of a technician in the art, a precedent, the advent of new technologies, or the like. Also, particular cases may include terms arbitrary selected by an applicant, and in this case, the meaning of the terms will be described in detail in the corresponding description. Therefore, the terms used in the present disclosure should be defined based on the meanings of the terms and the content throughout the present disclosure, rather than simply based on the titles of the terms.

[0020] It may be advantageous to set forth definitions of certain words and phrases used throughout this disclosure. Thus, the terms "transmit," "receive," and "communicate," as well as derivatives thereof, encompass both direct and indirect communication. The terms "include" and "comprise," as well as derivatives thereof, mean inclusion without limitation. The term "or" is inclusive, meaning and / or. The phrase "associated with," as well as derivatives thereof, means to include, be included within, interconnect with, contain, be contained within, connect to or with, couple to or with, be communicable with, cooperate with, interleave, juxtapose, be proximate to, be bound to or with, have, have a property of, have a relationship to or with, or the like.

[0021] Moreover, various functions described below may be implemented or supported by one or more computer programs, each of which is formed from a computer readable program code and embodied in a computer readable medium. The terms "application" and "program" refer to one or more computer programs, software components, sets of instructions, procedures, functions, objects, classes, instances, related data, or a portion thereof adapted for implementation in a suitable computer readable program code. The phrase "computer readable program code" includes any type of computer code, including source code, object code, and executable code. The phrase "computer readable medium" includes any type of medium capable of being accessed by a computer, such as read only memory (ROM), random access memory (RAM), a hard disk drive, a compact disc (CD), a digital video disc (DVD), or any other type of memory. A "non-transitory" computer readable medium excludes wired, wireless, optical, or other communication links that transport transitory electrical or other signals. A non-transitory computer readable medium includes media where data can be permanently stored and media where data can be stored and later overwritten, such as a rewritable optical disc or an erasable memory device.

[0022] As used here, terms and phrases such as "have," "may have,", "include," or "may include" a feature (like a number, function, operation, or component such as a part) indicate the existence of the feature and do not exclude the existence of other features. Also, as used here, the phrases "A or B," "at least one of A and / or B," or "one or more of A and / or B" may include all possible combinations of A and B. For example, "A or B," "at least one of A and B," and "at least one of A or B" may indicate all of (1) including at least one A, (2) including at least one B, or (3) including at least one A and at least one B. Further, as used here, the terms "first" and "second" may modify various components regardless of importance and do not limit the components. These terms are only used to distinguish one component from another. For example, a first user device and a second user device may indicate different user devices from each other, regardless of the order or importance of the devices. A first component may be denoted a second component and vice versa without departing from the scope of this disclosure.

[0023] It will be understood that, when an element (such as a first element) is referred to as being (operatively or communicatively) "coupled with / to" or "connected with / to" another element (such as a second element), the element can be coupled or connected with / to the other element directly or via a third element. In contrast, it will be understood that, when an element (such as a first element) is referred to as being "directly coupled with / to" or "directly connected with / to" another element (such as a second element), no other element (such as a third element) intervenes between the element and the other element.

[0024] As used here, the phrase "configured (or set) to" may be interchangeably used with the phrases "suitable for," "having the capacity to," designed to," "adapted to," "made to," or "capable of" depending on the circumstances. The phrase "configured (or set) to" does not essentially mean "specifically designed in hardware to." Rather, the phrase "configured to" may mean that a device can perform an operation together with another device or parts. For example, the phrase "processor configured (or set) to perform A, B, and C" may mean a generic-purpose processor (such as a CPU or application processor) that may perform the operations by executing one or more software programs stored in a memory device or a dedicated processor (such as an embedded processor) for performing the operations.

[0025] The terms and phrases as used here are provided merely to describe some embodiments of this disclosure but not to limit the scope of other embodiments of this disclosure. It is to be understood that the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. All terms and phrases, including technical and scientific terms and phrases, used here have the same meanings as commonly understood by one of ordinary skill in the art to which the embodiments of this disclosure belong. It will be further understood that terms and phrases, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and will not be interpreted in an idealized or overly formal sense unless expressly so defined here. In some cases, the terms and phrases defined here may be interpreted to exclude embodiments of this disclosure.

[0026] Examples of an "electronic device" according to embodiments of this disclosure may include at least one of a smartphone, a tablet personal computer (PC), a mobile phone, a video phone, an e-book reader, a desktop PC, a laptop computer, a netbook computer, a workstation, a personal digital assistant (PDA), a portable multimedia player (PMP), an MP3 player, a mobile medical device, a camera, or a wearable device (such as smart glasses, a head-mounted device (HMD), electronic clothes, an electronic bracelet, an electronic necklace, an electronic accessory, an electronic tattoo, a smart mirror, or a smart watch). Other examples of the "electronic device" include a smart home appliance. Examples of the smart home appliance may include at least one of a television, a digital video disc (DVD) player, an audio player, a refrigerator, an air conditioner, a cleaner, an oven, a microwave oven, a washer, a dryer, an air cleaner, a set-top box, a home automation control panel, a security control panel, a TV box (such as SAMSUNG HOMESYNC), a smart speaker or speaker with an integrated digital assistant (such as SAMSUNG GALAXY HOME), a gaming console, an electronic dictionary, an electronic key, a camcorder, or an electronic picture frame. Still other examples of the "electronic device" include at least one of various medical devices (such as diverse portable medical measuring devices (like a blood sugar measuring device, a heartbeat measuring device, or a body temperature measuring device), a magnetic resource angiography (MRA) device, a magnetic resource imaging (MRI) device, a computed tomography (CT) device, an imaging device, or an ultrasonic device), a navigation device, a global positioning system (GPS) receiver, an event data recorder (EDR), a flight data recorder (FDR), an automotive infotainment device, a sailing electronic device (such as a sailing navigation device or a gyro compass), avionics, security devices, vehicular head units, industrial or home robots, automatic teller machines (ATMs), point of sales (POS) devices, or Internet of Things (IoT) devices (such as a bulb, various sensors, electric or gas meter, sprinkler, fire alarm, thermostat, street light, toaster, fitness equipment, hot water tank, heater, or boiler). Other examples of the "electronic device" include at least one part of a piece of furniture or building / structure, an electronic board, an electronic signature receiving device, a projector, or various measurement devices (such as devices for measuring water, electricity, gas, or electromagnetic waves). Note that, according to various embodiments of this disclosure, the "electronic device" may be one or a combination of the above-listed devices. According to some embodiments of this disclosure, the "electronic device" may be a flexible electronic device. The electronic device disclosed here is not limited to the above-listed devices and may include new electronic devices depending on the development of technology.

[0027] In the following description, electronic devices are described with reference to the accompanying drawings according to various embodiments of this disclosure. As used here, the term "user" may denote a human or another device (such as an artificial intelligent electronic device) using the electronic device.

[0028] Definitions for other certain words and phrases may be provided throughout this patent document. Those of ordinary skill in the art should understand that in many if not most instances, such definitions apply to prior as well as future uses of such defined words and phrases.

[0029] Throughout the present disclosure, expressions such as "at least one of a, b or c", "at least one of a, b, or c", "at least one of a, b and c", "at least one of a, b, and c" may indicate "a," "b," "c," "a and b," "a and c," "b and c," "all of a, b, and c," or variations thereof. Similarly, expressions such as "at least one of a or b", "at least one of a, or b", "at least one of a and b", "at least one of a, and b " may indicate "a," "b,", "a and b," or variations thereof.

[0030] In this disclosure, the expression "and / or" includes a combination of a plurality of described components or any component of the plurality of described components. In this disclosure, terms such as "1st," "2nd," "first," and "second" may be merely used to distinguish a corresponding component from other corresponding components and do not limit the corresponding components in terms of other aspects (for example, the degree of importance or the order).

[0031] Throughout the present disclosure, when a part "includes" or "comprises" an element, the part may further include other elements, not excluding the other elements, unless there is a particular description contrary thereto.

[0032] Also, the terms "portion," "module," etc. described in this disclosure denote a unit configured to process at least one function or operation, and the "portion," and the "module" may be realized as hardware or software, such as a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC), or a combination of the hardware and the software. The term "portion" used in an embodiment of the present disclosure does not have a meaning limited to software or hardware. A "portion" described in the present disclosure may be configured to be in a storage medium which may be addressed or may be configured to play one or more processors. According to an embodiment of the present disclosure, a "portion" may include components, such as software components, object-oriented software components, class components, and task components, processes, functions, attributes, procedures, sub-routines, segments of a program code, drivers, firmware, a microcode, a circuit, data, a database, data structures, tables, arrays, and variables. Functions provided through a predetermined component or a predetermined "portion" may be combined to reduce the number of functions or may be divided into additional components. Also, according to an embodiment, a "portion" may include one or more processors.

[0033] According to an embodiment of the present disclosure, each of blocks of the flowcharts and combinations of the flowcharts may be performed by computer program instructions. The computer program instructions may be loaded on a general-purpose computer, a specialized computer, or a processor of other programmable data processing device. The instructions performed through the computer or the processor of the other programmable data processing device may generate a medium for performing the functions described in the flowchart block(s). The computer program instructions may also be stored in a computer-available or computer-readable memory oriented for the computer or the other programmable data processing device in order to realize the functions in a predetermined way. The instructions stored in the computer-available or computer-readable memory may also produce manufacturing items embedding an instruction medium for performing the functions described in the flowchart block(s). The computer program instructions may also be loaded on the computer or the other programmable data processing device.

[0034] Furthermore, each block in the flowcharts may represent a module, segment, or a part of code including one or more executable instructions to perform particular logic function(s). According to an embodiment of the present disclosure, the functions described with respect to the blocks may also be generated not according to an order. For example, two blocks illustrated in succession may be executed substantially concurrently or the blocks may sometimes be executed in a reverse order, depending on the functions involved therein.

[0035] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings, so that the embodiments of the present disclosure may be easily implemented by one of ordinary skill in the art. However, the embodiments of the present disclosure may have different forms and should not be construed as being limited to the embodiments of the present disclosure described herein. Also, in the drawings, parts not related to descriptions are omitted for the clear description of the embodiments of the present disclosure, and throughout the specification, like reference numerals are used for like elements.

[0036] Throughout the present disclosure, the term "electronic device" may be referred to as a system that includes functions for transmitting and / or managing transcripts generated through automatic speech recognition.

[0037] FIG. 1illustrates an environment 100 for generating a customized automatic speech (ASR) recognition transcript, according to an embodiment of the present disclosure. In an embodiment of the present disclosure, the electronic device and method for generating a speech recognition transcript may be referred to as a device (or a system) for performing a speech recognition (SR) function and an SR method.

[0038] The environment 100 may include a user 102a, an electronic device (or a user device) 102b, a first recipient device 104a, and a second recipient device 106a. The user device 102b, the first recipient device 104a, and the second recipient device 106a are connected to one another over a network.

[0039] The user 102a may provide an audio input 102 to the electronic device 102b. The audio input 102 may be a message which is to be sent to the first recipient device 104a or the second recipient device 106a. For example, if the user 102a is driving a vehicle and is unable to send a message to the first recipient device 104a, the user 102a may provide the audio input 102 such as, "Hi, I am driving. Talk to you later". The electronic device 102b may be configured to recognize the audio input 102 and convert the audio input 102 into a transcript and transmit the message including the transcript to the first recipient device 104a. The first recipient device 104a may receive the message sent by the user 102a via the user device 102b as "Hi, I am driving. Talk to you later" Here, the transcription process takes place without any personalization, producing a direct, word-for-word transcription of the user's audio input for transmission to a first recipient device 104a.

[0040] In addition, the user 102a may provide the audio input 102 to the electronic device 102b, intending to send a message to a different recipient, for instance, a second recipient device 106a. The user 102a is driving a vehicle and is unable to send a message to the second recipient device 106a. The user 102a may provide the audio input 102 such as "Hi, I am driving. Talk to you later." The electronic device 102b is configured to recognize the audio input 102 and convert the audio input 102 into a text transcript, which the text transcript then sends to the intended recipient, the second recipient device 106a. However, in this case, the transcription process may include a personalized modification, resulting in a message received by the second recipient device 106a that may include abbreviations and informal language, such as "Hi, m driving. Ttyl." It may include shorthand and informal language, showing that the electronic device may apply a layer of personalization to the transcription based on contextual factors, such as the user's relationship with the recipient or prior interaction patterns. This example represents a personalized transcription approach that differs from verbatim transcription.

[0041] Referring FIG. 1, the electronic device 102b may extract at least one audio feature from the received audio input 102, as shown in 105a. The electronic device 102b may retrieve recipient-specific context, as shown in 105b. For example, the electronic device 102b may retrieve recipient-specific context from the electronic device 102b, or an external device (for example, a server). The electronic device 102b may perform formality-aware automatic speech recognition (ASR), as shown in 105c. The electronic device 102b may analyze a user's behavior by using a behavioral customization engine, as shown in 105d. The electronic device 102b may generate recipient-specific customized transcripts and then transmit the generated recipient-specific customized transcripts to the second recipient device 106a, as shown in 105e.

[0042] FIG. 2a illustrates a detailed architecture 200 implemented in the electronic device 102b for generating the customized automatic speech recognition transcript, according to an embodiment of the present disclosure.

[0043] The architecture 200 may include an input module 202, an audio feature extraction module 204, a recipient-specific context retrieval module 206, a formality cognizant engine 208, a behavioral customization engine 210, and a recipient-specific customized output module 212. Since the architecture 200 may be implemented in the electronic device 102b, it may be considered that the electronic device 102b includes the input module 202, the audio feature extraction module 204, the recipient-specific context retrieval module 206, the formality cognizant engine 208, the behavioral customization engine 210, and the recipient-specific customized output module 212.

[0044] The input module 202 and the audio feature extraction module 204 may be configured to handle preprocessing and feature extraction. The input module 202 may be configured to receive the audio input 102 from the user 102a. The audio input 102 may include the user's voice, which is recorded as a waveform and subsequently processed for analysis. The audio input 102 may include various acoustic features, such as pitch, tone, and frequency, and may contain words (or spoken words), phrases, or non-verbal sounds. Once captured, the raw audio input or data undergoes preprocessing steps like filtering, noise reduction, and feature extraction to enhance its quality and suitable for tasks like automatic speech recognition, command interpretation, or other processing tasks.

[0045] Upon receiving the audio input 102 from the user 102a, the audio feature extraction module 204 may process the audio input 102 to output windowed and framed mel-spectrogram features, supporting a set number of frequency bins. To achieve this, the audio feature extraction module 204 may be configured to utilize a conventional audio feature extraction pipeline based primarily on Mel-frequency Cepstral Coefficients (MFCCs).

[0046] The process may begin with framing, where the audio input 102 is divided into short, overlapping frames. Each frame may typically be 30 milliseconds long with 50% overlap between consecutive frames. A Hamming window may be applied to each frame to reduce spectral leakage as the signal transitions to the frequency domain. The audio-feature extraction module 204 may use a Discrete Short Time Fourier Transform (STFT) to convert each frame into its frequency components.

[0047] The audio-feature extraction module 204 may be configured to use triangular filters spaced on the Mel scale to capture frequency components aligned with human auditory perception, producing filterbank energies to extract audio features. These energies are then logarithmically scaled to compress dynamic range, followed by a Discrete Cosine Transform (DCT) to isolate the most significant coefficients. The resulting DCT coefficients represent the Mel-Frequency Cepstral Coefficients (MFCCs), which serve as the extracted audio features, effectively capturing the essential characteristics of the user's voice command.

[0048] The recipient-specific context retrieval module 206 may be configured to gather relevant contextual information about the second recipient device 106a based on conversation history within a specific application. The recipient-specific context retrieval module 206 may function as a database retrieval and maintenance system, where it stores and retrieves past interactions with the recipient to provide a conversation context for the current exchange. The recipient-specific context retrieval module 206 may be configured to receive information about the recipient and search stored conversation history data, allowing it to access context data specific to prior interactions. The retrieved conversation historical data may enable formality analysis, allowing the electronic device 102b to gauge and adjust the tone or style of interaction.

[0049] The recipient-specific context retrieval module 206 may be configured to compute formality vectors i.e., summarized representations of formality levels in the conversation history data based on recipient-specific context. The formality vectors may be updated periodically to reflect the latest context data ensuring the response style remains aligned with any evolving conversational dynamics with the recipient.

[0050] For example, consider a customer support chatbot designed to assist users with account inquiries on a banking app. When a user, let's call her Sarah, initiates a conversation, the recipient-specific context retrieval module 206 may be configured to access Sarah's past interactions with the chatbot from the conversation history data. The conversation history data may be stored in a conversation history database (not shown), which may reside within the electronic device 102b or / and an external device. For example, if Sarah previously engaged in formal inquiries, such as balance checks and account statements, the electronic device 102b may compute a formality vector indicating a more professional tone. The formality vector is stored and pre-computed based on Sarah's prior messages, meaning the chatbot doesn't need to check the conversation history data in real time for every message she sends. Instead, the chatbot can quickly apply the stored formality setting, ensuring a consistent and formal tone throughout the conversation.

[0051] The formality cognizant engine 208 may be configured to generate a first customized ASR transcript based on at least one of the one or more audio features, the recipient context, and a formality vector. The first customized ASR transcript may capture the generic formality of the message. The first customized ASR transcript may be referred to as a first customized speech recognition (SR) transcript.

[0052] The formality cognizant engine 208 may be configured to process the first customized ASR transcript to align with the formality level of the conversational context. The formality cognizant engine 208 may receive both the audio features of the audio input of the user (or a voice command of the user) and context details of the conversation or application to generate a first customized ASR transcript that matches the desired formality level. The formality cognizant engine 208 may be configured to adjust the first customized ASR output based on a "formality priming vector" which is a computational vector derived from the recipient-specific context retrieval module 206, which can include parameters like vocabulary style, emoji usage, and the overall tone of the conversation. The first customized ASR transcript may be included in an automatic speech recognition (ASR) output.

[0053] For example, in a high-formality setting with traditional vocabulary and no emoji usage, the formality cognizant engine 208 may produce a conventional, formal ASR output, such as "Hi, How are you doing?" In contrast, for low-formality settings with non-traditional vocabulary and high emoji usage, the ASR output becomes more informal, such as ", hw r u doing". To achieve this flexibility, the formality cognizant engine 208 may use a fine-grained annotation of context.

[0054] For example, instead of general descriptors like "high" or "low" for emoji usage, the formality cognizant engine 208 may assign a continuous numerical range (e.g., [-1, 1]) to quantify the context elements precisely. An ASR model is statistically trained on various conversational contexts, meaning that the engine applies the most common adjustments for each formality level but may not capture individual user preferences. The output customization relies on statistical averages, ensuring that it broadly aligns with formal or informal contexts as understood by the ASR model, while ignoring minor idiosyncrasies. This approach allows the ASR model to adapt effectively to different conversation settings, providing an output that feels natural and appropriate to the context.

[0055] The formality cognizant engine 208 may include a formality vector extraction module 210 and a formality cognizant ASR module 212, but the present disclosure is not limited thereto. The formality vector extraction module 210 may include a pre-trained formality prediction AI model 210a, a lexical feature extraction module 210b, and a vector encoding 210c, as shown in FIG. 2b, but it is not limited thereto. The lexical feature extraction module 201b may be referred to as a lexical feature extraction engine.

[0056] FIG. 2b illustrates a structure of a formality vector extraction module according to an embodiment of the present disclosure.

[0057] Referring to FIG. 2b, the pre-trained formality prediction AI model 210a may comprise several key steps, including dataset procurement, tokenization, training of static word embeddings, and the application of masking and windowing techniques. These steps may work together to produce a formality-sensitive word embedding model capable of distinguishing between formal and informal language.

[0058] The dataset procurement step may be crucial for effective training of the word embedding model, a diverse dataset that captures a broad range of formality levels in language is required. This dataset is curated to include both formal and informal text sources to represent real-world usage patterns. Formal texts, such as Wikipedia articles and newspaper content, provide structured, grammatically correct sentences and formal vocabulary. Informal texts, gathered from social media posts, chat logs, and instant messages, introduce slang, abbreviations, emojis, and a more casual tone. The pre-trained formality prediction AI model 210a with the combination of these types of data, may learn the nuances of both formal and informal language styles, which is critical for creating embeddings that reflect the varied language encountered in real conversations. This diverse corpus allows the model to develop embeddings that are not biased towards any single language style, thus reducing issues with out-of-vocabulary (OOV) words, especially for informal expressions such as emojis.

[0059] The tokenization process may involve a wordpiece tokenization technique designed to handle both formal and informal language, including emojis, in a granular way. Instead of relying on simple word-level tokenization (which would break down text by spaces and punctuation), this approach uses sub-word tokenization, where words are divided into smaller units resembling morphemes. This method enables better handling of vocabulary variations, especially with informal or mixed content that includes abbreviations, slang, and emojis.

[0060] FIG. 2c illustrates an example of vocabulary creation in wordpiece tokenization according to an embodiment of the present disclosure.

[0061] As shown in FIG. 2c, the tokenization may begin by extracting initial character unigrams from the text corpus (e.g., "Hi, how are you doing?"), creating an initial vocabulary from single characters (like "h," "a," "y," etc.) (① of FIG. 2c). In the next step, pairs of characters or units may be scored based on a frequency-based formula, which is the frequency of the pair divided by the product of the frequencies of each element (② of FIG. 2c). High-scoring pairs may be then combined into a new unit in the vocabulary. For example, common letter pairs such as "hi" or "ar" may have high scores and be added to the vocabulary as single tokens (③ of FIG. 2c).

[0062] This iterative process of pairing and scoring may continue until the vocabulary reaches a predefined size (④ of FIG. 2c). Through this approach, a tokenizer may build a vocabulary that is flexible and inclusive of word variations, emojis, and other informal language elements. When tokenizing new text, a tokenizing algorithm may perform a greedy search through the vocabulary to match the longest tokens first, accelerating the process with a Trie data structure. This detailed tokenization approach may create a vocabulary robust enough to handle the nuances of both formal and informal language contexts in conversation.

[0063] The training of static word embeddings for the pre-trained formality prediction AI model 210a may involve a continuous bag of words (CBOW) approach. In this approach, the pre-trained formality prediction AI model 210a is trained to predict a target word based on the target word's surrounding context, which is represented by words within a specific "window" around the target word. Each word in the vocabulary may be represented as a unique vector, and through training, each word's vector may be adjusted to capture the semantic relationships between words.

[0064] In CBOW training, an input may consist of one-hot encodings of context words (words surrounding the target), and an output may be a prediction of the target word. For each word pair (context and target), the pre-trained formality prediction AI model 210a may use a two-layer neural network to learn which words frequently occur together in various contexts. During training, a softmax activation function and categorical cross-entropy loss may be employed to optimize the model's weights. The result is a set of static embeddings where each word has a fixed vector representation that captures each word's meaning based on each word's usage across the corpus. Since the dataset includes a variety of formal and informal texts, the embeddings will reflect formality levels in word usage.

[0065] Masking and windowing are techniques applied during the training process to manage the context provided to the pre-trained formality prediction AI model 210a. The windowing may refer to the selection of a fixed number of words surrounding a target word to define the target word's context. For example, with a window size of 3, the sentence "How are you doing?" would generate training samples such as "How <MASK> you" and "are <MASK> doing." Here, the <MASK> token may represent the target word that the model must predict based on the other words in the window.

[0066] Masking, in this context, involves replacing the target word in each training sample with a placeholder (e.g., <MASK>) so the pre-trained formality prediction AI model 210a learns to predict missing words based on the surrounding context. By repeatedly presenting the pre-trained formality prediction AI model 210a with masked words within various windows of context, it may learn to infer meaning from limited surrounding words. This approach strengthens the embeddings by forcing the model to consider the relationships between words rather than memorizing their positions in specific sentences. Through the combination of windowing and masking, the pre-trained formality prediction AI model 210a may become adept at predicting word meanings in different contexts, a critical skill for handling variations in formality. These embeddings may be then used to generate formality-sensitive vectors, allowing the pre-trained formality prediction AI model 210a to adjust pre-trained formality prediction AI model 210a's outputs based on the formality detected in the input text.

[0067] Referring to FIG. 2b, the lexical feature extraction module 210b is designed to analyze the surface-level linguistic properties of text to identify features associated with formality or informality in conversation. The lexical feature extraction module 210b may directly analyze text at the word level and morpheme level to extract measurable and intuitive features that explicitly indicate formality. This approach analyzes text characteristics by focusing on easily identifiable elements such as emojis, internet slang, and out-of-vocabulary terms, which are commonly found in informal communication.

[0068] The feature extraction process by the lexical feature extraction module 210b may begin by analyzing the text's components to identify specific markers of informal or formal language use.

[0069] One key feature the lexical feature extraction module 210b extracts is the emoji constant, which measures the proportion of emojis in the text. Emojis are identified using Unicode code points, which uniquely represent each emoji across platforms. The lexical feature extraction module 210b may be configured to count these emoji instances and then normalize the count by dividing it by the total number of words in the text, producing a metric that reflects the intensity of emoji usage relative to text length. Since emoji usage is commonly associated with a casual or informal tone, a higher normalized Emoji Constant may indicate a lower level of formality.

[0070] Another key feature extracted by the lexical feature extraction module 210b is the out of vocabulary (OOV) constant, which may be configured to detect the presence of abbreviations, slang, and other non-standard terms. The OOV constant may identify words in the text that do not appear in a formal vocabulary or dictionary, which typically includes only standard language usage. This helps capture instances of informal expressions, such as "u" for "you" or "btw" for "by the way," which are often excluded from formal vocabularies. The OOV count is normalized by the total word count, providing a ratio that indicates the density of non-standard language in the conversation. This metric serves as an indicator of informality when the proportion of OOV terms is high.

[0071] The lexical feature extraction module 210b may be configured to calculate a lingo constant, which measures the usage of internet slang or culturally specific terms whose meanings may vary by context. The lingo constant is derived using a predefined dictionary of internet slang terms, some of which may appear in formal vocabulary but have different meanings in informal contexts (e.g., "drip" or "cap" in modern slang). To quantify this feature, the lexical feature extraction module 210b may employ a weighted scoring system. For each word in the text, a binary value (Li) is assigned: "1" if the word is recognized as lingo and "0" otherwise. Each lingo term may be assigned a weight (Wi) within a range of [0,1], based on each lingo term's frequency of use in informal contexts, with higher weights that may be given to words strongly associated with informality. The final lingo score, or the lingo constant, may be calculated as a weighted average. In this calculation, the presence of informal slang terms may increase the score, which may indicate a more casual tone.

[0072] The lexical feature extraction module 210b may provide a comprehensive profile of the text's formality based on tangible lexical elements by aggregating these three constants, i.e., the emoji constant, the OOV constant, and the lingo constant. This approach to lexical feature extraction is particularly useful in real-time applications, where distinguishing between informal language and formal language can improve the quality and appropriateness of automated responses.

[0073] The vector encoding 210c serves as a bridge between the formality feature extraction process and the downstream ASR (automatic speech recognition) model by converting extracted features into a unified, multidimensional vector format. This vector, known as the formality vector, may be configured to integrate the various metrics of formality derived from both primary and secondary features to prepare the electronic device 102b for interpreting the conversation's formality level.

[0074] The primary features in this encoding process may include core metrics identified earlier, i.e., the formality polarity score, the emoji constant, the OOV constant, and the lingo constant. Each of these features may reflect different aspects of the conversation's tone.

[0075] For example, the formality polarity score may provide an overall measure of formality based on the linguistic context, while the emoji constant, the OOV constant, and the lingo constant may provide more specific insights into the level of casualness or informality, as evidenced by the use of emojis, slang, and abbreviations.

[0076] The secondary features may be configured to capture nuanced relationships between the primary features. These are calculated through weighted averages of positively correlated pairs and triples of primary features. For example, if high values for both the OOV constant and the lingo constant often indicate a highly informal context, a composite feature combining these two values with a specific weight can help the ASR model better understand contexts where informality is strongly indicated by unconventional language patterns. By aggregating these feature combinations, the model enhances the encoding's expressiveness and relevance for varied conversational contexts.

[0077] The final output of this process is the formality vector, a single, structured representation of all formality-revealing features. The formality vector effectively packages the contextual information needed to "prime" the downstream ASR model. When the formality vector is fed into the ASR model, the formality vector helps the electronic device 102b interpret spoken input according to the identified formality level, thereby enabling more context-sensitive transcription and response generation. The ASR model may be referred to as a speech recognition (SR) model. The Vector Encoding 201c ensures that all relevant formality features are synthesized and organized for maximum utility in adjusting the ASR's processing behavior to fit the conversation's tone.

[0078] FIG. 2d illustrates a structure of a formality cognizant automatic speech recognition (ASR) module according to an embodiment of the present disclosure.

[0079] Referring to FIG. 2d, the formality cognizant automatic speech recognition (ASR) module 212 may be configured to extract at least one of the one or more audio features of the received audio input and the formality vector for a prior conversation. The formality vector for the prior conversation is concatenated to a prediction window of one or more prediction windows. For example, the formality cognizant ASR module 212 may be designed to integrate formality considerations into an automatic speech recognition (ASR) process, producing outputs that are customized to match the conversational tone.

[0080] The formality cognizant ASR module 212 may be configured to receive both the pre-computed formality vector, which encapsulates the conversational context's level of formality and the acoustic audio features derived from the audio input (or input speech) 102. The Formality Cognizant ASR module 212 may be composed of several key components which may include framewise multiplexing of features 212a, a formality primed ASR model 212b, and a token-to-vocabulary mapper 212c.

[0081] The framewise multiplexing of features 212a may be configured to prepare the audio input (or the audio input data) 102 for model inference by aligning the acoustic features, which are recorded at a frame level, with the formality vector, which remains constant throughout the utterance. To accommodate this mismatch, the formality vector may be replicated for each prediction window within an utterance, ensuring that contextual formality information is continuously integrated at every prediction step.

[0082] The formality primed ASR model 212b may be configured to process the multiplexed features. In an embodiment, the formality primed ASR model 212b may be an artificial intelligent model built as a neural network. The formality primed ASR model 212b may be structured to function as a sequence-to-sequence (or transduction-based) system, where acoustic sequences are transformed into token sequences. A distinctive feature of the formality primed ASR model 212b is the formality primed ASR model 212b's hierarchical audio encoder, which processes inputs at both frame and prediction-window levels. The formality primed ASR model 212b allows for window-based processing, with output tokens representing morphemes rather than individual letters, which can better capture meaning and context. The formality primed ASR model 212b may operate an auto-regressive manner, allowing the formality primed ASR model 212b to fuction in streaming mode. The formality primed ASR model 212b may provide real-time output for each window without depending on future context. The inclusion of a "formality encoder" allows the formality primed ASR model 212b to adjust its language model based on the intended conversational tone, distinguishing the formality primed ASR model 212b from traditional ASR systems.

[0083] Finally, the token to vocabulary mapper 212c may be configured to transform the sequence of predicted token IDs into the final ASR transcript (or final SR transcript). A mapping process of the token to vocabulary mapper 212c not only assigns vocabulary elements to each token ID but also handles the collapsing of repeated tokens, i.e., a common characteristic of transduction-based ASR systems. The token to vocabulary mapper 212c may operate like an integer-to-string conversion mechanism, where each token is associated with its respective vocabulary element, and sequences are segmented by special non-emitting characters to delineate new tokens. An example of the token to vocabulary mapping is shown in FIG. 2e.

[0084] FIG. 2e illustrates an example of token-to-vocabulary mapping according to an embodiment of the present disclosure.

[0085] Referring to FIG. 2e, repeated tokens (4, 5, φ, as shown in FIG. 2e) are collapsed, non-emitting characters (φ, as shown in FIG. 2e) are handled, and the remaining token sequence (4, 5, 5, 6, as shown in FIG. 2e) is mapped to corresponding vocabulary entries to generate the final text output (Hi wassup").

[0086] The above steps show generating a first customized ASR transcript based on at least one of the one or more audio features, the recipient context, and a formality vector. The first customized SR transcript may capture the generic formality of the message.

[0087] FIG. 2f illustrates a structure of a behavioral customization engine according to an embodiment of the present disclosure.

[0088] Referring to FIG. 2f, the behavioral customization engine 213 may include a behavioral rule generator 214a, a rule support framework 214b, and a behavioral customization application 214c.

[0089] The behavioral rule generator 214a is a critical component of the behavioral customization engine 214 which may be configured to capture user-specific language preferences and idiosyncrasies based on their interactions with the ASR (or SR)-generated output. The behavioral rule generator 214a may take input from the feedback provided by the user on the customized text, which may involve corrections or modifications made to better align the output with the user's unique linguistic style. The behavioral rule generator 214a may identify two types of behavioral rules, i.e., replacement rules and addition rules, each corresponding to specific user-customized edits.

[0090] The replacement rules may be used to map phrases or terms in the ASR(or SR) transcript to preferred expressions specified by the user. For example, the user may prefer the shorthand "howru" in place of "hw r u." To create replacement rules, the behavioral rule generator 214a may analyze the smallest linguistic units, such as individual words or phrases, and establish mappings without making finer adjustments at the morpheme level. Addition rules may involve inserting new elements or expressions, such as emojis or additional words, based on the audio input (or a user input). For example, the user may add an emoji to convey tone, as shown in the example where the thumbs-up emoji is added after "Hi."

[0091] The behavioral rule generator 214a may be configured to compare the ASR (or SR) output with the user-confirmed text, locating the points at which the strings initially diverge and where they reconverge for both replacement and addition rules. Two pointers traverse the ASR (or SR) output and the customized user text, advancing until a mismatch is identified, after which the electronic device 102b adjusts the output to mirror the user's preference. The process may continue until convergence is achieved, and the relevant rule is logged in the Behavioral Rule Database.

[0092] The behavioral rule generator 214 may be a mechanism to account for rule reversion. If the user alters a previously customized phrase back to its original form, the behavioral rule generator 214a may be configured to detect this "reverted rule" by checking the presence of the modification in the user-confirmed text. If the rule is no longer applicable, it is temporarily removed from the database and assigned a support value of zero, ensuring that only active and relevant customizations influence future ASR (or SR) output.

[0093] In an embodiment, the rule support framework 214b may be configured to filter the one or more rules based on a frequency of the one or more rules. The frequency of the one or more rules may indicate a number of times each rule is applied for the intended recipient. The rule support framework 214b may be designed to filter and validate behavioral rules identified by the behavioral rule generator 214a, ensuring only rules with consistent user support are applied. When a rule is created based on user feedback, it is initially stored with a score representing the number of times the rule has been invoked. If a user repeatedly makes the same modification or correction, this score may increase.

[0094] When a rule reaches a certain threshold, demonstrating sufficient support, it is elevated to a "supported behavioral rule." This elevation indicates that the rule is reliable enough to be consistently applied to the ASR (or SR) output, reflecting a true user preference rather than a one-time correction.

[0095] The behavioral customization application 214c may be configured to use the filtered supported behavioral rules to adjust the ASR (or SR) transcripts, applying each rule as a mapping from specific phrases or patterns to their customized counterparts. The application may perform substring matching based on the key-value pairs stored in the rule map for replacement rules. When a match for the key substring is found in the ASR (or SR) transcript, it is replaced with the corresponding value, transforming the text to reflect user preferences.

[0096] The addition rules may require a slightly more complex process. Here, the application may identify both a left and a right context, as stored in the rule, and checks if these contexts align with the ASR (or SR) output. If they match, the specified addition (such as an emoji or additional text) is inserted between the left and right context, resulting in a customized transcript that incorporates the user's nuanced preferences.

[0097] Referring to FIG. 2a, the recipient-specific customized output module 216 may serve as the final stage where the user may review and make any necessary adjustments to the ASR (or SR)-generated text before it is confirmed and sent to the intended recipient or application.

[0098] Upon receiving the customized text after processing by the behavioral customization engine 214, the user may be provided with the option to modify or correct the generated ASR (or SR) transcript to ensure accuracy and personal preference. Once the user confirms the output, it is redirected to the behavioral rule generator 214a for further analysis. For this analysis, the original customized text is included in the feedback, enabling the electronic device 102b to compare the user-confirmed version with the initially generated text. This feedback loop helps refine by identifying user-specific preferences and common corrections, which can inform future behavioral rule updates.

[0099] FIG. 3 illustrates a block diagram of an electronic device for generating a customized automatic speech recognition transcript, according to an embodiment of the present disclosure.

[0100] The electronic device 102b may be configured to extract a one or more audio features associated with a received audio input from a user. The electronic device 102b may be configured to retrieve a recipient context from one or more messages of a recipient based on conversation history (or conversation history data) with the recipient. The electronic device 102b may be configured to generate a first customized SR transcript based on at least one of the one or more audio features, the recipient context, and a formality vector, wherein the first customized SR transcript may capture the generic formality of the message. The electronic device 102b may be configured to generate a second customized SR transcript based on at least one of the first customized SR transcript and user behavior. The second customized SR transcript may capture idiosyncrasies of an individual user with the recipient of the message, thus allowing the electronic device 102b to generate the customized automatic speech recognition transcript.

[0101] The electronic device 102b may include, but is not limited to, one or more processors 302, a memory 304, one or more modules 306, and data 308. The one or more modules 306 and the memory 304 may be coupled to the one or more processor 302.

[0102] The one or more processor 302 can be a single processing unit or several units, all of which could include multiple computing units. The one or more processor 302 may be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, state machines, logic circuitries, and / or any devices that manipulate signals based on operational instructions. Among other capabilities, the one or more processors 302 are adapted to fetch and execute computer-readable instructions and data stored in the memory 304. The one or more processor 302 may be electrically connected to the memory 304, for example, through a bus. the one or more processor 302 may be referred to as at least one processor. The one or more processor 302 according to an embodiment of the disclosure may include at least one circuitry, such as a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), a many integrated core (MIC), a digital signal processor (DSP), or a neural processing unit (NPU). The one or more processor 302 may be implemented in the form of an integrated system-on-chip (SoC) including one or more electronic components. When the one or more processor 302 is configured as a plurality of processors, the plurality of processors may be realized as separate pieces of hardware (H / W). The one or more processor 302 may be referred to as a microprocessor controller (MICOM), a micro-processor unit (MPU), or a micro-controller unit (MCU).

[0103] The memory 304 may include any non-transitory computer-readable medium known in the art including, for example, volatile memory, such as static random access memory (SRAM) and dynamic random access memory (DRAM), and / or non-volatile memory, such as read-only memory (ROM), erasable programmable ROM, flash memories, hard disks, optical disks, and magnetic tapes.

[0104] The one or more modules 306, amongst other things, include routines, programs, objects, components, data structures, etc., which perform particular tasks or implement data types. The one or more modules 306 may be implemented as, signal processor(s), state machine(s), logic circuitries, and / or any other device or component that manipulates signals based on operational instructions.

[0105] The one or more modules 306 may be implemented in hardware, instructions executed by a processing unit, or by a combination thereof. The processing unit may comprise a computer, a processor, such as the one or more processor 302, a state machine, a logic array, or any other suitable devices capable of processing instructions. The processing unit may be a general-purpose processor which executes instructions to cause the general-purpose processor to perform the required tasks, or the processing unit may be dedicated to performing the required functions. In an embodiment of the present disclosure, the one or more modules 306 may be machine-readable instructions (software) which, when executed by a processor / processing unit, perform any of the described functionalities.

[0106] In an embodiment, the one or more modules 306 may include the input module 202, the audio feature extraction module 204, the recipient-specific context retrieval module 206, the formality cognizant engine 208, the formality vector extraction module 210, the formality cognizant ASR module 212, the behavioral customization engine 214, and the recipient specific customized output module 216. The data 308 serves, amongst other things, as a repository for storing data processed, received, and generated by one or more of the one or more modules 306.

[0107] In an embodiment, the input module 202 may be configured to receive the audio input 102 from the user 102a. The audio input 102 may include the sound of the user's voice (the user's voice command), which is recorded as a waveform and subsequently processed for analysis. The input module 202 may be configured to extract a one or more audio features associated with a received audio input 102 from the user 102a.

[0108] In an embodiment, upon receiving the audio input 102 from the user 102a through the input module 202, the audio feature extraction module 204 may process the audio input to output windowed and framed mel-spectrogram features, supporting a set number of frequency bins. To achieve this, the audio feature extraction module 204 may be configured to utilize a conventional audio feature extraction pipeline based primarily on Mel-frequency Cepstral Coefficients (MFCCs). The audio feature extraction module 204 may be configured to retrieve the recipient context from one or more messages of the recipient based on conversation history with the recipient.

[0109] The formality cognizant engine 208 may be configured to generate a first customized ASR transcript based on at least one of the one or more audio features, the recipient context, and a formality vector. The first customized ASR transcript captures the generic formality of the message. The formality cognizant engine 208 may be configured to generate the first customized ASR transcript based on at least one of the one or more audio features, the recipient context, and the formality vector. The first customized ASR transcript captures the generic formality of the message.

[0110] In an embodiment, the formality cognizant engine 208 may utilize the formality vector extraction module 210 to extract formality vector. The formality vector extraction module 210 may be configured to determine a semantic formality score based on a pre-trained formality prediction artificial intelligence (AI) model based on the recipient context. The formality vector extraction module 210 may be configured to extract one or more lexical features based on analyzing the recipient context. The one or more lexical features indicate characteristics of at least one of words, phrases, and sentences. The formality vector extraction module may be configured to determine the formality vector based on at least one of, the one or more lexical features and the semantic formality score. In an embodiment, the formality vector may include one or more of a formality polarity, an emoji constant, an OOV constant, and a lingo constant.

[0111] Upon determining the formality vector using the formality vector extraction module 210, the formality cognizant ASR module 212 may be configured to extract at least one of the one or more audio features of the received audio input and the formality vector for a prior conversation. The formality vector for the prior conversation is concatenated to a prediction window of one or more prediction windows. The formality cognizant ASR module 212 may be configured to generate one or more tokens using a formality primed ASR AI model based on the one or more audio features and the formality vector for each of the prediction windows. The one or more tokens indicate an identifiable unit of information. The formality cognizant ASR module 212 may be configured to map the one or more tokens to vocabulary elements such that the first customized ASR transcript is generated. The vocabulary elements are stored in the database.

[0112] In an embodiment, upon generating the first customized ASR transcript by the formality cognizant engine 208, the behavioral customization engine 214 may be configured to generate a second customized ASR transcript based on at least one of the first customized ASR transcript and user behavior. The second customized ASR transcript captures idiosyncrasies of an individual user with the recipient of the message.

[0113] The formality cognizant engine 208 may be configured to receive the user feedback on presenting the generated first customized ASR transcript to the user. The formality cognizant engine 208 may be configured to identify one or more rules based on analyzing the feedback. The one or more rules indicate an intent to replace or add tokens to the output of the first customized ASR transcript. The formality cognizant engine 208 may be configured to filter the one or more rules based on a frequency of the one or more rules. The frequency of the one or more rules indicates a number of times each rule is applied to the intended recipient. The formality cognizant engine 208 may be further configured to generate the second customized ASR transcript based on at least one of the filtered one or more rules and the first customized ASR transcript.

[0114] The one or more rules include at least one of an addition rule and a replacement rule. The addition rule indicates the insertion of a user-specified substring at a specific position in the first customized ASR transcript. The replacement rule indicates substituting a substring with a user-specified customization.

[0115] The recipient-specific customized output module 216 may be configured to generate a second customized ASR transcript based on at least one of the first customized ASR transcript and user behavior. The second customized ASR transcript captures idiosyncrasies of an individual user with the recipient of the message.

[0116] FIG. 4 illustrates a flowchart depicting a method for generating the customized automatic speech recognition transcript, according to an embodiment of the present disclosure. The flowchart outlines a sequence of steps involved in the implementation of the electronic device 102b by illustrating how the customized automatic speech recognition transcript is generated.

[0117] At step 402, the method 400 may include extracting a one or more audio features associated with a received audio input from a user.

[0118] At step 404, the method 400 may include retrieving a recipient context from one or more messages of a recipient based on conversation history with the recipient.

[0119] At step 406, the method 400 may include generating a first customized SR (or ASR) transcript based on at least one of the one or more audio features, the recipient context, and a formality vector. The first customized SR transcript captures the generic formality of the message.

[0120] At step 408, the method 400 may include generating a second customized SR transcript based on at least one of the first customized SR transcript and user behavior. The second customized SR transcript captures idiosyncrasies of the user with the recipient of the message.

[0121] FIG.5 illustrates a scenario for generating a customized automatic speech recognition transcript, according to an embodiment of the present disclosure.

[0122] The FIG.5 illustrates how the electronic device 102b, with the behavioral customization engine 214, enhances user experience by adapting voice-to-text output to meet specific formality and personalization preferences. The figure consists of three panels showing a chat interface in different stages of customization.

[0123] In the first panel 510 of the scenario 500, labeled "Original Way of Writing," the user is shown typing messages in his informal, personalized writing style, using shorthand and emojis, which represents his natural way of communicating with friends or close contacts. Here, expressions like "hi..m f9 wht abt u?" and "wil cal u in eveng" reflect the user's personal linguistic preferences, blending informality with abbreviation and casual language.

[0124] In the middle panel 520 of the scenario 500, labeled "Without personalization" the ASR (or SR) output is shown without any customization or behavioral adaptation. Here, the voice assistant converts the user's speech into a more standardized, formal text, changing informal expressions to complete, conventional sentences. For example, the informal greeting is transformed to "Hi, I am fine what about you?" and the phrase "Sure, will cal u in eveng" becomes "Sure, I will call you in the evening." Although the response is grammatically correct, it loses the user's unique conversational style and preferences.

[0125] The third panel 530 of the scenario 500, labeled "With personalization" demonstrates the effect of the proposed invention. Here, the ASR (or SR) output preserves the user's personalized style, including informal spelling, abbreviations, and the use of emojis, thus reflecting the user's original voice input more authentically. The behavioral customization engine has successfully adapted the ASR transcription to align with the user's customary writing style. This adaptation is achieved by applying previously learned user-specific behavioral rules, making the ASR output more natural and familiar to the user and their conversation partners.

[0126] FIG. 6 illustrates a scenario for generating a customized automatic speech recognition transcript, according to an embodiment of the present disclosure.

[0127] The FIG. 6 demonstrates how the electronic device 102b provides customized text for different users based on detected user preferences. In an embodiment, the electronic device 102b customizes subtitles for a video clip based on the detected user, adapting the text format and language style to align with each user's personal preferences.

[0128] In the FIG. 6, a central "Personalized Voice" module 610 is shown, representing capability of the electronic device 102b to modify text output based on user-specific rules. On the left 620, the original video clip displays a subtitle in a more formal style: "Oho!! Someone helps me I'm so frightened." This is the default subtitle text before any user-specific customization is applied.

[0129] The middle and right images 630 illustrate the personalized subtitle adjustments for two different users, "User X" and "User Y." When "User X" is detected, the electronic device 102b modifies the subtitle to match this user's informal, abbreviated style, changing the text to "Oho!! Some1 helps me.. m so wow (emoji)." This modified text uses abbreviations ("Some1" for "Someone" and "m" for "I'm") and includes an emoji, aligning with User X's casual language preferences.

[0130] For "User Y," detected by the system in the right image, the text is adjusted to match a slightly more formal style, but still with a level of personalization that reflects User Y's unique language choices. The subtitle reads, "Oho !! Someone helps me I am so afraid." Here, the language is less formal than the original but does not include abbreviations or emojis, indicating a preference for simpler, direct text without excessive informality.

[0131] The present disclosure advantageously overcomes one or more technical problems associated with the existing systems, such as:

[0132] Firstly, the present disclosure personalizes ASR (or SR) output by generating formality vectors specific to each user, ensuring that responses match their preferred communication style. This customization leads to more natural and personalized interactions, enhancing user satisfaction.

[0133] The present disclosure automates the application of learned behavioral rules, minimizing the need for manual edits. This results in a more streamlined experience, allowing users to focus on effective communication without spending time adjusting the output.

[0134] The present disclosure translates formal or complex language into a simpler, user-specified formats, content is made more accessible to a wider range of users. Individuals with limited language proficiency, literacy challenges, or specific communication preferences are especially benefited.

[0135] The present disclosure applies behavioral rules over time, reducing the need for users to repeatedly edit or adjust their messages. Once a user's preference is recognized and established as a supported rule, it can be consistently applied, saving users time and effort in reformatting text.

[0136] Further numerous advantages of the present disclosure include a user-centric approach, efficiency enhancement, communication optimization, adaptability to user behavior, competitive advantage, alignment with industry trends, market differentiation, and future-proofing capabilities.

[0137] In an embodiment, to determine the formality vector, the method 400 comprises determining a semantic formality score using a pre-trained formality prediction artificial intelligence (AI) model (210a) based on the recipient context, extracting one or more lexical features through analyzing the recipient context, wherein the one or more lexical features indicate characteristics of at least one of words, phrases, and sentences, and determining the formality vector based on at least one of, the one or more lexical features and the semantic formality score.

[0138] In an embodiment, the formality vector includes one or more of a formality polarity, an emoji constant, an out-of-vocabulary (OOV) constant, and a lingo constant.

[0139] In an embodiment, the generating 406 the first customized ASR transcript, extracting at least one of the one or more audio features of the received audio input and the formality vector for prior conversation, wherein the formality vector for the prior conversation is concatenated with a prediction window of one or more prediction windows, generating one or more tokens using a pre-trained formality prediction AI model 210a based on the one or more audio features and the formality vector for each of the one or more prediction windows, wherein the one or more tokens indicate an identifiable unit of information, and mapping the one or more tokens to vocabulary elements to generate the first customized ASR transcript, wherein the vocabulary elements are stored in a database.

[0140] In an embodiment, the generating 408 the second customized ASR transcript comprises receiving user feedback on presenting of the generated first customized ASR transcript to the user, identifying one or more rules through analyzing the feedback, wherein the one or more rules indicate intent to replace tokens in the output or add tokens to the output of the first customized ASR transcript, filtering the one or more rules based on a frequency of the one or more rules, wherein the frequency of the one or more rules indicates a number of times each rule is applied to the recipient, and generating the second customized ASR transcript based on at least one of the filtered one or more rules and the first customized ASR transcript.

[0141] In an embodiment, the one or more rules include at least one of an addition rule or a replacement rule, and wherein the addition rule indicates insertion of a user-specified substring at a specific position in the first customized ASR transcript, and wherein the replacement rule indicates substituting of a substring by a user-specified customization.

[0142] In an embodiment, the extracting (402) of the one or more audio features associated with the received audio input comprises framing the received audio input within a specified duration with a specified percentage of overlap, wherein the framing indicates segmentation of the received audio input, applying a discrete Short-Time Fourier transform (STFT) to each frame, and extracting the one or more audio features based on frequency-domain representation of each frame, wherein the extracted one or more audio features include Mel-frequency Cepstral Coefficients (MFCCs).

[0143] In an embodiment, the method 400 comprises updating the conversation history between the user and the recipient periodically.

[0144] In an embodiment, an electronic device 102b, to determine the formality vector, the at least one processor (302) is configured to execute the at least one instruction to perform determining a semantic formality score using a pre-trained formality prediction artificial intelligence (AI) model (210a) based on the recipient context, extracting one or more lexical features through analyzing the recipient context, wherein the one or more lexical features indicate characteristics of at least one of words, phrases, and sentences, and determining the formality vector based on at least one of, the one or more lexical features and the semantic formality score.

[0145] In an embodiment, to generate the first customized ASR transcript, the at least one processor 302 is configured to execute the at least one instruction to perform extracting at least one of the one or more audio features of the received audio input, and the formality vector for prior conversation, wherein the formality vector for the prior conversation is concatenated with a prediction window of one or more prediction windows, generating one or more tokens using a pre-trained formality prediction AI model 210a based on the one or more audio features and the formality vector for each of the one or more prediction windows, wherein the one or more tokens indicate an identifiable unit of information, and mapping the one or more tokens to vocabulary elements to generate the first customized ASR transcript, wherein the vocabulary elements are stored in a database.

[0146] In an embodiment, to generate the second customized ASR transcript, the at least one processor (302) is configured to execute the at least one instruction to perform receiving user feedback on presenting of the generated first customized ASR transcript to the user, identifying one or more rules through analyzing the feedback, wherein the one or more rules indicate intent to replace tokens in the output or add tokens to the output of the first customized ASR transcript, filtering the one or more rules based on a frequency of the one or more rules, wherein the frequency of the one or more rules indicates a number of times each rule is applied to the recipient, and generating the second customized ASR transcript based on at least one of the filtered one or more rules and the first customized ASR transcript.

[0147] In an embodiment, the one or more rules include at least one of an addition rule or a replacement rule, and wherein the addition rule indicates insertion of a user-specified substring at a specific position in the first customized ASR transcript, and wherein the replacement rule indicates substituting of a substring by a user-specified customization.

[0148] In an embodiment, to extract the one or more audio features o associated with the received the audio input, the at least one processor 302 is configured to execute the at least one instruction to perform framing the received audio input with a specified duration with a specified percentage of overlapping, wherein the framing indicates segmentation of the received audio input, applying a discrete short time Fourier transform (STFT) to each frame, and extracting the one or more audio features based on frequency domain representation of each frame, wherein the extracted one or more audio features include Mel-frequency Cepstral Coefficients (MFCCs).

[0149] In an embodiment, the at least one processor is configured to update the historical conversation between the user and the recipient periodically.

[0150] While specific language has been used to describe the present disclosure, any limitations arising on account thereto, are not intended. As would be apparent to a person in the art, various working modifications may be made to the method to implement the inventive concept as taught herein. The drawings and the foregoing description give examples of embodiments. Those skilled in the art will appreciate that one or more of the described elements may well be combined into a single functional element. Alternatively, certain elements may be split into multiple functional elements. Elements from one embodiment may be added to another embodiment.

Claims

1.A method 400 performed by an electronic device (102b) for generating a customized automatic speech recognition (ASR) transcript, the method 400 comprising:extracting 402one or more audio features associated with a received audio input from a user;retrieving 404 a recipient context from one or more messages of a recipient based on conversation history with the recipient;generating 406 a first customized ASR transcript based on at least one of the one or more audio features, the recipient context, and a formality vector, wherein the first customized ASR transcript captures the generic formality of the message; andgenerating 408 a second customized ASR transcript based on at least one of the first customized ASR transcript and user behavior, wherein the second customized ASR transcript captures idiosyncrasies of a user with the recipient of the message.2.The method 400 as claimed in claim 1, wherein to determine the formality vector, the method comprises:determining a semantic formality score using a pre-trained formality prediction artificial intelligence (AI) model 210a based on the recipient context;extracting one or more lexical features through analyzing the recipient context, wherein the one or more lexical features indicate characteristics of at least one of words, phrases, and sentences; anddetermining the formality vector based on at least one of, the one or more lexical features and the semantic formality score.3.The method 400 as claimed in claim 1 or 2, wherein the formality vector includes one or more of a formality polarity, an emoji constant, an out-of-vocabulary (OOV) constant, and a lingo constant.4.The method 400 as claimed in claim 1, wherein the generating 406 the first customized ASR transcript comprises:extracting at least one of the one or more audio features of the received audio input and the formality vector for prior conversation, wherein the formality vector for the prior conversation is concatenated with a prediction window of one or more prediction windows;generating one or more tokens using a pre-trained formality prediction AI model 210a based on the one or more audio features and the formality vector for each of the one or more prediction windows, wherein the one or more tokens indicate an identifiable unit of information; andmapping the one or more tokens to vocabulary elements to generate the first customized ASR transcript, wherein the vocabulary elements are stored in a database.5.The method 400 as claimed in any one of claims 1 to 4, wherein the generating 408 the second customized ASR transcript comprises:receiving user feedback on presenting of the generated first customized ASR transcript to the user;identifying one or more rules through analyzing the feedback, wherein the one or more rules indicate intent to replace tokens in the output or add tokens to the output of the first customized ASR transcript;filtering the one or more rules based on a frequency of the one or more rules, wherein the frequency of the one or more rules indicates a number of times each rule is applied to the recipient; andgenerating the second customized ASR transcript based on at least one of the filtered one or more rules and the first customized ASR transcript.6.The method 400 as claimed in claim 5, wherein the one or more rules include at least one of an addition rule or a replacement rule, wherein the addition rule indicates insertion of a user-specified substring at a specific position in the first customized ASR transcript, and wherein the replacement rule indicates substituting of a substring by a user-specified customization.7.The method 400 as claimed in any one of claims 1 to 6, wherein the extracting 402 of the one or more audio features associated with the received audio input comprises:framing the received audio input within a specified duration with a specified percentage of overlap, wherein the framing indicates segmentation of the received audio input;applying a discrete Short-Time Fourier transform (STFT) to each frame; andextracting the one or more audio features based on frequency-domain representation of each frame, wherein the extracted one or more audio features include Mel-frequency Cepstral Coefficients (MFCCs).8.The method 400 as claimed in any one of claims 1 to 7, the method further comprising:updating the conversation history between the user and the recipient periodically.9.An electronic device 102b for generating a customized automatic speech recognition (ASR) transcript, the electronic device 102b comprising:a memory 304 configured to store at least one instruction a computer program;at least one processor 302 configured to execute the at least one instruction to perform:extracting one or more audio features associated with a received audio input from a user;retrieving a recipient context from one or more messages of a recipient based on conversation history with the recipient;generating a first customized ASR transcript based on at least one of the one or more audio features, the recipient context, and a formality vector, wherein the first customized ASR transcript captures the generic formality of the message; andgenerating a second customized ASR transcript based on at least one of the first customized ASR transcript, a user behavior, wherein the second customized ASR transcript captures anidiosyncrasies of the user with the recipient of the message.10.The electronic device 102b as claimed in claim 9, wherein to determine the formality vector, the at least one processor 302 is configured to execute the at least one instruction to perform:determining a semantic formality score using a pre-trained formality prediction artificial intelligence (AI) model 210a based on the recipient context;extracting one or more lexical features through analyzing the recipient context, wherein the one or more lexical features indicate characteristics of at least one of words, phrases, and sentences; anddetermining the formality vector based on at least one of, the one or more lexical features and the semantic formality score.11.The electronic device 102b as claimed in claim 9 or 10, wherein the formality vector includes one or more of a formality polarity, an emoji constant, an out-of-vocabulary (OOV) constant, and a lingo constant.12.The electronic device 102b as claimed in claim 9, wherein to generate the first customized ASR transcript, the at least one processor 302 is configured to execute the at least one instruction to perform:extracting at least one of the one or more audio features of the received audio input, and the formality vector for prior conversation, wherein the formality vector for the prior conversation is concatenated with a prediction window of one or more prediction windows;generating one or more tokens using a pre-trained formality prediction AI model 210a based on the one or more audio features and the formality vector for each of the one or more prediction windows, wherein the one or more tokens indicate an identifiable unit of information; andmapping the one or more tokens to vocabulary elements to generate the first customized ASR transcript, wherein the vocabulary elements are stored in a database.13.The electronic device 102b as claimed in any one of claims 9 to 12, wherein to generate the second customized ASR transcript, the at least one processor 302 is configured to execute the at least one instruction to perform:receiving user feedback on presenting of the generated first customized ASR transcript to the user;identifying one or more rules through analyzing the feedback, wherein the one or more rules indicate intent to replace tokens in the output or add tokens to the output of the first customized ASR transcript;filtering the one or more rules based on a frequency of the one or more rules, wherein the frequency of the one or more rules indicates a number of times each rule is applied to the recipient; andgenerating the second customized ASR transcript based on at least one of the filtered one or more rules and the first customized ASR transcript.14.The electronic device 102b as claimed in any one of claims 9 to 13, wherein the one or more rules include at least one of an addition rule or a replacement rule, wherein the addition rule indicates insertion of a user-specified substring at a specific position in the first customized ASR transcript, and wherein the replacement rule indicates substituting of a substring by a user-specified customization.15.A computer readable storage medium having stored thereon a computer program that, when executed by at least one processor, performs the method of any one of claims 1-8.