A voice call real-time transcription system and method
The real-time speech transcription system, which utilizes layered compression and multimodal feature extraction, solves the problems of noise filtering difficulties and insufficient personalization in existing technologies, achieving efficient and accurate speech transcription while meeting the requirements for real-time performance and stability.
Patent Information
- Application Number
- CN202511028384.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-07-25
AI Technical Summary
Existing voice call transcription technology struggles to effectively filter out background noise in noisy environments, leading to difficulties in speech recognition, high error rates in transcribed text, a lack of personalized customization features, insufficient real-time performance and stability, and an inability to meet the demands for high-quality transcription.
The real-time speech transcription system employing layered compression and multimodal feature extraction includes a network element module, a speech streaming engine, a speech engine, and an analysis and optimization module. It performs audio compression using a perceptual weighted vector quantization algorithm, performs text transcription by combining multimodal feature data and a preset speech recognition model, and makes personalized adjustments using a preset vocabulary.
It improves the accuracy and real-time performance of speech transcription, meets users' personalized needs, enhances the stability and robustness of the system, reduces computing resource consumption and transmission costs, and achieves efficient transcription in unstable network environments.
Smart Images

Figure CN120526774B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and more specifically, to a real-time voice call transcription system and method. Background Technology
[0002] In today's era of increasingly frequent digital communication, the importance of voice call transcription technology is becoming increasingly prominent, and it is widely used in many fields such as meeting minutes, customer service, and smart office. However, existing technologies have obvious limitations and urgently need innovation.
[0003] Existing systems are weak in audio preprocessing, making it difficult to effectively filter out background noise in noisy environments, leading to difficulties in speech engine recognition and a high error rate in transcribed text. Furthermore, relying solely on single speech features for text writing significantly compromises the semantic integrity and contextual coherence of the transcribed text, failing to meet the demands for high-quality transcription.
[0004] Furthermore, the personalized requirements for transcription results from different user groups and application scenarios are increasing, but existing systems lack flexible customization capabilities, making it difficult to optimize and adjust them according to specific needs and preferences. In practical applications, real-time performance and stability issues are also prominent. Due to system resource limitations and low algorithm efficiency, high transcription latency occurs in the voice data transmission, processing, and transcription process, which greatly affects the user experience. Especially in scenarios with extremely high real-time requirements, such as real-time caption generation and online meeting recording, existing technologies struggle to achieve ideal results. Summary of the Invention
[0005] The problem solved by this invention is one or more of the aforementioned related technical problems.
[0006] To address the aforementioned problems, this invention provides a real-time voice call transcription system and method.
[0007] In a first aspect, the present invention provides a real-time voice call transcription system, including a network element module, a voice streaming engine, a voice engine, and an analysis and optimization module;
[0008] The network element module is used to obtain corresponding audio data when a call request from a user terminal is detected.
[0009] The voice streaming engine is used to perform layered compression of the audio data based on a preset perceptual weighted vector quantization algorithm to obtain compressed audio data, and to perform format conversion processing on the compressed audio data to obtain temporary audio data.
[0010] The speech engine is used to extract features from the temporary audio data to obtain multimodal feature data, and to process the multimodal feature data based on a preset speech recognition model to obtain text information.
[0011] The analysis and optimization module is used to obtain corresponding real-time transcribed text data based on the preset large model, the text information, and the preset vocabulary.
[0012] Optionally, the multimodal feature data includes voiceprint feature data and acoustic parameter data; the step of extracting features from the temporary audio data to obtain multimodal feature data includes:
[0013] Feature extraction is performed on the temporary audio data to obtain the voiceprint feature data;
[0014] The temporary audio data is identified to obtain the acoustic parameter data.
[0015] Optionally, the step of extracting features from the temporary audio data to obtain the voiceprint feature data includes:
[0016] The temporary audio data is denoised to obtain processed audio data, and the processed audio data is then divided into frames to obtain multiple framed audio data.
[0017] A Fourier transform is performed on each frame of audio data to obtain the corresponding signal spectrum data. Then, the MFCC extraction method is used to extract features from each signal spectrum data to obtain the corresponding MFCC feature data. All the MFCC feature data are then concatenated to obtain the voiceprint feature data.
[0018] Optionally, the acoustic parameter data includes short-time energy, zero-crossing rate, and fundamental frequency standard deviation; the step of identifying the temporary audio data to obtain the acoustic parameter data includes:
[0019] Windowing is applied to each frame of audio data to obtain windowed signal data, and the corresponding short-time energy is calculated from the windowed signal data.
[0020] The number of zero crossings is determined for each frame of audio data, and the zero crossing rate is determined based on the number of zero crossings.
[0021] Based on a preset algorithm, the corresponding baseband data is determined according to each frame of audio data, and the baseband standard deviation is determined according to all the baseband data.
[0022] Optionally, the real-time voice call transcription system further includes a mixing scheduling engine, which is specifically used for:
[0023] Obtain current call demand data and engine load status, and determine current load data based on the engine load status;
[0024] The current call demand data is input into a preset load prediction model to obtain load prediction data;
[0025] Based on preset weighted data, target load data is obtained according to the current load data and the predicted load data;
[0026] The target load data is compared with the preset load data to obtain a comparison result, and the real-time voice call transcription system is dynamically allocated resources based on the comparison result; wherein, the load prediction model is constructed based on machine learning algorithms.
[0027] Optionally, the step of dynamically allocating resources to the real-time voice call transcription system based on the comparison result includes:
[0028] When the target load data is less than the preset load data, the real-time voice call transcription system is dynamically allocated resources according to the polling allocation algorithm.
[0029] When the target load data is greater than or equal to the preset load data, the QoS level and corresponding current link latency of each call request are determined; the weight data corresponding to the call request is determined according to each QoS level and the corresponding current link latency; and the voice call real-time transcription system is dynamically allocated resources according to each weight data.
[0030] Optionally, the process of constructing the preset vocabulary database includes:
[0031] Obtain the corresponding historical call text data, and identify the historical call text data based on the BiLSTM-CRF model to obtain keyword information, which includes terminology data and keyword data;
[0032] Based on the keyword information, the corresponding weight information is determined, and the keyword information is sorted and filtered according to the corresponding weight information to obtain high-frequency terms and personalized keywords.
[0033] The corresponding preset vocabulary library is determined based on the high-frequency terms and the personalized keywords.
[0034] Optionally, the real-time voice call transcription system further includes a voice segmentation module:
[0035] The speech segmentation module is used to segment each frame of audio data according to a preset segmentation trigger condition;
[0036] The preset sentence segmentation triggering conditions include silence duration triggering conditions, fundamental frequency standard deviation triggering conditions, and interjection triggering conditions.
[0037] Optionally, the real-time transcribed text data includes highlighted text data with timestamps; the real-time speech transcription system further includes a subtitle streaming module, which is specifically used for:
[0038] The highlighted text data is packaged to obtain independent subtitle frames, and independent subtitle layers are obtained based on the independent subtitle frames and preset personalized configurations;
[0039] The independent subtitle layer is parsed and rendered to obtain subtitle data, and the subtitle data is pushed to the display device.
[0040] Secondly, the present invention provides a real-time voice call transcription method, applied to the aforementioned real-time voice call transcription system, the transcription method comprising:
[0041] When a call request is detected from the user's end, the corresponding audio data is retrieved.
[0042] Based on a preset perceptual weighted vector quantization algorithm, the audio data is compressed in layers to obtain compressed audio data, and the compressed audio data is then converted to obtain temporary audio data.
[0043] Feature extraction is performed on the temporary audio data to obtain multimodal feature data, and the multimodal feature data is processed based on a preset speech recognition model to obtain text information;
[0044] Based on a pre-set large model, corresponding real-time transcribed text data is obtained according to the text information and a pre-set vocabulary.
[0045] The beneficial effects of the real-time voice call transcription system and method of the present invention are:
[0046] When a user initiates a call request, the network element module retrieves the audio data corresponding to the call request from the communication network. The retrieved audio data is then passed to the voice streaming engine.
[0047] The audio data is compressed in layers using a pre-defined Perceptual Weighted Vector Quantization (PWVQ) algorithm by the voice streaming engine. This layered compression typically includes a base layer (preserving the core speech frequency band of 300-3400Hz with a compression ratio of at least 50%), an enhancement layer (supplementing high-frequency details of 3400-8000Hz and dynamically allocating redundant packets to resist packet loss), and metadata layer processing (embedding timestamps and engine identifiers). The compressed audio data undergoes format conversion to obtain temporary audio data. This layered compression approach ensures both the core quality of the speech data and improves data transmission efficiency and resilience against packet loss. Format conversion ensures the audio data can be correctly processed by subsequent modules, while the voice streaming engine effectively removes interference and enhances the clarity of the speech signal.
[0048] The temporary audio data is processed by a speech engine to extract features, including but not limited to acoustic features (such as MFCC), intonation features, and rhythm features. These extracted features are then fused to generate multimodal feature data. This process comprehensively represents speech information by extracting rich features, and the fusion of multimodal feature data effectively utilizes information from multiple modalities. The multimodal feature data is then processed based on a pre-defined speech recognition model, which possesses high-accuracy speech-to-text conversion capabilities.
[0049] Finally, the analysis and optimization module analyzes and optimizes the text information based on a pre-set large model and a pre-set vocabulary to obtain real-time transcribed text data. This process, based on the pre-set large model and pre-set vocabulary, provides rich language knowledge and vocabulary information, improving the accuracy and real-time performance of the transcribed text.
[0050] In summary, this invention utilizes multimodal feature data to comprehensively represent speech information, enabling speech recognition models to more accurately convert speech to text. This not only significantly improves the accuracy of transcribed text data, providing a more realistic reflection of the speech content, but also, through an efficient audio acquisition, compression, preprocessing, feature extraction, and text generation process, allows the system to complete the transcription task in a short time, achieving real-time transcription of speech calls. This feature meets the needs of users in application scenarios with high real-time requirements, such as meeting minutes and real-time captions. Furthermore, the preset vocabulary can be customized according to different user needs, including terminology from specific fields and personalized vocabulary, allowing the transcribed text to better meet users' individual needs and improving the practicality and value of the transcription results.
[0051] Furthermore, this system demonstrates powerful speech understanding capabilities. By utilizing multimodal feature data, it can not only recognize speech content but also understand information such as emotion and intonation, resulting in more complete and in-depth transcribed text data. The real-time speech transcription system consists of network element modules, a speech streaming engine, a speech engine, and an analysis and optimization module. The collaborative work between these modules optimizes resource allocation and utilization, improves the system's efficiency in processing speech transcription tasks, reduces computational resource consumption, and enhances system stability and reliability. Through layered compression and format conversion, the system reduces the transmission bandwidth and storage requirements of audio data, lowering transmission and storage costs and improving overall efficiency while ensuring speech quality. Finally, dynamic allocation of redundant packets in the enhancement layer to resist packet loss improves robustness in unstable network environments, ensuring the integrity of speech data and the coherence of transcribed text. Attached Figure Description
[0052] Figure 1 This is one of the structural schematic diagrams of a real-time voice call transcription system according to an embodiment of the present invention;
[0053] Figure 2 This is a second schematic diagram of the structure of a real-time voice call transcription system according to an embodiment of the present invention;
[0054] Figure 3 This is a flowchart illustrating a real-time voice call transcription method according to an embodiment of the present invention. Detailed Implementation
[0055] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present invention. It should be understood that the accompanying drawings and embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of protection of the present invention.
[0056] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.
[0057] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to"; the term "based on" means "at least partially based on"; the term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments"; and the term "optionally" means "optional embodiments". Definitions of other terms will be given in the following description. It should be noted that the concepts of "first," "second," etc., mentioned in this invention are used only to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.
[0058] It should be noted that the terms "a" and "a plurality of" used in this invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0059] The names of the messages or information exchanged between the multiple devices in the embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0060] To address the problems existing in the aforementioned related technologies, embodiments of the present invention provide a real-time voice call transcription system and method.
[0061] like Figure 1 As shown in the figure, an embodiment of the present invention provides a real-time voice call transcription system, including a network element module, a voice streaming engine, a voice engine, and an analysis and optimization module;
[0062] The network element module is used to obtain the corresponding audio data when a call request from a user terminal is detected.
[0063] Specifically, when the system detects a call request from a user, such as when a mobile phone makes a call and initiates a call request or invitation signal, it connects through the network element module using the QUIC protocol, triggers the network element instantiation process, and then establishes a dedicated voice data channel, so that the connection establishment time does not exceed 20 milliseconds.
[0064] When a user initiates a call request, it is typically done by dialing a phone number or clicking a call button. Upon detecting a call request, a response mechanism is immediately activated. Efficient signaling processing ensures the system can quickly identify and respond to call requests. The network element module interacts with relevant equipment in the communication network (such as base stations and switching centers) to obtain the audio data stream corresponding to the call request. This ensures the integrity and real-time nature of the audio data, providing a reliable data foundation for subsequent processing.
[0065] During the acquisition of audio data, the network element module may employ optimized data transmission protocols, such as the QUIC protocol, to establish a connection, triggering the network element instantiation process and establishing a dedicated voice data channel. This ensures that the connection establishment time does not exceed 20 milliseconds, thereby reducing latency and improving data transmission stability. In other words, through optimized network configuration and resource scheduling, it ensures that audio data can be transmitted quickly and stably to subsequent processing modules.
[0066] The above process can quickly detect and respond to call requests from the user, effectively reducing user waiting time and thus improving user experience. Simultaneously, through the efficient operation of the network element module, the system can acquire high-quality audio data in a timely manner, providing a solid foundation for subsequent voice processing. Furthermore, the adoption of advanced data transmission protocols and network optimization technologies ensures fast and stable transmission of audio data, significantly reducing data loss and transmission latency. In summary, through its efficient call request detection and audio data acquisition mechanism, the system lays a solid foundation for the entire voice processing flow, improving system response speed and data processing quality, and providing users with a smoother and more reliable voice call transcription service.
[0067] The voice streaming engine is used to perform layered compression of the audio data based on a preset perceptual weighted vector quantization algorithm to obtain compressed audio data, and to perform format conversion processing on the compressed audio data to obtain temporary audio data.
[0068] Specifically, based on the preset perceptual weighted vector quantization (PWVQ) algorithm: PWVQ is an efficient audio compression technique that can significantly reduce the size of audio data while ensuring speech quality.
[0069] Layered compression process: Base layer: Retains the core frequency band of speech content (300-3400Hz), with a compression rate of no less than 50%. This layer ensures that the basic content and intelligibility of the speech are preserved.
[0070] Enhancement layer: This layer supplements high-frequency details (3400-8000Hz) and dynamically allocates redundant packets to combat packet loss. This layer improves speech quality and naturalness.
[0071] Metadata layer: Embedded timestamps and engine identifiers (Engine-ID). Timestamps are used to synchronize and identify the timing information of voice data, and engine identifiers are used to identify the specific engine instance that processes the data.
[0072] The compressed data stream is processed through a jitter buffer to eliminate network jitter and then converted into a temporary audio data format. This format is typically an intermediate format used internally by the system to improve the efficiency of subsequent processing. This process supports multiple encoding formats, ensuring that the converted audio data is compatible with and can be processed by subsequent speech engines and other modules.
[0073] The PWVQ algorithm significantly reduces audio data size while maintaining speech quality. The base layer compression rate is no less than 50%, preserving the core content of the speech and reducing storage and transmission requirements. The enhancement layer supplements high-frequency details and dynamically allocates redundant packets to combat packet loss, improving the naturalness and integrity of the speech. The timestamps and engine identifiers embedded in the metadata layer ensure the synchronization and traceability of the speech data, facilitating subsequent processing and analysis. Format conversion processing ensures that audio data can be efficiently processed by subsequent modules, reducing the time and resource consumption of data conversion and adaptation.
[0074] The combination of layered compression and format conversion reduces the bandwidth requirements for audio data transmission and improves the overall system efficiency, especially under network conditions with limited bandwidth. Dynamic allocation of redundant packets enhances the system's robustness in unstable network environments, ensuring the integrity of voice data and the coherence of transcribed text.
[0075] In summary, the voice streaming engine, through efficient layered compression and format conversion, lays a solid foundation for subsequent voice processing, improves the overall system performance and reliability, and brings numerous benefits to users.
[0076] The speech engine is used to extract features from the temporary audio data to obtain multimodal feature data, and to process the multimodal feature data based on a preset speech recognition model to obtain text information.
[0077] Specifically, the speech engine extracts features from temporary audio data. These features include, but are not limited to, acoustic features (such as Mel-frequency cepstral coefficients (MFCCs)), intonation features, and rhythm features. The extracted features are then fused to generate multimodal feature data. This multimodal feature data comprehensively represents speech information, including not only the acoustic properties of speech but also information such as intonation and emotion.
[0078] The pre-defined speech recognition model is a deep learning-based model (such as recurrent neural networks, recurrent neural networks (RNNs), convolutional neural networks (CNNs), etc.) that processes multimodal feature data. Based on the knowledge gained during training, the speech recognition model maps the multimodal feature data into corresponding text information. By analyzing patterns and rules in the feature data, the model identifies words and phrases in the speech, ultimately generating text information.
[0079] By extracting multimodal feature data to comprehensively represent speech information, speech recognition models can more accurately convert speech to text. The accuracy of transcribed text data is significantly improved, and the content of spoken conversations can be more realistically reproduced. Multimodal feature data covers the acoustic characteristics of speech as well as information such as intonation and emotion, so that transcribed text data not only contains the literal meaning of the speech, but also reflects the speaker's emotions and tone, enhancing the completeness and depth of transcription.
[0080] The speech recognition model can capture complex patterns in speech signals, such as intonation variations and emotional expressions, thereby enhancing the system's ability to understand speech, better handle various speech scenarios, and further improve transcription quality. Furthermore, the model has been trained on a large amount of data and can adapt to different speech characteristics and environmental conditions. The feature extraction and speech recognition model processing flow is efficient and fast, quickly converting audio data into text information, enabling transcription tasks to be completed in a short time, meeting the needs of applications with high real-time requirements.
[0081] In summary, by leveraging the feature extraction of the speech engine and the processing of the speech recognition model, the system generates text data rich in semantic information, significantly improving the accuracy and completeness of transcribed text, enhancing the system's ability to understand speech, optimizing processing efficiency, and enabling it to adapt to various speech scenarios and meet real-time transcription requirements.
[0082] The analysis and optimization module is used to obtain corresponding real-time transcribed text data based on the preset large model, the text information, and the preset vocabulary.
[0083] Specifically, the analysis and optimization module receives text information from the speech engine and performs preliminary integration to prepare for subsequent processing. Using a pre-defined large model (such as a Transformer-based deep learning language model), deep semantic understanding and contextual analysis are performed on the text information. This large model can capture complex semantic relationships, grammatical structures, and potential contextual information in the text, correcting potential grammatical errors, supplementing missing context, and making the text more fluent, natural, and linguistically compliant.
[0084] The text, optimized by a large model, is matched and compared with a preset vocabulary database. This database includes user-specific high-frequency terms, personalized keywords, and domain-specific professional terms. Based on the vocabulary and phrases in the database, the text undergoes further personalized adjustments and optimizations to ensure that the transcription accurately reflects the user's language habits, domain-specific characteristics, and specific expressions.
[0085] The system integrates the semantic analysis results of the large-scale model with personalized word matching from the vocabulary database to generate the final real-time transcribed text data. This transcribed text data is then converted into a user-friendly format (such as JSON, XML, or plain text) for easy viewing, editing, or integration into other applications. The large-scale model corrects grammatical errors and supplements contextual information through deep semantic analysis, combined with personalized word matching from the pre-set vocabulary database, ensuring the transcribed text accurately reflects the user's intended meaning, reducing ambiguity and errors, and improving transcription quality. It also understands the contextual relationships and semantic structure of the text, generating more coherent and natural transcribed text, enhancing the user's reading experience. The pre-set vocabulary database is customized according to specific user needs and professional fields, containing commonly used terms and keywords, making the transcription results more aligned with the user's language habits and professional background, enhancing the text's practicality and relevance. The combination of the pre-set large-scale model and the vocabulary database allows the system to adapt to the speech transcription needs of different users and fields, compatible with various language styles and professional domains, enhancing the system's versatility and adaptability. Furthermore, the efficient processing flow of the analysis and optimization module can complete text optimization and transcription tasks in a short time, ensuring the system outputs transcription results in real time, meeting users' real-time requirements.
[0086] In summary, by analyzing and optimizing the module's deep semantic analysis and personalized word matching, high-precision and personalized real-time transcribed text data is generated, effectively improving the system's transcription quality and user experience.
[0087] In this embodiment, when a user initiates a call request, the real-time voice call transcription system obtains the audio data corresponding to the call request from the communication network through a network element module. The obtained audio data is then transmitted to the voice streaming engine.
[0088] The audio data is compressed in layers using a pre-defined Perceptual Weighted Vector Quantization (PWVQ) algorithm by the voice streaming engine. This layered compression typically includes a base layer (preserving the core speech frequency band of 300-3400Hz with a compression ratio of at least 50%), an enhancement layer (supplementing high-frequency details of 3400-8000Hz and dynamically allocating redundant packets to resist packet loss), and metadata layer processing (embedding timestamps and engine identifiers). The compressed audio data undergoes format conversion to obtain temporary audio data. This layered compression approach ensures both the core quality of the speech data and improves data transmission efficiency and resilience against packet loss. Format conversion ensures the audio data can be correctly processed by subsequent modules, while the voice streaming engine effectively removes interference and enhances the clarity of the speech signal.
[0089] The temporary audio data is processed by a speech engine to extract features, including but not limited to acoustic features (such as MFCC), intonation features, and rhythm features. These extracted features are then fused to generate multimodal feature data. This process comprehensively represents speech information by extracting rich features, and the fusion of multimodal feature data effectively utilizes information from multiple modalities. The multimodal feature data is then processed based on a pre-defined speech recognition model, which possesses high-accuracy speech-to-text conversion capabilities.
[0090] Finally, the analysis and optimization module analyzes and optimizes the text information based on a pre-set large model and a pre-set vocabulary to obtain real-time transcribed text data. This process, based on the pre-set large model and pre-set vocabulary, provides rich language knowledge and vocabulary information, improving the accuracy and real-time performance of the transcribed text.
[0091] In summary, this invention utilizes multimodal feature data to comprehensively represent speech information, enabling speech recognition models to more accurately convert speech to text. This not only significantly improves the accuracy of transcribed text data, providing a more realistic reflection of the speech content, but also, through an efficient audio acquisition, compression, preprocessing, feature extraction, and text generation process, allows the system to complete the transcription task in a short time, achieving real-time transcription of speech calls. This feature meets the needs of users in application scenarios with high real-time requirements, such as meeting minutes and real-time captions. Furthermore, the preset vocabulary can be customized according to different user needs, including terminology from specific fields and personalized vocabulary, allowing the transcribed text to better meet users' individual needs and improving the practicality and value of the transcription results.
[0092] Furthermore, this system demonstrates powerful speech understanding capabilities. By utilizing multimodal feature data, it can not only recognize speech content but also understand information such as emotion and intonation, resulting in more complete and in-depth transcribed text data. The real-time speech transcription system consists of network element modules, a speech streaming engine, a speech engine, and an analysis and optimization module. The collaborative work between these modules optimizes resource allocation and utilization, improves the system's efficiency in processing speech transcription tasks, reduces computational resource consumption, and enhances system stability and reliability. Through layered compression and format conversion, the system reduces the transmission bandwidth and storage requirements of audio data, lowering transmission and storage costs and improving overall efficiency while ensuring speech quality. Finally, dynamic allocation of redundant packets in the enhancement layer to resist packet loss improves robustness in unstable network environments, ensuring the integrity of speech data and the coherence of transcribed text.
[0093] Optionally, the multimodal feature data includes voiceprint feature data and acoustic parameter data; the step of extracting features from the temporary audio data to obtain multimodal feature data includes:
[0094] Feature extraction is performed on the temporary audio data to obtain the voiceprint feature data;
[0095] The temporary audio data is identified to obtain the acoustic parameter data.
[0096] Specifically, the temporary audio data is first denoised to eliminate background noise and interference, thereby obtaining a clearer speech signal. Then, the temporary audio data is subjected to feature extraction and recognition processing to obtain voiceprint feature data and acoustic parameter data (such as fundamental frequency (F0), zero crossing rate, other acoustic parameters, etc.).
[0097] By extracting voiceprint feature data and acoustic parameter data, the system can more comprehensively capture the characteristics of speech signals, thereby improving the accuracy of the speech recognition system. Voiceprint feature data provides unique biometric information, helping to distinguish different speakers and enhancing personalized recognition and verification capabilities. Analyzing acoustic parameter data allows for better simulation of the characteristics of natural speech, resulting in more natural and realistic speech output in speech synthesis. Accurate voiceprint recognition and acoustic parameter analysis can improve the system's response speed and accuracy, providing users with a smoother and more satisfying interactive experience.
[0098] Optionally, the step of extracting features from the temporary audio data to obtain the voiceprint feature data includes:
[0099] The temporary audio data is denoised to obtain processed audio data, and the processed audio data is then divided into frames to obtain multiple framed audio data.
[0100] A Fourier transform is performed on each frame of audio data to obtain the corresponding signal spectrum data. Then, the MFCC extraction method is used to extract features from each signal spectrum data to obtain the corresponding MFCC feature data. All the MFCC feature data are then concatenated to obtain the voiceprint feature data.
[0101] Specifically, the temporary audio data is first denoised to eliminate background noise and other non-speech components, improving the quality of the audio signal. The processed audio data is then segmented into multiple short frames (typically 20-40ms) to facilitate subsequent feature extraction. Frame segmentation simulates the short-term memory characteristics of the human auditory system, helping to capture the dynamic changes in the speech signal. A Fourier transform is performed on each frame of audio data, converting it from the time domain to the frequency domain to obtain the signal's spectral data, displaying the frequency components of that frame.
[0102] Mel-frequency cepstral coefficients (MFCC) are used to extract features from the signal spectral data. MFCC can simulate the sensitivity of the human ear to different frequencies, thereby extracting features useful for speech recognition. The MFCC feature data of each frame are concatenated to form complete speaker signature feature data.
[0103] By accurately extracting voiceprint features and acoustic parameters, speech recognition models can more precisely distinguish between different speakers, significantly improving the reliability of recognition results. Voiceprint feature data, as unique biometric information of the speaker, can be widely used in personalized services, such as voice-activated locks and personalized recommendations. Simultaneously, accurate identification of acoustic parameters can optimize the response speed of voice interaction systems, making the communication experience more natural and fluid.
[0104] The multimodal feature data extraction process is highly efficient, enabling the system to quickly process large amounts of audio data, allocate computing resources rationally, and effectively reduce processing latency. This not only improves the accuracy and reliability of speech recognition but also provides strong support for the expansion of personalized services and the continuous improvement of user experience.
[0105] Optionally, the acoustic parameter data includes short-time energy, zero-crossing rate, and fundamental frequency standard deviation; the step of identifying the temporary audio data to obtain the acoustic parameter data includes:
[0106] Windowing is applied to each frame of audio data to obtain windowed signal data, and the corresponding short-time energy is calculated from the windowed signal data.
[0107] The number of zero crossings is determined for each frame of audio data, and the zero crossing rate is determined based on the number of zero crossings.
[0108] Based on a preset algorithm, the corresponding baseband data is determined according to each frame of audio data, and the baseband standard deviation is determined according to all the baseband data.
[0109] Specifically, acoustic parameter calculation is a key step in speech signal processing, and real-time calculation of these parameters can provide characteristic information of the speech signal.
[0110] Short-time energy calculation: The speech signal is divided into frames, typically 20-40 milliseconds long. In real-time processing, short-time variations in the signal are extracted frame by frame. Each frame is windowed (e.g., a Hamming window) to reduce spectral leakage. The sum of squares of the windowed signal is then calculated using the following formula:
[0111] ;
[0112] in, It is the first i The nth sampling point of the frame, It is the frame length. For window functions (such as Hamming windows). Let be the short-time energy of the nth sampling point.
[0113] Zero-crossing rate calculation: The speech signal is processed by framing, consistent with the framing method used in short-time energy calculation. Within each frame, the number of symbol changes between adjacent sampling points is the number of zero-crossings. If the product of adjacent sampling points is negative, the symbol changes once.
[0114] Zero-crossing rate calculation: The zero-crossing rate is determined by dividing the number of zero-crossings by the frame length.
[0115] Fundamental frequency (F0) standard deviation calculation: Based on a preset algorithm (such as autocorrelation method, cepstral method, etc.), fundamental frequency (F0) data is extracted from each frame of audio data to reflect the pitch information of speech.
[0116] Fundamental frequency standard deviation: Calculates the standard deviation of fundamental frequency data for all frames, measures the dispersion of fundamental frequency changes, and reflects the stability of the speech signal.
[0117] Optionally, such as Figure 2 As shown, the real-time voice call transcription system also includes a mixing scheduling engine, which is specifically used for:
[0118] Obtain current call demand data and engine load status, and determine current load data based on the engine load status;
[0119] The current call demand data is input into a preset load prediction model to obtain load prediction data;
[0120] Based on preset weighted data, target load data is obtained according to the current load data and the predicted load data;
[0121] The target load data is compared with the preset load data to obtain a comparison result, and the real-time voice call transcription system is dynamically allocated resources based on the comparison result; wherein, the load prediction model is constructed based on machine learning algorithms.
[0122] Specifically, this real-time voice call transcription system optimizes resource allocation and processing efficiency based on the engine's load status and call demand data through a dynamic scheduling mechanism of a mixed-stream scheduling engine. First, it collects relevant information about the current call (such as call duration, number of participants, and voice data volume) as call demand data and monitors system resource usage (such as CPU, memory, and network bandwidth) to determine the current load data. Then, this data is input into a load prediction model built based on machine learning algorithms to obtain predicted load data. Next, the target load data is calculated by combining the current load and predicted load according to preset weights. Finally, the target load data is compared with a preset threshold, and system resource allocation is dynamically adjusted based on the result.
[0123] In some embodiments, call demand data includes:
[0124] User request data: By deploying a request capture interface in the access layer of the network element module, we can listen to user-initiated call requests (including voice, video, and text calls) in real time and parse parameters such as service type, target terminal address, and media encoding format in the request protocol (such as SIP and WebRTC).
[0125] Session status data: Get the number of call sessions currently in the process of being established, in progress, or waiting, as well as the network access type (4G / 5G / Wi-Fi) and geographical location (resolved via IP or GPS) of the corresponding terminal.
[0126] Engine load status: Monitors the resource usage of the real-time voice call transcription system, including indicators such as CPU, memory, and network bandwidth, to determine the current load data.
[0127] In some embodiments, engine load states include:
[0128] Hardware resource monitoring: Collect CPU utilization (single-core / multi-core), memory usage (physical memory / swap partition), disk I / O throughput (read / write speed), and network interface traffic (inbound / outbound bandwidth) through operating system APIs (such as procfs in Linux and PDH in Windows).
[0129] Engine kernel metrics collection: Implant probes at the engine framework layer to obtain thread pool status (number of idle threads / number of busy threads), task queue depth (number of pending encoding / decoding tasks), session processing latency (time difference from receiving a request to transmitting a media stream), and error log counts (such as the number of encoding / decoding failures and the number of resource preemption conflicts).
[0130] Distributed tracing system: For microservice architectures, OpenTelemetry or Jaeger is used to trace cross-node call chains and calculate the load balancing factor (such as the number of node connections and the percentile of request processing time) of each node in the distributed engine cluster.
[0131] Call demand data is input into a pre-defined load prediction model. This model, often based on machine learning algorithms, predicts future load demands based on historical and real-time call demand data. Current call demand data is then input into the model to obtain load prediction data, helping the system anticipate changes in resource demand.
[0132] Target load data is calculated based on preset weighted data: the weights set according to the system strategy are used to balance the importance of the current load and the predicted load.
[0133] Based on the combined current load data and predicted load data, the target load data is calculated according to the preset weights. The target load data = α × current load data + β × predicted load data, where α and β are weighting coefficients, and α + β = 1.
[0134] The target load data is compared with a preset load threshold. If the target load exceeds the threshold, resource allocation is automatically optimized, such as adjusting parameters like sampling rate, bit rate, and frame length, or allocating more resources. Conversely, resource input is reduced to improve efficiency. In other words, scheduling is dynamically adjusted in real time based on the comparison results to ensure efficient and stable operation.
[0135] By monitoring engine load status in real time and dynamically adjusting resource allocation, overload can be effectively prevented, significantly improving stability. Simultaneously, resources are precisely allocated according to actual needs, avoiding waste and thus improving resource utilization efficiency. This mechanism can quickly respond to load changes, enabling the system to adapt to different call scenarios and enhancing flexibility. Furthermore, even under high load, it maintains stable operation, providing high-quality transcription services and thus enhancing user satisfaction. Avoiding over-provisioning of resources also helps reduce hardware and energy costs, maximizing cost-effectiveness.
[0136] In summary, this mechanism, through real-time monitoring and dynamic adjustment of resource allocation, not only improves system stability and resource utilization efficiency, but also enhances system flexibility, optimizes user experience, and reduces operating costs.
[0137] Optionally, the step of dynamically allocating resources to the real-time voice call transcription system based on the comparison result includes:
[0138] When the target load data is less than the preset load data, the real-time voice call transcription system is dynamically scheduled according to the round-robin allocation algorithm.
[0139] When the target load data is greater than or equal to the preset load data, the QoS level and corresponding current link latency of each call request in the real-time voice call transcription system are determined.
[0140] The weight data corresponding to the call request is determined based on each QoS level and the corresponding current link latency; the real-time voice call transcription system is dynamically scheduled based on each weight data.
[0141] Specifically, the calculated target load data is compared with a preset load threshold. If the target load data is less than the preset load data, a round-robin allocation algorithm will be used for dynamic scheduling. The round-robin algorithm ensures that all requests receive a fair chance of processing and is suitable for scenarios with light loads.
[0142] When the target load data is greater than or equal to the preset load data, the Quality of Service (QoS) level and corresponding link latency for each call request are first determined. The QoS level reflects the priority and importance of the call request, and the link latency represents the delay time for data transmission. Based on the QoS level and current link latency of each call request, corresponding weight data is calculated. The weight data is used to measure the priority of each request. The calculated weight data is used to dynamically schedule the real-time voice call transcription system, prioritizing requests with higher weights. This scheduling strategy ensures that critical requests receive priority service while effectively managing resources to cope with high load conditions.
[0143] Dynamically adjusting resource allocation optimizes resource usage based on actual load, avoiding waste. Under high load, high-weight requests are prioritized to expedite critical task response, thereby improving user satisfaction and service quality. The dynamic scheduling mechanism prevents system overload, enhances stability and reliability, and quickly adapts to load changes, flexibly responding to various call scenarios and needs.
[0144] In summary, this dynamic scheduling mechanism based on load comparison results can optimize resource allocation, improve system response speed and stability, enhance user experience and system adaptability, and ensure efficient system operation under different load conditions through precise weight calculation and scheduling strategies.
[0145] Quality of Service (QoS) level: This is an attribute related to a task or flow, reflecting the task's service quality requirements. It may include different requirements for latency, bandwidth, reliability, etc. For example, applications like real-time video calls typically have high QoS requirements because they cannot tolerate high latency; while file downloads may have relatively lower latency requirements.
[0146] Current link delay: This is a network state-related attribute that indicates the transmission delay of data within the current network link. Lower latency means faster data transmission speeds, which is crucial for applications requiring low latency (such as online games and real-time video streaming).
[0147] In some embodiments,
[0148] ;
[0149] in, It is the weight of task j, used to determine its scheduling priority. Representative task j The service quality level. This represents the current link latency. and It is an adjustment factor used to balance the impact of QoS and latency on the weight.
[0150] Optionally, the process of constructing the preset vocabulary database includes:
[0151] Obtain the corresponding historical call text data, and identify the historical call text data based on the BiLSTM-CRF model to obtain keyword information, which includes terminology data and keyword data;
[0152] Based on the keyword information, the corresponding weight information is determined, and the keyword information is sorted and filtered according to the corresponding weight information to obtain high-frequency terms and personalized keywords.
[0153] The corresponding preset vocabulary library is determined based on the high-frequency terms and the personalized keywords.
[0154] Specifically, historical call text data of users is collected, which will serve as the foundation for building a vocabulary database. The BiLSTM-CRF model is used to process the historical call text data to identify keyword information. Keyword information includes terminology data (such as professional terms and industry jargon) and general keyword data (such as frequently occurring words expressing the user's core intent).
[0155] Based on certain rules or algorithms, a corresponding weight is assigned to each keyword. This weight reflects the importance and relative priority of the keyword.
[0156] In some implementations, the BiLSTM-CRF model is used to deeply mine terminology and keyword data from users' historical call texts. The training extraction process standards include: the frequency of terminology data must be 2 times or more per 1,000 words; the keyword data must have a TF-IDF value greater than 0.8 and not be included in general dictionaries.
[0157] Based on this, weight information corresponding to keyword information (terminology data and keyword data) is generated. The formula for the weight information is:
[0158] ;
[0159] in, This provides the weight information for the t-th term or keyword data. The frequency of the t-th term or keyword data; For contextual information entropy, and It is the adjustment coefficient.
[0160] Based on keyword weight information, keywords are sorted, and high-frequency terms and personalized keywords with higher weights are selected. This process involves sorting the terminology and keyword data separately, and then filtering according to preset criteria such as the number of terms. A preset vocabulary is constructed based on the selected high-frequency terms and personalized keywords. This vocabulary will be used in subsequent speech-to-text and recognition processes to improve the accuracy and personalization of the transcription.
[0161] The pre-built vocabulary database aims to reflect users' specific language habits and needs, thereby improving the accuracy and personalization of speech transcription and recognition. This database contains personalized keywords to meet the specific needs of different users, significantly enhancing the personalization of the service and the user experience. Through weighted sorting and filtering, high-frequency and key keywords are retained, effectively reducing redundancy in the vocabulary database and optimizing resource utilization efficiency.
[0162] Furthermore, the preset vocabulary database is dynamically updated, adjusting based on the user's historical call data to adapt to changes in the user's language habits. This dynamic update mechanism enhances the system's adaptability and flexibility. During speech recognition, the preset vocabulary database helps accelerate the recognition and understanding of keywords, thereby improving overall recognition efficiency.
[0163] In conclusion, the process of building a pre-defined vocabulary database is crucial for improving the performance and user experience of a real-time speech transcription system, and is a key step in achieving efficient, accurate, and personalized speech recognition.
[0164] Optionally, the default vocabulary database uses a two-level tree indexing mechanism for its storage architecture:
[0165] The purpose of tree-structured indexes is to solve the problems of efficient storage and dynamic retrieval of large-scale personalized data, while the purpose of hierarchical structure is to determine term priority.
[0166] Level 1 (L1): Industry-specific terminology, which is systematically managed through the establishment of standardized industry classification codes, and users have the right to choose independently;
[0167] The second level (L2) consists of personalized vocabulary, which can be generated using the learning capabilities of large language models. It also incorporates a dynamic learning algorithm that automatically updates the weight matrix of the vocabulary database every time 20 new call records are added.
[0168] Specifically, a tree-structured index addresses the challenges of efficient storage and dynamic retrieval of large-scale personalized data. The hierarchical design helps prioritize terms. The vocabulary is divided into two layers, each responsible for managing different types of data.
[0169] Level 1 (L1): Industry Terminology Content: Storage industry terminology, ensuring coverage of professional vocabulary from different fields.
[0170] Management method: Systematic management is carried out through standardized industry classification codes, which facilitates classification and retrieval.
[0171] User permissions: Users can choose whether or not to use these technical terms according to their own needs.
[0172] Level 2 (L2): Personalized vocabulary content: Stores users' personalized vocabulary, reflecting their specific language habits and needs.
[0173] Leveraging the learning capabilities of large language models, new personalized vocabulary can be generated based on users' historical data. The dynamic update process utilizes a built-in dynamic learning algorithm, automatically updating the weight matrix of the vocabulary database every 20 new call records to reflect the latest changes in users' language habits.
[0174] The tree-structured indexing mechanism efficiently stores large amounts of data and supports rapid retrieval. By storing and dynamically updating personalized vocabulary, it better adapts to specific user needs, providing personalized speech-to-text and recognition services. Users have the autonomy to choose industry-specific terminology, which can be adjusted and optimized according to actual needs. Furthermore, dynamic learning algorithms enable the vocabulary database to be continuously updated as user habits change, ensuring its accuracy and relevance. Combining industry-specific terminology with personalized vocabulary allows for more accurate recognition and transcription of speech content, effectively reducing errors and misunderstandings.
[0175] In summary, the pre-defined vocabulary's two-layer tree-structured indexing mechanism not only improves storage and retrieval efficiency but also enhances adaptability and transcription accuracy through dynamic updates of personalized vocabulary and flexible management of industry terminology. This architecture provides users with an efficient and personalized speech-to-text solution.
[0176] Optionally, such as Figure 2 As shown, the real-time voice call transcription system also includes a voice sentence segmentation module:
[0177] The speech segmentation module is used to segment each frame of audio data according to a preset segmentation trigger condition;
[0178] The preset sentence segmentation triggering conditions include silence duration triggering conditions, fundamental frequency standard deviation triggering conditions, and interjection triggering conditions.
[0179] Specifically, the preset sentence segmentation triggering conditions include silent duration triggering conditions, which are speech pauses detected based on short-time energy. The short-time energy calculation process: Short-time energy is the energy of a signal within a certain short period of time, usually obtained by summing the squares of the signal over that time period.
[0180] Its calculation formula:
[0181] ;
[0182] in, Let i be the signal value in the i-th frame. For window functions, is the frame length, i.e., the total number of sampling points, and n is the nth sampling point of the signal in the current frame.
[0183] Point detection thresholds are used: Sound: short-time energy greater than 0.1. Unsound: short-time energy between 0.01 and 0.1. Silence: short-time energy less than 0.01.
[0184] Sentence segmentation trigger condition: When the duration of silence exceeds the preset configuration (e.g., silence duration ≥ 200ms), sentence segmentation is triggered.
[0185] The preset sentence segmentation trigger conditions include the fundamental frequency standard deviation trigger condition, which determines intonation changes based on the fundamental frequency (F0) standard deviation.
[0186] Fundamental frequency (F0) detection: The fundamental frequency is the lowest frequency component in a speech signal, reflecting the pitch of the speech. Fundamental frequency standard deviation calculation: Calculate the standard deviation of the fundamental frequency (F0) values within consecutive frames.
[0187] Sentence segmentation trigger condition: When the fundamental frequency standard deviation is greater than 25Hz, it indicates that there is a significant change in intonation in the speech, and sentence segmentation is performed at this time.
[0188] The preset sentence segmentation trigger conditions include the modal particle trigger condition, that is, recognizing modal particles to segment speech.
[0189] Interjections are words that express the speaker's attitude or emotion, such as "ah" and "um." Their recognition can be achieved through speech recognition models.
[0190] Sentence segmentation trigger condition: When an interjection is detected at the end of a sentence, the speech is segmented, that is, the current sentence is considered to have ended and a new sentence begins.
[0191] By combining the results of the three methods (silence duration trigger condition, fundamental frequency standard deviation trigger condition, and interjection trigger condition), the system intelligently determines the punctuation points in the speech stream. This multimodal punctuation method can more accurately simulate human perception of speech pauses and intonation changes, thereby improving the naturalness and accuracy of speech transcription and speech recognition.
[0192] By simulating human perception of pauses and intonation changes in speech, the above process significantly improves the naturalness of speech processing, making speech transcription and recognition results more natural and fluent. This method can accurately identify sentence breaks, thereby enhancing the overall accuracy of speech recognition in complex speech scenarios. Furthermore, more natural and accurate speech processing significantly improves the user experience, allowing users to experience more convenient and efficient services when using speech recognition and transcription services. These improvements not only optimize the performance of speech recognition technology but also greatly enhance user satisfaction with related applications.
[0193] Optionally, the real-time transcribed text data includes highlighted text data with timestamps; such as... Figure 2 As shown, the real-time voice call transcription system also includes a subtitle streaming module, which is specifically used for:
[0194] The highlighted text data is packaged to obtain independent subtitle frames, and independent subtitle layers are obtained based on the independent subtitle frames and preset personalized configurations;
[0195] The independent subtitle layer is parsed and rendered to obtain subtitle data, and the subtitle data is pushed to the display device.
[0196] Specifically, the above process is achieved through a subtitle streaming module. Real-time transcribed text data includes timestamped, highlighted text data used to emphasize important information or keywords. The highlighted text data is packaged to generate independent subtitle frames. Each subtitle frame contains text content and display style at a specific point in time. Based on the independent subtitle frames and the user's preset personalized configurations (such as font, color, background, etc.), an independent subtitle layer is generated. This layer can be managed and displayed independently of other video or image content. The independent subtitle layer is parsed and rendered, converting it into displayable subtitle data. The subtitle data is then pushed to a display device (such as a monitor, projector, etc.) for real-time display.
[0197] The combination of highlighting markers and timestamps effectively emphasizes important information, helping users quickly locate key content. The system allows users to customize subtitle styles according to their preferences, such as font, color, and background, thereby optimizing the viewing experience. The independent subtitle frame and layer generation and management mechanism ensures accurate presentation of subtitle information, greatly improving information delivery efficiency. Furthermore, the real-time push display function for subtitle data allows users to instantly access transcription information, effectively enhancing the system's interactivity.
[0198] In summary, this method can efficiently transform real-time transcribed text data into intuitive and easy-to-understand subtitle information. By leveraging various methods such as highlighting, personalized configuration, and real-time push notifications, it significantly improves the efficiency of users acquiring and understanding information, comprehensively optimizing the user experience.
[0199] In some embodiments, the subtitle layer first needs to parse the language data. This step involves converting the raw subtitle data (which may be JSON, XML, or other formats) into a format that can be used for rendering. The parsing process may include extracting and formatting the data structure to ensure that the subtitle information can be correctly identified and processed.
[0200] The parsed data is converted into HTML and CSS formats. HTML defines the subtitle tags and structure, while CSS defines the subtitle styles, such as font, color, and position. This step is crucial for transforming the raw data into a webpage-readable format. Through HTML and CSS, the subtitle data is formatted into a structured document, easy to display on a webpage, and can be beautified and layout adjusted using CSS.
[0201] The WebSocket protocol was chosen to enable bidirectional communication between the server and the client. Compared to the traditional HTTP protocol, WebSocket supports full-duplex communication, making it suitable for real-time update applications, such as real-time caption transmission.
[0202] The formatted HTML / CSS subtitle data is pushed to the client (user end) in real time via the WebSocket protocol. This means that as soon as there is new subtitle data on the server, the client will immediately receive these updates.
[0203] When the client receives HTML / CSS data pushed via WebSocket, the HTML content is parsed to extract key information such as text content and style information. After parsing, this text information is displayed on the terminal screen (such as a computer, mobile phone, or tablet). This step may involve DOM manipulation, adding HTML content to specific parts of the webpage, and applying CSS styles.
[0204] By using the WebSocket protocol, subtitles can be displayed on the client's screen in real time, ensuring users don't miss any key content. This not only enhances the user's viewing experience but also increases their engagement. Furthermore, the subtitles displayed in HTML / CSS format support flexible style customization, allowing users to adjust the appearance of the subtitles according to their personal preferences. In short, this process improves information delivery efficiency and user experience by converting subtitle data into HTML / CSS format and pushing it to the client in real time via the WebSocket protocol, enabling users to view subtitles instantly on their terminal devices.
[0205] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, data stored, data displayed, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties. The collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0206] like Figure 3 As shown, this embodiment of the invention provides a real-time voice call transcription method, applied to the aforementioned real-time voice call transcription system. The transcription method includes:
[0207] When a call request is detected from the user's end, the corresponding audio data is retrieved.
[0208] Based on a preset perceptual weighted vector quantization algorithm, the audio data is compressed in layers to obtain compressed audio data, and the compressed audio data is then converted to obtain temporary audio data.
[0209] Feature extraction is performed on the temporary audio data to obtain multimodal feature data, and the multimodal feature data is processed based on a preset speech recognition model to obtain text information;
[0210] Based on a pre-set large model, corresponding real-time transcribed text data is obtained according to the text information and a pre-set vocabulary.
[0211] While the present invention has been disclosed above, its scope of protection is not limited thereto. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention, and all such changes and modifications will fall within the scope of protection of the present invention.
Claims
1. A real-time voice call transcription system, characterized in that, It includes network element modules, voice streaming engine, voice engine, and analysis and optimization modules; The network element module is used to obtain corresponding audio data when a call request from a user terminal is detected. The voice streaming engine is used to perform layered compression of the audio data based on a preset perceptual weighted vector quantization algorithm to obtain compressed audio data, and to perform format conversion processing on the compressed audio data to obtain temporary audio data. The speech engine is used to extract features from the temporary audio data to obtain multimodal feature data, and to process the multimodal feature data based on a preset speech recognition model to obtain text information. The multimodal feature data includes voiceprint feature data and acoustic parameter data; the step of extracting features from the temporary audio data to obtain multimodal feature data includes: Feature extraction is performed on the temporary audio data to obtain the voiceprint feature data; The temporary audio data is identified to obtain the acoustic parameter data; The analysis and optimization module is used to obtain corresponding real-time transcribed text data based on the text information and the preset vocabulary library, according to the preset large model. The construction process of the preset vocabulary database includes: Obtain the corresponding historical call text data, and identify the historical call text data based on the BiLSTM-CRF model to obtain keyword information, which includes terminology data and keyword data; Based on the keyword information, the corresponding weight information is determined, and the keyword information is sorted and filtered according to the corresponding weight information to obtain high-frequency terms and personalized keywords. The corresponding preset vocabulary library is determined based on the high-frequency terms and the personalized keywords.
2. The real-time voice call transcription system according to claim 1, characterized in that, The step of extracting features from the temporary audio data to obtain the voiceprint feature data includes: The temporary audio data is denoised to obtain processed audio data, and the processed audio data is then divided into frames to obtain multiple framed audio data. A Fourier transform is performed on each frame of audio data to obtain the corresponding signal spectrum data. Then, the MFCC extraction method is used to extract features from each signal spectrum data to obtain the corresponding MFCC feature data. All the MFCC feature data are then concatenated to obtain the voiceprint feature data.
3. The real-time voice call transcription system according to claim 2, characterized in that, The acoustic parameter data includes short-time energy, zero-crossing rate, and fundamental frequency standard deviation; the process of identifying the temporary audio data to obtain the acoustic parameter data includes: Windowing is applied to each frame of audio data to obtain windowed signal data, and the corresponding short-time energy is calculated from the windowed signal data. The number of zero crossings is determined for each frame of audio data, and the zero crossing rate is determined based on the number of zero crossings. Based on a preset algorithm, the corresponding baseband data is determined according to each frame of audio data, and the baseband standard deviation is determined according to all the baseband data.
4. The real-time voice call transcription system according to claim 1, characterized in that, The real-time voice call transcription system also includes a streaming scheduling engine, which is specifically used for: Obtain current call demand data and engine load status, and determine current load data based on the engine load status; The current call demand data is input into a preset load prediction model to obtain load prediction data; Based on preset weighted data, target load data is obtained according to the current load data and the predicted load data; The target load data is compared with the preset load data to obtain a comparison result, and the real-time voice call transcription system is dynamically allocated resources based on the comparison result; wherein, the load prediction model is constructed based on machine learning algorithms.
5. The real-time voice call transcription system according to claim 4, characterized in that, The step of dynamically allocating resources to the real-time voice call transcription system based on the comparison results includes: When the target load data is less than the preset load data, the real-time voice call transcription system is dynamically allocated resources according to the polling allocation algorithm. When the target load data is greater than or equal to the preset load data, the QoS level and corresponding current link latency of each call request are determined; the weight data corresponding to the call request is determined according to each QoS level and the corresponding current link latency; and the voice call real-time transcription system is dynamically allocated resources according to each weight data.
6. The real-time voice call transcription system according to claim 3, characterized in that, The real-time voice call transcription system also includes a voice segmentation module: The speech segmentation module is used to segment each frame of audio data according to a preset segmentation trigger condition; The preset sentence segmentation triggering conditions include silence duration triggering conditions, fundamental frequency standard deviation triggering conditions, and interjection triggering conditions.
7. The real-time voice call transcription system according to claim 1, characterized in that, The real-time transcribed text data includes highlighted text data with timestamps; the real-time speech transcription system also includes a subtitle streaming module, which is specifically used for: The highlighted text data is packaged to obtain independent subtitle frames, and independent subtitle layers are obtained based on the independent subtitle frames and preset personalized configurations; The independent subtitle layer is parsed and rendered to obtain subtitle data, and the subtitle data is pushed to the display device.
8. A method for real-time transcription of voice calls, characterized in that, The real-time voice call transcription system as described in any one of claims 1 to 7, wherein the transcription method comprises: When a call request is detected from the user's end, the corresponding audio data is obtained; Based on a preset perceptual weighted vector quantization algorithm, the audio data is compressed in layers to obtain compressed audio data, and the compressed audio data is then converted to obtain temporary audio data. Feature extraction is performed on the temporary audio data to obtain multimodal feature data, and the multimodal feature data is processed based on a preset speech recognition model to obtain text information; Based on a pre-set large model, corresponding real-time transcribed text data is obtained according to the text information and a pre-set vocabulary.
Citation Information
Patent Citations
Scheduling call real-time transferring system and method based on regulation and control voice recognition
CN115567637A
Talent evaluation management method and system based on AI intelligence
CN120338741A
Language learning auxiliary application system based on speech recognition
CN120356458A