Intelligent terminal conference voice transcription method and system

CN122551775APending Publication Date: 2026-08-11SHENZHEN ZHILIANMAO TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-21
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

这种关键信息的错漏和丢失,导致转写文本语义混乱,用户不得不花费大量时间进行人工校对和修正,使得语音转写技术难以在专业会议场景中真正落地应用,无法有效支撑后续的会议纪要生成、待办事项提取等智能化工作的问题

Benefits of technology

[0009]The beneficial effects of this application are as follows: This application first acquires the voice information of the intelligent terminal conference, including user voice information and environmental voice information. Then, it extracts acoustic features and initial text features based on the user voice information, integrates the two to identify conference domain labels and calls the corresponding professional dictionary library to achieve domain adaptation. Then, it extracts environmental noise features and conference scene context features based on the environmental voice information, combines the domain labels to generate adaptive noise reduction parameters, performs targeted noise reduction and event gain adjustment on the user voice, and obtains the initial transcribed text. Next, it matches and verifies the initial transcribed text with the professional dictionary library to locate the misidentified segment sequence. Finally, it uses contextual semantics and acoustic features to generate candidate word sequences, intelligently replaces the misidentified segments, and outputs the final conference transcribed text. This invention significantly improves the accuracy of speech transcription in professional conference scenarios by integrating multi-source information, domain-adaptive dictionaries, and context-aware error correction, avoiding the problem of misidentification of professional terms by general models. At the same time, it enhances robustness in complex acoustic environments through environmental adaptive noise reduction, so that the transcription results can be directly used for subsequent intelligent processing such as conference minutes generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122551775A_ABST
    Figure CN122551775A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of voice processing, and particularly discloses a conference voice transcription method and system for an intelligent terminal. The application first acquires user voice information and environmental voice information, then extracts acoustic features and initial text features according to the user voice information, fuses the two to identify conference field labels and call corresponding professional dictionary libraries, realizes field self-adaptation, extracts environmental noise features and conference scene context features according to the environmental voice information, combines the field labels to generate self-adaptive noise reduction parameters, obtains initial transcription text, then matches and verifies the initial transcription text with the professional dictionary library, positions a misrecognized segment sequence, finally generates a candidate word sequence by using context semantics and acoustic features, intelligently replaces the misrecognized segment, and outputs a final conference transcription text. The application significantly improves the voice transcription accuracy in a professional conference scene by fusing multi-source information, a field self-adaptive dictionary and context perception error correction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech processing technology, and in particular to a method and system for transcribing conference speech on a smart terminal. Background Technology

[0002] With the widespread adoption of remote work and intelligent conferencing, speech-to-text functionality on smart devices (such as mobile phones, tablets, and dedicated conference tablets) has become a key tool for improving meeting efficiency. Existing speech-to-text methods typically employ general speech recognition models to convert speech into text in real time. However, in practical applications, it has been found that while these general models perform reasonably well in everyday conversations, their recognition accuracy drops sharply when faced with specialized meetings in vertical fields (such as law, healthcare, IT architecture, and finance).

[0003] Specifically, professional conferences are filled with a large number of domain-specific technical terms (such as "epigenetics" and "equivalence relationship"), English abbreviations (such as "API" and "SLA"), and mixed Chinese-English expressions (such as "this bug needs a hotfix"). General-purpose speech recognition models, lacking domain knowledge, often misidentify these technical terms as similar-sounding general terms. For example, they might misidentify "Transformer architecture" in the IT field as "transmitter architecture," or "market maker" in the financial field as "doing business." This omission and loss of key information leads to semantic confusion in the transcribed text, forcing users to spend considerable time on manual proofreading and correction. This makes it difficult for speech-to-text technology to be truly applied in professional conference scenarios and effectively support subsequent intelligent tasks such as meeting minutes generation and to-do list retrieval. Therefore, a smart terminal conference speech-to-text method and system are needed to solve these problems. Summary of the Invention

[0004] The purpose of this invention is to provide a method and system for transcribing voice in smart terminal conferences, so as to solve the technical problems mentioned in the background art.

[0005] To achieve the above objectives, this application provides a method for transcribing conference speech on a smart terminal, comprising: Acquire voice information from a smart terminal conference, wherein the voice information includes user voice information and environmental voice information; Acoustic features and initial text features are obtained based on the user's voice information, and conference domain tags are obtained based on the acoustic features and initial text features, and corresponding professional dictionary databases are obtained based on the conference domain tags; Environmental noise features and meeting scenario context features are obtained based on the environmental speech information, and adaptive noise reduction parameters are obtained based on the environmental noise features and the meeting domain labels. Initial transcribed text is obtained based on the adaptive noise reduction parameters and the meeting scenario context features. The initial transcribed text is matched and verified against the professional dictionary database to obtain a sequence of misidentified fragments; Contextual semantics are obtained based on the initial text features, and candidate word sequences are obtained based on the contextual semantics and the acoustic features. The candidate word sequences are then used to replace the misidentified fragment sequences in the initial transcribed text to obtain the final conference transcribed text.

[0006] This application also provides a smart terminal conference voice transcription system, including: A voice information acquisition module is used to acquire voice information from smart terminal conferences, wherein the voice information includes user voice information and environmental voice information; The domain recognition module is used to obtain acoustic features and initial text features based on the user's voice information, obtain conference domain labels based on the acoustic features and initial text features, and obtain the corresponding professional dictionary based on the conference domain labels; The initial transcription module is used to obtain environmental noise features and meeting scene context features based on the environmental speech information, obtain adaptive noise reduction parameters based on the environmental noise features and the meeting domain label, and obtain the initial transcribed text based on the adaptive noise reduction parameters and the meeting scene context features. The matching and verification module is used to match and verify the initial transcribed text with the professional dictionary database to obtain a sequence of misidentified segments; The error correction module is used to obtain contextual semantics based on the initial text features, and obtain candidate word sequences based on the contextual semantics and the acoustic features. The candidate word sequences are then used to replace the misidentified segment sequences in the initial transcribed text to obtain the final conference transcribed text.

[0007] This application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of any of the methods described above.

[0008] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the methods described above.

[0009] The beneficial effects of this application are as follows: This application first acquires the voice information of the intelligent terminal conference, including user voice information and environmental voice information. Then, it extracts acoustic features and initial text features based on the user voice information, integrates the two to identify conference domain labels and calls the corresponding professional dictionary library to achieve domain adaptation. Then, it extracts environmental noise features and conference scene context features based on the environmental voice information, combines the domain labels to generate adaptive noise reduction parameters, performs targeted noise reduction and event gain adjustment on the user voice, and obtains the initial transcribed text. Next, it matches and verifies the initial transcribed text with the professional dictionary library to locate the misidentified segment sequence. Finally, it uses contextual semantics and acoustic features to generate candidate word sequences, intelligently replaces the misidentified segments, and outputs the final conference transcribed text. This invention significantly improves the accuracy of speech transcription in professional conference scenarios by integrating multi-source information, domain-adaptive dictionaries, and context-aware error correction, avoiding the problem of misidentification of professional terms by general models. At the same time, it enhances robustness in complex acoustic environments through environmental adaptive noise reduction, so that the transcription results can be directly used for subsequent intelligent processing such as conference minutes generation. Attached Figure Description

[0010] Figure 1 This is a schematic diagram of a method flow according to an embodiment of this application.

[0011] Figure 2 This is a schematic diagram of the system structure according to an embodiment of this application.

[0012] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0013] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0014] like Figure 1 As shown, this application provides a method and system for transcribing voice messages in a smart terminal conference, including: S1. Obtain voice information from the smart terminal conference, wherein the voice information includes user voice information and environmental voice information; S2. Obtain acoustic features and initial text features based on the user's voice information, obtain conference domain tags based on the acoustic features and initial text features, and obtain the corresponding professional dictionary based on the conference domain tags; S3. Obtain environmental noise features and meeting scene context features based on the environmental speech information, obtain adaptive noise reduction parameters based on the environmental noise features and the meeting domain label, and obtain initial transcribed text based on the adaptive noise reduction parameters and the meeting scene context features; S4. Match and verify the initial transcribed text with the professional dictionary database to obtain the sequence of misidentified segments; S5. Obtain contextual semantics based on the initial text features, and obtain candidate word sequences based on the contextual semantics and the acoustic features. Replace the misidentified segment sequences with the candidate word sequences and output the final conference transcription text.

[0015] As described in steps S1-S5 above, this invention acquires voice information from a smart terminal conference, including user voice information and environmental voice information. This step is the signal acquisition and separation stage of the entire transcription process. Its physical significance lies in separating the target signal for semantic transcription and the interference signal for environmental analysis from the original continuous audio signal acquired from the microphone array of the smart terminal, providing a clear input source for subsequent feature extraction and processing. The voice information is acquired by the multi-channel microphone built into the smart terminal at a sampling rate of 16kHz and a quantization precision of 16bit. After acquisition, the effective human voice segments are marked as user voice information, and the silent segments and non-human voice interference segments are marked as environmental voice information through a voice activity detection algorithm. This separation process can avoid feature interference caused by the mixed processing of effective voice and environmental noise, providing an independent and accurate data foundation for subsequent domain recognition and environmental noise reduction. Compared with the method of directly inputting mixed audio into the model, it can reduce the interference of noise on acoustic features and text feature extraction, and improve the stability of subsequent domain determination.

[0016] Acoustic features and initial text features are obtained from user voice information, and then conference domain labels are obtained based on these features. A corresponding professional dictionary is then retrieved based on the conference domain labels. This step is the core of domain-adaptive matching. Its physical meaning lies in extracting dual features from user voice that characterize both pronunciation and text semantics. Feature fusion enables accurate determination of the conference domain, and then the professional dictionary of the matching domain is invoked to reduce the probability of misidentification of professional terms at the decoding baseline level. In specific implementation, the user voice information is first processed by framing and windowing with a 25ms frame length and a 10ms frame shift. Hamming windows are used to suppress spectral leakage. The time-domain and frequency-domain energy values ​​of each voice frame are calculated to form a time-domain-frequency-domain energy vector as the acoustic feature vector. Then, the user voice is input into an end-to-end speech recognition model based on the Conformer architecture. Greedy decoding generates an initial recognition text string, which is then processed... Maximum positive matching word segmentation is performed, and part-of-speech tagging is completed using a Hidden Markov Model. The segmented word sequence and the part-of-speech tag sequence are used as the initial text feature vector. The acoustic feature vector is concatenated with the initial text feature vector and then input into a pre-trained support vector machine classification model. This model adopts a one-to-one multi-classification strategy and is trained on labeled datasets covering fields such as law, medicine, IT, and finance. The model calculates the conference domain label corresponding to the highest probability and then retrieves the corresponding professional dictionary from a pre-set dictionary library based on the label. This dual-feature fusion classification method can avoid the domain judgment bias caused by initial recognition errors due to relying solely on text features. At the same time, the professional dictionary library provides a domain-specific benchmark for subsequent text verification, directly solving the problem that general models do not have the ability to recognize professional terms. For example, in IT conferences, it can accurately match professional terms such as Transformer architecture and API, reducing the erroneous substitution of homophones.

[0017] The process involves obtaining environmental noise features and meeting scenario context features from environmental speech information, acquiring adaptive noise reduction parameters based on these features and meeting domain labels, and then obtaining initial transcribed text based on these parameters and the meeting scenario context features. This step is the environmental adaptive noise reduction and initial decoding stage. Its physical meaning lies in constructing a noise reduction strategy adapted to the current meeting based on the noise distribution, scene events, and domain features of the meeting scenario. This strategy aims to suppress noise while preserving effective audio segments in the corresponding domain, thereby improving the quality of the initial transcribed text. Specifically, the process first involves detecting silent segments in the environmental speech information, extracting pure environmental noise samples, and performing power spectral density analysis to obtain the noise distribution at each frequency point as environmental noise features. Then, an acoustic event detection model with a hybrid architecture of CNN and GRU is used to identify events such as multiple people speaking simultaneously, applause, projector noise, and their corresponding time intervals, forming meeting scenario context features. Finally, the signal-to-noise ratio at each frequency point and the pre-set domain-sensitive frequency band weights based on the meeting domain labels are used to determine the final transcribed text.

[0018] For example, in Case 1, legal conferences are characterized by formal and rigorous language, relying heavily on formal legal terminology ("consideration," "apparent agency," "force majeure"). Pronunciation is clear, the speaking speed is relatively steady, and information is primarily concentrated in the mid-frequency range. High-frequency noise (such as paper rubbing or keyboard tapping) significantly interferes with the identification of key terms. Sensitive frequency band weighting: Mid-frequency band (300Hz-3kHz): Weight = 0.9-1.0. This is the area most sensitive to human hearing, containing the vast majority of the energy and intelligibility information of legal terms; Low frequency range (<300Hz): Weight = 0.5-0.6. A portion is retained to ensure natural speech, while moderately suppressing room echo. High-frequency band (>3kHz): Weight = 0.3-0.5. High-frequency noise (such as keyboard sounds) is strongly suppressed here. The distinction of legal terms relies less on extremely high-frequency information, so noise reduction is safe.

[0019] When someone types on a keyboard (high-frequency noise), the noise reduction parameters will form a very low filtering coefficient in the high-frequency band, effectively eliminating the "click-click-click" interference sound. At the same time, key terms such as "apparent agency" in the mid-frequency band are fully protected, ensuring the legal accuracy of the transcription and preventing "apparent agency" from being misinterpreted as "apparent subrogation".

[0020] Example 2: In IT (Information Technology) conferences, the characteristics of the field include: a large number of English abbreviations (API, SDK, CPU), mixed English and Chinese terms ("deploy a service", "this bug needs to be fixed"), and specific technical terms (Transformer, Kubernetes, React). The pronunciation of these words usually contains clear high-frequency consonants (such as / p / , / t / , / k / , / s / ) and fricatives; Sensitive frequency band weighting: High frequency band (2kHz-6kHz): Weight = 0.9-1.0. This frequency band contains a large amount of consonant intelligibility information, which is crucial for distinguishing between "API" and "ABI", "Kubernetes" and "cooperatives"; Mid-frequency band (300Hz-2kHz): Weight = 0.6-0.8. Contains vowels and most semantic information, relatively important; Low frequency band (<300Hz): weight = 0.2-0.4. Mainly consists of ambient noise (such as air conditioners and fans) and the fundamental frequency of speech, contributing little to distinguishing technical terms.

[0021] When encountering low-frequency noise from the projector fan, the noise reduction algorithm significantly suppresses the frequency band <300Hz (because of its low weight), but hardly suppresses the high-frequency band of 2kHz-6kHz. This ensures that the high-frequency pronunciation of the "former" part in "Transformer" is completely preserved, so that it can be correctly captured by the recognition model, rather than being mistakenly identified as the "transmitter" after being filtered out as noise.

[0022] Then, a frequency-domain adaptive filtering coefficient matrix is ​​generated as an adaptive noise reduction parameter. This parameter is used to perform frequency-domain filtering on the user's speech information. Then, the gain is adjusted for specific time periods based on the contextual features of the meeting scenario. For example, the gain is increased for multi-person speaking segments and decreased for applause segments. The processed speech signal is input into a general speech recognition model for decoding to obtain the initial transcribed text. This adaptive noise reduction combined with domain features can adjust the filtering intensity according to the sensitive frequency bands of speech in different domains, avoiding the loss of effective speech details caused by general fixed noise reduction algorithms. At the same time, the gain adjustment triggered by scene events can cope with sudden interference, so that the initial transcribed text still maintains a high basic accuracy in complex meeting environments.

[0023] The initial transcribed text is matched and verified against a professional dictionary database to obtain a sequence of misidentified segments. This step is the localization stage for misidentified professional terms. Its physical meaning lies in using a domain-specific dictionary as a benchmark, combined with recognition confidence levels, to filter out segments in the initial transcribed text that have low credibility and do not conform to professional expressions, forming an ordered sequence of misidentified segments. This provides precise targets for subsequent error correction. Specifically, the initial transcribed text is first segmented using the maximum positive matching method, consistent with the previous steps, to obtain an initial word sequence arranged by time. Each word is then precisely matched against the professional dictionary database. Words that fail to match are retrieved... The posterior probability score generated during the decoding process of the general speech recognition model is used as the confidence level. Words with confidence levels below the threshold are marked as candidate words for misidentification. Then, candidate words within adjacent time windows of 5 word lengths are combined into continuous misidentified segments, which are arranged in chronological order to obtain a sequence of misidentified segments. This method of dual verification by dictionary matching and confidence level can accurately locate homophones and near-homophones of professional terms. For example, it can effectively detect segments in the financial field that are misidentified as "market maker" instead of "business maker". Compared with pure grammatical rule verification, it is more in line with the transcription needs of professional conferences and avoids missing core professional vocabulary errors.

[0024] The process involves obtaining contextual semantics from initial text features, generating candidate word sequences based on contextual semantics and acoustic features, replacing misidentified segment sequences with these candidate word sequences, and outputting the final conference transcription text. This step is the final stage of contextual semantics and acoustic feature fusion and error correction. Its physical significance lies in combining the semantic logic before and after the misidentified segment with the acoustic pronunciation features of the original speech to generate replacement words that conform to domain semantics and pronunciation rules, thus achieving accurate correction of the transcribed text. Specifically, this involves first extracting 10 words before and after the misidentified segment from the initial text features as context, inputting them into the BERT-base Chinese language model to obtain a contextual semantic vector, mapping it to a contextual attention weight distribution through a two-layer multilayer perceptron network, and simultaneously extracting segments temporally aligned with the misidentified segment from the acoustic features, and then performing phoneme posterior probabilistic analysis. The process involves rate decoding to obtain an acoustic candidate phoneme sequence, reconstructing the phoneme sequence using contextual attention weights as coefficients, and then generating a probability-sorted candidate word sequence using a beam search decoder with a beam width of 10. The top 5 candidate words with the highest probabilities are selected, and their edit distances to misidentified segments are calculated. Candidate words exceeding a preset threshold of 3 are removed. This threshold is determined based on statistics of edit distances for misidentification of homophones and near-homophones in Chinese. The remaining candidate words with the highest probabilities are used as replacement words. After replacing all misidentified segments, the final transcribed text is output. This error correction method, which integrates contextual semantics and acoustic features, ensures that the replacement words conform to the logic of the sentence, match the original pronunciation features, and adhere to the expression norms of professional fields. Compared with pure text error correction, it can significantly improve the accuracy of correcting professional terms, making the final transcribed text meet the requirements of professional conferences.

[0025] In this embodiment, step S2, which involves obtaining acoustic features and initial text features based on the user's voice information, obtaining conference domain tags based on the acoustic features and the initial text features, and obtaining the corresponding professional dictionary based on the conference domain tags, includes: S21. The user voice information is segmented and windowed to obtain multiple voice frames, and multiple time-domain-frequency-domain energy values ​​are obtained based on the multiple voice frames. S22. Obtain a time-frequency energy vector based on the multiple time-frequency energy values ​​and a preset time axis, and use the time-frequency energy vector as an acoustic feature vector; S23. Input the user's voice information into a pre-trained end-to-end speech recognition model, perform greedy decoding, and generate an initial recognition text string; S24. Perform maximum positive matching word segmentation on the initial identified text string to obtain a word segmentation sequence, and perform part-of-speech tagging on each word in the word segmentation sequence based on a hidden Markov model to obtain a part-of-speech tag sequence. Use the word segmentation sequence and the part-of-speech tag sequence together as the initial text feature vector. S25. Concatenate the acoustic feature vector and the initial text feature vector to obtain a fused feature vector; S26. Input the fused feature vector into a pre-trained support vector machine classification model, calculate the probability that the fused feature vector belongs to each of the preset multiple conference domain categories, and output the category corresponding to the highest probability as the conference domain label; S27. Retrieve the corresponding professional dictionary from the preset dictionary database set according to the conference domain label.

[0026] As described in steps S21-S27 above, this invention performs frame segmentation and windowing on user speech information to obtain multiple speech frames, and obtains multiple time-domain-frequency domain energy values ​​based on these multiple speech frames. This step is a basic preprocessing step for acoustic feature extraction. Its physical significance lies in converting continuous analog speech signals into discrete, easily computed frame sequences while suppressing spectral leakage and ensuring the accuracy of energy calculation. This processing uses a 25-millisecond frame length and a 10-millisecond frame shift configuration, and the window function used is a Hamming window. This configuration is determined based on the short-time stationary characteristics of speech signals, which can ensure the stability of signals within a frame while taking into account the continuity of adjacent frames. After frame segmentation and windowing, a fast Fourier transform is performed on each speech frame to obtain frequency domain energy information. Combined with the time-domain frame energy values, the time-domain-frequency domain energy value corresponding to each speech frame is obtained. This step can convert the original speech signal into computable energy features, providing stable and reliable underlying data for subsequent acoustic feature vector construction, and avoiding the problems of high computational complexity and unstable features caused by direct processing of continuous signals.

[0027] The time-frequency energy vector is obtained based on multiple time-frequency energy values ​​and a preset time axis, and this time-frequency energy vector is used as the acoustic feature vector. This step is the vectorization and normalization of acoustic features. The physical meaning is to normalize the scattered frame-level time-frequency energy values ​​into a fixed-dimensional vector according to the time order, so that it can be directly input into the classification model for calculation. In the processing, the time-frequency energy values ​​corresponding to each frame are arranged sequentially along the preset time axis according to the order of the speech frames, forming a one-dimensional feature vector with a length matching the total number of speech frames. This vector completely preserves the energy distribution characteristics of speech in the time and frequency dimensions, and can reflect the speaker's pronunciation habits, speech rate, intonation, and domain-related pronunciation features. Compared with the traditional method of only using Mel frequency cepstral coefficients, this vector retains more complete spectral energy information and has a better representation ability to distinguish the speech characteristics of different domain conferences.

[0028] User voice information is input into a pre-trained end-to-end speech recognition model, and greedy decoding is performed to generate an initial recognized text string. This step is a prerequisite for obtaining text features. Its physical significance lies in converting the speech signal into a text string, providing raw text data for subsequent text feature extraction. The model used is an end-to-end speech recognition model based on the Conformer architecture. This model combines the advantages of convolutional neural networks in extracting local features and Transformers in modeling global dependencies, making it suitable for long speech recognition in meetings. Greedy decoding uses the output with the highest probability at each step as the current recognition result, resulting in high computational efficiency and fast processing speed, making it suitable for real-time operation on smart terminals. Although the initial recognized text string obtained through this step contains some misrecognition of technical terms, it is sufficient to reflect the overall semantic direction of the meeting and can meet the needs of subsequent text feature extraction.

[0029] The initial identified text string is segmented using maximum forward matching to obtain a segmented word sequence. Each word in the segmented word sequence is then labeled with part-of-speech tags based on a Hidden Markov Model (HMM) to obtain a part-of-speech tag sequence. The segmented word sequence and the part-of-speech tag sequence are used together as the initial text feature vector. This step is the text feature structure extraction stage, which physically decomposes continuous text into semantic units and labels grammatical attributes, transforming text information into computable structured features. Maximum forward matching segmentation uses a general Chinese lexicon combined with a basic professional vocabulary to ensure segmentation accuracy. The HMM part-of-speech tagging can assign corresponding tags to nouns, verbs, professional terms, etc. The segmented word sequence and the part-of-speech tag sequence are combined in sequence to form the initial text feature vector. This vector contains both text content information and grammatical structure information, which can effectively represent the domain tendency of the conference text and provide a reliable text-level basis for subsequent fusion classification.

[0030] The acoustic feature vector and the initial text feature vector are concatenated to obtain the fused feature vector. This step is a two-dimensional feature fusion process. Its physical meaning is to combine the physical acoustic features of speech with the semantic features of text to form a comprehensive feature that can fully represent the domain attributes of the conference. The concatenation method is to directly connect along the feature dimension, maintaining the independence and integrity of the acoustic features and text features. The fused feature vector carries both speech pronunciation characteristics and text semantic attributes, which can effectively avoid the classification bias caused by a single feature. For example, when the initial text recognition is incorrect, the acoustic features can provide compensation information. When the acoustic features are affected by slight noise, the text features can stabilize the classification basis and improve the overall robustness of the domain judgment.

[0031] The fused feature vector is input into a pre-trained support vector machine (SVM) classification model. The probability of the fused feature vector belonging to each of the preset multiple meeting domain categories is calculated, and the category corresponding to the highest probability is output as the meeting domain label. This step is the accurate meeting domain determination stage. Its physical meaning is to classify the comprehensive features through the machine learning model and output the domain identifier that the current meeting most likely belongs to. The SVM model adopts a one-to-one multi-classification strategy. The training data is annotated voice data covering common meeting domains such as law, medicine, IT, and finance. The sample size of a single category is no less than 100,000 hours, which can ensure the reliability of the classification results. The model outputs a probability value for each preset category, and selects the category corresponding to the maximum value as the meeting domain label. Compared with the traditional domain determination method based on keyword matching, this method makes judgments based on the learned classification boundary, which is more adaptable and more accurate, and can deal with scenarios where professional terms are misidentified.

[0032] Based on the conference domain tags, the system retrieves the corresponding professional dictionary from the preset dictionary database. This step is the domain dictionary matching and retrieval process, which aims to match a specific set of professional vocabulary for the current conference, providing a benchmark for subsequent text verification and error correction. Each domain tag in the preset dictionary database corresponds to an independent professional dictionary database, which includes professional terms, industry abbreviations, and commonly used expressions in both Chinese and English. Accurate indexing through domain tags enables rapid matching and retrieval. After matching, the professional dictionary database is loaded into the recognition system for subsequent verification and misidentification correction of the initial transcribed text. This ensures that professional terms are not misidentified as similar-sounding general terms from a vocabulary benchmark perspective, thereby improving overall transcription accuracy.

[0033] In this embodiment, step S3, which involves obtaining environmental noise features and meeting scene context features based on the environmental speech information, obtaining adaptive noise reduction parameters based on the environmental noise features and the meeting domain label, and obtaining initial transcribed text based on the adaptive noise reduction parameters and the meeting scene context features, includes: S31. Detect silence segments in the environmental speech information, extract the audio signals of non-silent segments as pure environmental noise samples, and perform power spectral density analysis on the pure environmental noise samples to obtain environmental noise characteristics. S32. Perform acoustic event detection on the speech activity segments in the environmental speech information to identify specific scene events and their time intervals. Specific scene events include events where multiple people speak at the same time, applause events, or projector noise events. The identified event types and their time intervals are used as contextual features of the meeting scene. S33. Calculate the signal-to-noise ratio of each frequency point based on the environmental noise characteristics, and combine it with the preset domain sensitive frequency band weights in the conference domain label to generate a frequency domain adaptive filtering coefficient matrix as an adaptive noise reduction parameter. S34. The user's voice information is frequency-domain filtered using the adaptive noise reduction parameters to obtain a noise-reduced voice signal. Then, based on the specific scene events detected in the context features of the meeting scene, the noise-reduced voice signal is adjusted by event-triggered gain adjustment. Finally, the processed voice signal is input into a general speech recognition model to decode and obtain the initial transcribed text.

[0034] As described in steps S31-S34 above, this invention performs silent segment detection on environmental speech information, extracts non-silent audio signals as pure environmental noise samples, and performs power spectral density analysis on the pure environmental noise samples to obtain environmental noise characteristics. This step is the foundation for accurate environmental noise modeling. Its physical significance lies in separating pure noise signals without effective human voices from mixed audio, obtaining the energy distribution law of noise across the entire frequency band through frequency domain analysis, and providing an objective basis for subsequent adaptive noise reduction. The silent segment detection adopts a frame-level judgment mechanism based on WebRTCVAD, with a frame length set to 10 milliseconds, which can accurately distinguish between human voice segments and silent noise segments. The extracted pure environmental noise samples cover typical conference noises such as projector fan noise, air conditioner background noise, and keyboard typing noise. The power spectral density analysis adopts the Welch algorithm, which suppresses spectral fluctuations through piecewise averaging and window functions, and outputs the noise power value of each frequency point. These values ​​are arranged in frequency order to form environmental noise characteristics. This feature fully characterizes the noise physical characteristics of the current conference environment and avoids the adaptation bias caused by the traditional use of fixed noise models.

[0035] Acoustic event detection is performed on speech activity segments in the environmental speech information to identify specific scene events and their time intervals. The identified event types and time intervals are used as the context features of the meeting scene. This step is a key link in the dynamic perception of the meeting scene. Its physical significance lies in locating sudden interference or multiple voice superposition events that occur during the meeting, clarifying the time and type of interference, and providing a trigger basis for subsequent dynamic gain adjustment. The acoustic event detection adopts a hybrid architecture model of convolutional neural network combined with gated recurrent unit. This model extracts local spectral features through multi-layer convolution and models temporal dependencies through gated recurrent units. It can accurately identify events such as multiple people speaking at the same time, applause, and projector noise. The model output includes the event type and the corresponding start and end timestamps. This information is combined in chronological order to form the context features of the meeting scene. This feature can reflect the dynamic changes of the meeting scene in real time, so that subsequent speech processing can make differentiated adjustments for specific time periods, rather than using a uniform processing method for the entire speech.

[0036] The signal-to-noise ratio (SNR) of each frequency point is calculated based on the environmental noise characteristics. Combined with the preset domain-sensitive frequency band weights in the conference domain labels, a frequency-domain adaptive filtering coefficient matrix is ​​generated as an adaptive noise reduction parameter. This step is the core of the dual adaptive noise reduction parameter generation process. Its physical meaning lies in assigning optimal filtering weights to each frequency point based on noise strength and domain speech characteristics, thereby suppressing noise while preserving the effective speech segments of the corresponding domain to the greatest extent. The SNR calculation uses the ratio of the user speech signal power to the environmental noise power at the corresponding frequency point. The higher the value, the clearer the speech at that frequency point. The domain-sensitive frequency band weights corresponding to the conference domain labels are determined by the domain speech spectrum statistics. For example, the IT domain contains a large number of English abbreviations, so higher weights are set for the mid-to-high frequency bands. The legal domain mainly uses formal Chinese pronunciation, so higher weights are set for the mid-to-low frequency bands. The SNR and domain-sensitive frequency band weights are weighted and calculated for each frequency point to generate filtering coefficients that correspond one-to-one with the frequency points. All coefficients are arranged by frequency to form a frequency-domain adaptive filtering coefficient matrix. This matrix serves as an adaptive noise reduction parameter, achieving synchronous adaptation between noise levels and domain characteristics, which is different from the traditional fixed filtering coefficient processing method.

[0037] The user's speech information is frequency-domain filtered using adaptive noise reduction parameters to obtain a denoised speech signal. Then, based on specific scene events detected in the context features of the meeting scenario, event-triggered gain adjustment is performed on the denoised speech signal. Finally, the processed speech signal is input into a general speech recognition model for decoding to obtain the initial transcribed text. This step is the execution stage of speech enhancement and initial recognition. Its physical meaning is to apply the generated adaptive noise reduction parameters to the speech signal, optimize the signal amplitude in combination with scene events, and finally complete the basic transcribed text generation. The frequency domain filtering adopts a linear filtering method, multiplying the filter coefficient by frequency point to suppress noise while retaining the energy of the domain-sensitive frequency band. The event-triggered gain adjustment is executed according to preset rules. When multiple people speak at the same time, the overall signal gain is increased to avoid the speech being drowned out, while the gain is reduced during periods of impact noise such as applause to prevent amplitude saturation. The processed speech signal has the characteristics of low noise, low distortion, and complete preservation of domain features. After being input into the general speech recognition model, the initial transcribed text is obtained through beam search decoding. This text can still maintain a high basic accuracy in complex environments, providing reliable input for subsequent dictionary verification and semantic error correction.

[0038] In this embodiment, step S4, which involves matching and verifying the initial transcribed text against the professional dictionary database to obtain a sequence of misidentified segments, includes: S41. The initial transcribed text is segmented to obtain an initial word sequence arranged in chronological order, and each word in the initial word sequence is matched precisely with the entries in the professional dictionary database. S42. For words that fail to match, the posterior probability score generated by the general speech recognition model during the decoding process is further called, and the posterior probability score is used as the confidence level of the words that fail to match. S43. Determine whether the confidence level is lower than a preset confidence threshold; If the confidence level is lower than the preset confidence level threshold, the corresponding failed matching words will be marked as misidentified candidate words. S44. Combine the misidentified candidate words in adjacent time windows according to their original order in the initial transcribed text to form one or more consecutive misidentified segments. Each misidentified segment contains at least one misidentified candidate word. Arrange all misidentified segments in chronological order to obtain a sequence of misidentified segments.

[0039] As described in steps S41-S44 above, this invention performs word segmentation on the initial transcribed text to obtain an initial word sequence arranged in chronological order. Each word in the initial word sequence is then matched precisely with entries in a professional dictionary. This step is the basic screening step for misidentification and localization. Its physical meaning is to split the continuous transcribed text into independent word units aligned with the speech time, and to complete the first round of explicit error screening based on a domain-specific dictionary, ensuring that all words that do not conform to professional expressions can be initially detected. The word segmentation process uses the maximum positive matching algorithm consistent with the feature extraction mentioned above to ensure the consistency of word boundaries and timestamps. The initial word sequence is strictly arranged according to the order of speech pronunciation. Precise string matching uses a complete equivalence comparison method, without fuzzy matching or synonym replacement, ensuring that only standard terms, English abbreviations, and mixed Chinese-English words included in the professional dictionary can pass the matching. For example, in a financial conference scenario, market makers can pass the matching, while words with similar pronunciations, such as "doing business," will be directly judged as failing the matching. This step can quickly filter out words that obviously do not conform to the current conference domain norms, narrowing the scope of subsequent error judgments.

[0040] For words that fail to match, the posterior probability score generated by the general speech recognition model during the decoding process is further invoked, and this posterior probability score is used as the confidence level of the words that failed to match. This step is the error confidence quantification step, which physically means introducing the model's own deterministic score of the recognition result, and conducting numerical confidence assessment on the initially screened abnormal words. This avoids misjudging low-frequency professional words or matching failures caused by temporary fluctuations in the model as real errors. The posterior probability score is output synchronously by the general speech recognition model during the bundle search decoding stage. This score reflects the model's internal confidence level for the current word decoding result, and the value ranges from 0 to 1. The lower the score, the weaker the model's support for the recognition result, and the higher the possibility of misidentification. Using this score directly as the confidence basis can make full use of the decoding information inside the model to form a more reliable error judgment standard than external text rules. This step upgrades the binary judgment of string matching to probabilistic confidence judgment, improving the stability of the overall localization mechanism.

[0041] The system determines whether the confidence level is lower than a preset confidence threshold. If it is, the corresponding failed matching words are marked as misidentified candidate words. This step is the precise screening of misidentified candidate words. Its physical meaning is that by setting a quantitative threshold, words with insufficient confidence and failed dictionary matching are identified as high-probability errors, eliminating accidental matching failures with high confidence and ensuring the accuracy of candidate words. When the confidence level of a failed matching word is lower than this threshold, it means that the model itself is uncertain about the result, and the content does not conform to professional expression standards. It can be marked as a misidentified candidate word with high confidence. Threshold screening can effectively reduce false detections, allowing subsequent correction operations to focus on the segments that truly need adjustment.

[0042] The misidentified candidate words within adjacent time windows are combined according to their original order in the initial transcribed text to form one or more consecutive misidentified segments. Each misidentified segment contains at least one misidentified candidate word. All misidentified segments are then arranged chronologically to obtain a sequence of misidentified segments. This step is the error fragment aggregation and serialization output stage. Its physical significance lies in combining scattered, isolated erroneous words into semantically complete correction units according to temporal continuity, avoiding contextual semantic breaks caused by fragmented correction. At the same time, it forms a batch-processable correction target sequence in chronological order. This step uses adjacent time windows with a length of 5 words. The length of this window is determined based on the common continuous distribution statistics of professional terms and collocations in conference speech, which can cover most professional term combination scenarios without introducing non-erroneous words due to an excessively large window. The combined misidentified segments maintain the original order in the initial transcribed text and have complete timestamps and location information. All segments are arranged chronologically to form a sequence of misidentified segments. This sequence can be directly input into the subsequent error correction module to achieve accurate and efficient batch correction and improve the overall processing efficiency of the transcription system.

[0043] In an embodiment, step S5, which involves obtaining contextual semantics based on the initial text features and obtaining candidate word sequences based on the contextual semantics and the acoustic features, includes: S51. Extract the fixed-length text to the left of each misidentified segment from the initial text features as the preceding text, extract the fixed-length text to the right of each misidentified segment as the following text, and input the preceding text and the following text into the pre-trained language model to obtain the preceding text semantic vector and the following text semantic vector, respectively. S52. The preceding semantic vector and the following semantic vector are concatenated to obtain a joint context semantic vector, and the joint context semantic vector is mapped to a context attention weight distribution through a multilayer perceptron network. S53. Extract acoustic feature segments that are time-aligned with the misidentified segment from the acoustic features, and perform phoneme posterior probability decoding on the acoustic feature segments to obtain the acoustic candidate phoneme sequence of the misidentified segment. S54. Using the context attention weight distribution as a weighting coefficient, the acoustic candidate phoneme sequence is reconstructed by weighting. Then, a list consisting of multiple words ordered by probability is generated by a phoneme-to-word decoder based on beam search, which serves as the candidate word sequence.

[0044] As described in steps S51-S54 above, this invention extracts a fixed-length text to the left of each misidentified segment as the preceding text and a fixed-length text to the right of each misidentified segment as the following text from the initial text features. The preceding and following texts are then input into a pre-trained language model to obtain the preceding semantic vector and the following semantic vector, respectively. This step is the foundational step for extracting contextual semantic features. Its physical significance lies in obtaining the semantic context before and after the misidentified segment, transforming natural language text into a computable distributed semantic vector, and providing data support for subsequent semantic constraints. The length of the extracted context is set to 10 words. This length is determined based on the statistical analysis of the effective semantic coverage of official documents and meeting statements, which can fully cover the core semantics of the sentence containing the misidentified segment. The pre-trained language model adopts the BERT-base Chinese model, which contains a 12-layer Transformer structure and can effectively capture the deep semantics and syntactic relationships of the text. The preceding and following texts are encoded by the model and output fixed-dimensional semantic vectors. The two vectors respectively represent the semantic direction and topic constraints before and after the misidentified segment, avoiding semantic deviation in subsequent corrections.

[0045] The preceding and following semantic vectors are concatenated to obtain a joint context semantic vector. This joint context semantic vector is then mapped to a context attention weight distribution using a multilayer perceptron network. This step is the core of the transformation from context semantics to attention constraints. Physically, it integrates and transforms context semantic information into attention weights that can be used for acoustic feature weighting, enabling subsequent acoustic feature reconstruction to prioritize pronunciation information that conforms to the context semantics. The concatenation method involves direct connection along the feature dimension, maintaining the integrity of the preceding and following semantics. The multilayer perceptron network uses a two-layer fully connected structure with a ReLU activation function in between, which can map the high-dimensional semantic vector to a weight distribution that matches the length of the phoneme sequence. This attention weight distribution represents the degree of attention the context pays to different pronunciation positions, making the use of acoustic features more targeted rather than treating all acoustic information equally, effectively improving the matching degree between candidate words and context.

[0046] The acoustic feature segments that are time-aligned with the misidentified segments are extracted from the acoustic features, and phoneme posterior probability decoding is performed on the acoustic feature segments to obtain the acoustic candidate phoneme sequence of the misidentified segments. This step is the precise extraction of pronunciation features and phoneme-level decoding. Its physical significance lies in locating the original speech acoustic information corresponding to the misidentified segments and converting the acoustic features into phoneme-level pronunciation sequences, providing the underlying pronunciation basis for candidate word generation. Time alignment is completed based on the speech timestamps corresponding to the words in the initial transcribed text, ensuring that the extracted acoustic feature segments are completely consistent with the pronunciation time of the misidentified segments. Phoneme posterior probability decoding adopts a phoneme alignment algorithm based on a hidden Markov model, which converts continuous acoustic features into discrete phoneme probability sequences. This sequence completely preserves the pronunciation details of the original speech, including initials, finals, stress, and linking features, and is not affected by previous recognition errors. It can truly reflect the user's actual pronunciation content, such as restoring the acoustic features corresponding to the misidentified transmission architecture to the phoneme sequence corresponding to the correct pronunciation.

[0047] The contextual attention weight distribution is used as a weighting coefficient to reconstruct the acoustic candidate phoneme sequence. Specifically, this involves combining the positional information of each phoneme in the phoneme sequence to map a fixed-dimensional joint contextual semantic vector into a weight vector with the same length as the current acoustic candidate phoneme sequence. Each weight coefficient in this weight vector corresponds to a phoneme position in the acoustic candidate phoneme sequence, thus achieving direct weighting of the phoneme level by contextual semantic information.

[0048] First, obtain the length information of the acoustic candidate phoneme sequence: denote the length of the acoustic candidate phoneme sequence as T, where T is a natural number representing the number of phonemes contained in the sequence. At the same time, generate a position encoding vector pos_i for each phoneme position i (i=1,2,…,T). The position encoding vector adopts the sine-cosine position encoding formula or a learnable position embedding vector, and its dimension is set to D_pos (e.g., 128 dimensions).

[0049] Secondly, a context-phoneme weight mapping network is constructed: a two-layer fully connected network is set as the weight mapping module, which shares parameters across all phoneme positions. For the i-th phoneme position, the joint context semantic vector v_ctx (with dimension D, e.g., 768-dimensional) is concatenated with the position encoding vector pos_i to obtain the input vector input_i=[v_ctx,pos_i], where the dimension is D+D_pos. input_i is then passed through the first fully connected layer, the ReLU activation function, and the second fully connected layer, and finally output as a scalar weight w_i between 0 and 1 via the Sigmoid function, where the calculation formula is: h_i=ReLU(W1*input_i+b1), w_i=Sigmoid(W2*h_i+b2); Where W1, b1, W2, and b2 are trainable network parameters, W1 has a dimension of H×(D+D_pos), b1 is an H-dimensional vector, W2 has a dimension of 1×H, b2 is a scalar, and H is the hidden layer dimension (e.g., 128).

[0050] By traversing all phoneme positions i=1~T, we obtain the weight vector w=[w_1,w_2,...,w_T], which is the context attention weight distribution with the same length as the acoustic candidate phoneme sequence. The acoustic candidate phoneme sequence is reconstructed by weighting: the acoustic candidate phoneme sequence is P=[p_1,p_2,...,p_T], where each p_i is a posterior probability vector of length K (total number of phoneme categories) and the sum of its components is 1.

[0051] Each weight coefficient w_i in the context attention weight distribution w is applied to the corresponding p_i to perform weighted reconstruction, resulting in the reconstructed phoneme sequence P'=[p'_1,p'_2,...,p'_T].

[0052] The specific weighted calculation method is as follows: For each i, multiply each component of p_i by w_i, i.e., p'_i = w_i·p_i, and then renormalize p'_i so that the sum of all components is 1, i.e., p'_i ← p'_i / sum(p'_i). This operation is equivalent to using contextual semantic information to enhance or suppress the original acoustic probability at each phoneme position, so that phonemes that conform to the contextual semantic expectation get a higher effective probability, while those that do not are weakened. This achieves weighted reconstruction of acoustic candidate phoneme sequences using phonemes as the basic unit and contextual attention weights, completely resolving the issues of inconsistent units and length mismatches between word / character-level weights and phoneme-level sequences. Furthermore, the reconstructed phoneme sequence P' can be input into the beam search's phoneme-to-word decoder to generate a list of multiple words ordered by probability as a candidate word sequence. This step involves semantic and acoustic fusion decoding and candidate word generation. Its physical significance lies in using contextual semantic constraints to guide the reconstruction of the acoustic phoneme sequence, and then using beam search to transform the phoneme sequence into a domain-compliant sequence. The standardized vocabulary list and weighted reconstruction multiply the attention weights by the phoneme probabilities at corresponding positions point by point, strengthening the weights of phonemes that conform to the contextual semantics and weakening the weights of irrelevant phonemes. The beam width of the beam search is set to 10, which can balance decoding accuracy and computing speed on smart terminal devices. The decoder has a built-in professional dictionary library that has been matched in the previous text, ensuring that the output vocabulary prioritizes the professional expressions of the current conference field. The final output candidate word sequence is sorted from high to low according to the comprehensive probability, which meets the triple requirements of pronunciation matching, semantic fluency and professional standardization, and can be directly used for subsequent replacement of misidentified segments.

[0053] For example, in an IT architecture design meeting, participants were discussing deep learning models. The original audio content was: "We plan to adopt the Transformer architecture to improve the model performance." During preprocessing, a sequence of misidentified segments was generated in the initial transcribed text. The first misidentified segment was "transformer architecture," which corresponds to the pronunciation range of "Transformer architecture" in the original audio (timestamp: 1.5 seconds - 2.3 seconds). The first step is to extract the context semantic vector: extract the preceding text ("We plan to adopt") of a fixed length (10 words) to the left of the misidentified segment "transmission architecture" from the initial text features, and the following text ("To improve model performance") of a fixed length (10 words) to the right. Then, input the preceding and following text strings into a pre-trained BERT-base Chinese language model (12-layer Transformer, output dimension 768). Model output: preceding semantic vector and following semantic vector; Among them, the two vectors encode the semantic direction and topic constraints before and after the misidentified segment, respectively, providing a semantic basis for the subsequent generation of attention weights; The second step is to generate the context attention weight distribution, and directly concatenate the preceding semantic vector and the following semantic vector along the feature dimension to obtain the joint context semantic vector. The acoustic candidate phoneme sequence has a length of T=12 (corresponding to the 12 phonemes in the "Transformer architecture"). For each phoneme position i=1,2,…,12, a position encoding vector is generated, and a two-layer fully connected network (weight mapping module, sharing parameters across all phoneme positions) is constructed: First layer: Input dimension 1536+128=1664, output dimension H=128, activation function ReLU; Second layer: Input dimension 128, output dimension 1, activation function Sigmoid; For the i-th phoneme position, the joint context semantic vector and the position encoding vector are concatenated and then input into the network; h_i=ReLU(w1*input_i+b1), w_i=Sigmoid(w2*h_i+b2); Secondly, iterating through i=1 to 12, we obtain the context attention weight distribution: W=[w1, w2, ..., w 12 For example, w=0.95,0.92,0.88,0.85,0.80,0.78,0.75,0.70,0.68,0.65,0.60,0.55; The weight distribution reflects the degree of attention the context semantics pays to the phoneme position: since the stress of "Transformer" is in the first half, the first few phonemes receive higher weights, while the weights in the second half decrease slightly, but still retain sufficient effective information. The third step involves extracting time-aligned acoustic feature segments from the original acoustic features based on the time interval (1.5 seconds to 2.3 seconds) of the misidentified segment's "transmitter architecture." A phoneme alignment algorithm based on a Hidden Markov Model is then used to perform posterior probability decoding on this segment, yielding an acoustic candidate phoneme sequence. P = [P1, P2, ..., P] 12 ]; Here, each P is a posterior probability vector for a phoneme (total number of phoneme categories K=128), with the sum of its components being 1. However, the actual decoded initial phoneme posterior probabilities are biased towards the pronunciation of the "Transformer architecture." For example, in the first phoneme position p1: / t / has a probability of 0.60, / d / has a probability of 0.30, and the rest have a probability of 0.10; in the fifth phoneme position, the probability of / f / , which should be / f / , is blurred by noise and becomes / s / , which dominates. This sequence fully preserves the pronunciation details of the original speech, but due to noise and the initial decoding bias of the model, it differs from the correct pronunciation of the "Transformer architecture." The fourth step involves weighted reconstruction and generation of a candidate word sequence. The context attention weight distribution *w* is used as the weighting coefficient to perform positional weighting on the acoustic candidate phoneme sequence: *p'_i* = *w_i* *p_i*. Then, *p'_i* is re-normalized so that the sum of all components is 1, i.e., *p'_i* ← *p'_i* / *sum(p'_i)*. The reconstructed phoneme sequence is *p'* = [*p'1*, *p'2*, ..., *p'*]. 12 This more accurately reflects the pronunciation characteristics of the "Transformer architecture"; Finally, beam search decoding: The reconstructed phoneme sequence p' is input into the beam search decoder (beam width set to 10). This decoder has a built-in professional dictionary library in the IT field (containing entries such as "Transformer architecture", "API", and "convolutional neural network"), and performs phoneme-to-word conversion; The decoder combines phoneme matching probabilities with dictionary constraints and outputs a sequence of candidate words in descending order of combined probability: 1. "Transformer architecture" (combined probability 0.92), 2. "Transformer architecture" (0.08), 3. "Transmission architecture" (0.03), 4. "Transmitter architecture" (0.01), 5. "Transformer structure" (0.01). Among them, "Transformer architecture" ranked first with a significantly higher probability than other candidate words. This candidate word not only matches the pronunciation features of the original speech, but also conforms to the professional expression in the IT field, and is semantically coherent with the context "we plan to use... to improve the model performance". This candidate word sequence will be output to steps S55-S58 to replace the misidentified segment "transformer architecture", completing the correction of the final transcribed text.

[0054] In an embodiment, step S5, which involves replacing the misidentified segment sequence in the initial transcribed text with the candidate word sequence to obtain the final conference transcribed text, includes: S55. Select the top N candidate words with the highest probability from the candidate word sequence to form a candidate word list, and locate the position of the current misidentified segment in the misidentified segment sequence in the initial transcribed text; S56. Calculate the edit distance between each candidate word in the candidate word list and the current misidentified segment, and remove candidate words whose edit distance exceeds a preset distance threshold to obtain a filtered candidate word list; S57. The candidate word with the highest probability in the filtered candidate word list is used as the replacement word. The replacement word is used to replace the corresponding position of the current misidentified segment in the initial transcribed text to generate the locally corrected transcribed text. S58. Traverse all misidentified segments in the misidentified segment sequence and repeat the step of selecting the top N candidate words with the highest probability from the candidate word sequence to generate the locally corrected transcribed text, complete the replacement operation for all misidentified segments, and output the complete transcribed text after replacement as the final conference transcribed text.

[0055] As described in steps S55-S58 above, this invention selects the top N candidate words with the highest probability from the candidate word sequence to form a candidate word list, and locates the position of the current misidentified segment in the misidentified segment sequence in the initial transcribed text. This step is a preparatory step for correction execution. Its physical significance is to narrow down the candidate word range, focus on high-confidence correction options, and accurately lock the position of the segment to be corrected in the text to ensure that the replacement operation does not destroy the original text structure. In this step, N is set to 5. This value is determined based on the statistical analysis of the effective candidate coverage range of conference speech terminology correction, which can reduce the subsequent computational load while ensuring the sufficiency of candidates. The candidate word sequence comes from the semantic and acoustic fusion candidate results generated in S51 to S54 above, and is sorted from high to low according to the model output probability. The position is located based on the timestamp and vocabulary index of the initial transcribed text, which can accurately match the start and end positions of the current misidentified segment in the complete text, providing a positional basis for subsequent accurate replacement.

[0056] The edit distance between each candidate word in the candidate word list and the current misidentified segment is calculated, and candidate words whose edit distance exceeds a preset distance threshold are removed to obtain a filtered candidate word list. This step is a character similarity compliance verification step. Its physical meaning is to measure the similarity between the pronunciation and characters of the candidate words and the misidentified segment through the string edit distance, and to remove candidate words that are too different or do not meet the homophone and near-homophone error characteristics, so as to avoid unreasonable replacement. The edit distance is used to represent the minimum number of single-character operations required to convert two strings from one to another, which can objectively reflect the similarity between pronunciation and spelling. Secondly, the preset distance threshold can effectively retain reasonable candidates and remove abnormal candidates, so that the remaining words all meet the error characteristics of real speech transcription, thereby improving the rationality of replacement.

[0057] The candidate word with the highest probability in the filtered candidate word list is used as the replacement word. This replacement word replaces the corresponding position of the currently misidentified segment in the initial transcribed text, generating a locally corrected transcribed text. This step is the single-segment precise replacement execution stage. Its physical significance lies in selecting the word with the highest comprehensive confidence and the best semantic and acoustic matching to complete the local correction, eliminating errors without affecting other text content. The filtered candidate word list has simultaneously met the two conditions of highest probability and edit distance compliance. Therefore, selecting the top-ranked candidate word as the replacement word has the highest correction reliability. The replacement process is strictly executed according to the S55 positioning position, only covering the interval where the misidentified segment is located, without changing the correct content of the context. The generated locally corrected text maintains the original text structure and semantic coherence, while correcting erroneous terms. For example, the misidentified transmission architecture is replaced with the Transformer architecture that conforms to professional expression, restoring the correct semantics of the local content.

[0058] The process iterates through all misidentified segments in the misidentified segment sequence, repeating the steps from candidate word selection to local text correction generation. This completes the replacement of all misidentified segments, outputting the final transcribed text as the final meeting transcript. This step involves full error correction and final output. Its physical significance lies in performing standardized corrections on all located error segments one by one according to chronological order, ensuring the entire text is complete and error-free. The final result is a professionally usable meeting transcript. The traversal process follows the chronological order of the misidentified segment sequence, ensuring the correction order matches the speech order. Each segment repeats the screening, verification, and replacement process from S55 to S57, ensuring all errors are corrected according to a unified, high-reliability standard. After all replacements are completed, misidentifications of technical terms and homophones in the text are corrected, resulting in a fluent, formatted text that conforms to domain expression habits. It can be directly used for subsequent applications such as meeting minutes and task extraction, without requiring extensive manual proofreading.

[0059] This application also provides a smart terminal conference voice transcription system, including: A voice information acquisition module is used to acquire voice information from smart terminal conferences, wherein the voice information includes user voice information and environmental voice information; The domain recognition module is used to obtain acoustic features and initial text features based on the user's voice information, obtain conference domain labels based on the acoustic features and initial text features, and obtain the corresponding professional dictionary based on the conference domain labels; The initial transcription module is used to obtain environmental noise features and meeting scene context features based on the environmental speech information, obtain adaptive noise reduction parameters based on the environmental noise features and the meeting domain label, and obtain the initial transcribed text based on the adaptive noise reduction parameters and the meeting scene context features. The matching and verification module is used to match and verify the initial transcribed text with the professional dictionary database to obtain a sequence of misidentified segments; The error correction module is used to obtain contextual semantics based on the initial text features, and obtain candidate word sequences based on the contextual semantics and the acoustic features. The candidate word sequences are then used to replace the misidentified segment sequences in the initial transcribed text to obtain the final conference transcribed text.

[0060] In this embodiment, the domain identification module includes: The feature extraction unit is used to perform frame segmentation and windowing processing on the user's voice information to obtain multiple voice frames, and to obtain multiple time-domain-frequency-domain energy values ​​based on the multiple voice frames. The feature vector acquisition unit is used to acquire a time-frequency energy vector based on multiple time-frequency energy values ​​and a preset time axis, and to use the time-frequency energy vector as an acoustic feature vector. The recognition and decoding unit is used to input the user's voice information into a pre-trained end-to-end speech recognition model, perform greedy decoding, and generate an initial recognition text string; The text feature extraction unit is used to perform maximum positive matching word segmentation on the initial identified text string to obtain a word segmentation sequence, and to perform part-of-speech tagging on each word in the word segmentation sequence based on a hidden Markov model to obtain a part-of-speech tag sequence. The word segmentation sequence and the part-of-speech tag sequence are used together as the initial text feature vector. The feature fusion unit is used to concatenate the acoustic feature vector and the initial text feature vector to obtain a fused feature vector; The conference domain classification unit is used to input the fused feature vector into a pre-trained support vector machine classification model, calculate the probability that the fused feature vector belongs to each of the preset multiple conference domain categories, and output the category corresponding to the highest probability as the conference domain label. The professional dictionary retrieval module is used to retrieve the corresponding professional dictionary from a preset dictionary set based on the conference field tags. This application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described method.

[0061] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method.

[0062] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in this application and in the embodiments can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0063] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, apparatus, article, or method. Unless otherwise specified, an element defined by the phrase "comprising..." does not exclude the presence of other identical elements in the process, apparatus, article, or method that includes that element.

[0064] The above description is merely a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. A method for converting conference voice into text by an intelligent terminal, characterized in that, include: Acquire voice information from a smart terminal conference, wherein the voice information includes user voice information and environmental voice information; Acoustic features and initial text features are obtained based on the user's voice information, and conference domain tags are obtained based on the acoustic features and initial text features, and corresponding professional dictionary databases are obtained based on the conference domain tags; Environmental noise features and meeting scenario context features are obtained based on the environmental speech information, and adaptive noise reduction parameters are obtained based on the environmental noise features and the meeting domain labels. Initial transcribed text is obtained based on the adaptive noise reduction parameters and the meeting scenario context features. The initial transcribed text is matched and verified against the professional dictionary database to obtain a sequence of misidentified fragments; Contextual semantics are obtained based on the initial text features, and candidate word sequences are obtained based on the contextual semantics and the acoustic features. The candidate word sequences are then used to replace the misidentified fragment sequences in the initial transcribed text to obtain the final conference transcribed text. 2.The intelligent terminal conference voice transcription method of claim 1, wherein, The steps of obtaining acoustic features and initial text features based on the user's voice information, obtaining conference domain tags based on the acoustic features and initial text features, and obtaining corresponding professional dictionary databases based on the conference domain tags include: The user voice information is segmented and windowed to obtain multiple voice frames, and multiple time-domain-frequency-domain energy values ​​are obtained based on the multiple voice frames. A time-frequency energy vector is obtained based on multiple time-frequency energy values ​​and a preset time axis, and the time-frequency energy vector is used as an acoustic feature vector. The user's voice information is input into a pre-trained end-to-end speech recognition model, and greedy decoding is performed to generate an initial recognition text string. The initial identified text string is segmented by maximum positive matching to obtain a segmented word sequence. Each word in the segmented word sequence is labeled with part-of-speech tags based on a hidden Markov model to obtain a part-of-speech tag sequence. The segmented word sequence and the part-of-speech tag sequence are used together as the initial text feature vector. The acoustic feature vector and the initial text feature vector are concatenated to obtain a fused feature vector; The fused feature vector is input into a pre-trained support vector machine classification model, and the probability of the fused feature vector belonging to each of the preset multiple conference domain categories is calculated. The category corresponding to the highest probability is output as the conference domain label. The corresponding professional dictionary is retrieved from a preset dictionary database based on the conference domain label. 3.The intelligent terminal conference voice transcription method of claim 1, wherein, The steps of obtaining environmental noise features and meeting scene context features based on the environmental speech information, obtaining adaptive noise reduction parameters based on the environmental noise features and the meeting domain labels, and obtaining initial transcribed text based on the adaptive noise reduction parameters and the meeting scene context features include: The environmental speech information is subjected to silence segment detection, and the audio signal of the non-silent segment is extracted as a pure environmental noise sample. The pure environmental noise sample is then subjected to power spectral density analysis to obtain the environmental noise characteristics. Acoustic event detection is performed on the speech activity segments in the environmental speech information to identify specific scene events and their time intervals. Specific scene events include events where multiple people speak at the same time, applause events, or projector noise events. The identified event types and their time intervals are used as contextual features of the meeting scene. The signal-to-noise ratio of each frequency point is calculated based on the environmental noise characteristics, and the frequency domain adaptive filtering coefficient matrix is ​​generated by combining the preset domain sensitive frequency band weights in the conference domain label as adaptive noise reduction parameters. The user's voice information is frequency-domain filtered using the adaptive noise reduction parameters to obtain a noise-reduced voice signal. Then, based on specific scene events detected in the context features of the meeting scene, the noise-reduced voice signal is subjected to event-triggered gain adjustment. Finally, the processed voice signal is input into a general speech recognition model to decode and obtain the initial transcribed text.

4. The intelligent terminal conference voice transcription method of claim 1, wherein, The step of matching and verifying the initial transcribed text with the professional dictionary database to obtain the misidentified segment sequence includes: The initial transcribed text is segmented to obtain an initial word sequence arranged in chronological order, and each word in the initial word sequence is matched precisely with the entries in the professional dictionary database. For words that fail to match, the posterior probability score generated by the general speech recognition model during the decoding process is further invoked, and this posterior probability score is used as the confidence level of the words that fail to match. Determine whether the confidence level is lower than a preset confidence threshold; If the confidence level is lower than the preset confidence level threshold, the corresponding failed matching words will be marked as misidentified candidate words. The misidentified candidate words within adjacent time windows are combined according to their original order in the initial transcribed text to form one or more consecutive misidentified segments. Each misidentified segment contains at least one misidentified candidate word. All misidentified segments are then arranged in chronological order to obtain a sequence of misidentified segments.

5. The intelligent terminal conference voice transcription method of claim 1, wherein, The steps of obtaining contextual semantics based on the initial text features and obtaining candidate word sequences based on the contextual semantics and the acoustic features include: Extract the fixed-length text to the left of each misidentified segment as the preceding text from the initial text features, and extract the fixed-length text to the right of each misidentified segment as the following text. Input the preceding text and the following text into a pre-trained language model to obtain the preceding text semantic vector and the following text semantic vector, respectively. The preceding semantic vector and the following semantic vector are concatenated to obtain a joint context semantic vector, and the joint context semantic vector is mapped to a context attention weight distribution through a multilayer perceptron network. From the acoustic features, an acoustic feature segment that is time-aligned with the misidentified segment is extracted, and the acoustic feature segment is decoded using phoneme posterior probability to obtain the acoustic candidate phoneme sequence of the misidentified segment. The contextual attention weight distribution is used as a weighting coefficient to reconstruct the acoustic candidate phoneme sequence. Then, a list consisting of multiple words ordered by probability is generated by a beam search-based phoneme-to-word decoder as a candidate word sequence. 6.The intelligent terminal conference voice transcription method of claim 1, wherein, The step of replacing the misidentified segment sequence in the initial transcribed text with the candidate word sequence to obtain the final conference transcribed text includes: The top N candidate words with the highest probability are selected from the candidate word sequence to form a candidate word list, and the position of the current misidentified segment in the misidentified segment sequence is located in the initial transcribed text. Calculate the edit distance between each candidate word in the candidate word list and the current misidentified segment, and remove candidate words whose edit distance exceeds a preset distance threshold to obtain a filtered candidate word list; The candidate word with the highest probability in the filtered candidate word list is used as the replacement word. The replacement word is used to replace the corresponding position of the currently misidentified segment in the initial transcribed text to generate a locally corrected transcribed text. Traverse all misidentified segments in the misidentified segment sequence and repeat the step of selecting the top N candidate words with the highest probability from the candidate word sequence to generate the locally corrected transcribed text, complete the replacement operation for all misidentified segments, and output the complete transcribed text after replacement as the final conference transcribed text.

7. An intelligent terminal conference voice transcription system, characterized in that, include: A voice information acquisition module is used to acquire voice information from smart terminal conferences, wherein the voice information includes user voice information and environmental voice information; The domain recognition module is used to obtain acoustic features and initial text features based on the user's voice information, obtain conference domain labels based on the acoustic features and initial text features, and obtain the corresponding professional dictionary based on the conference domain labels; The initial transcription module is used to obtain environmental noise features and meeting scene context features based on the environmental speech information, obtain adaptive noise reduction parameters based on the environmental noise features and the meeting domain label, and obtain the initial transcribed text based on the adaptive noise reduction parameters and the meeting scene context features. The matching and verification module is used to match and verify the initial transcribed text with the professional dictionary database to obtain a sequence of misidentified segments; The error correction module is used to obtain contextual semantics based on the initial text features, and obtain candidate word sequences based on the contextual semantics and the acoustic features. The candidate word sequences are then used to replace the misidentified segment sequences in the initial transcribed text to obtain the final conference transcribed text.

8. The intelligent terminal conference voice transcription system of claim 7, wherein, The domain identification module includes: The feature extraction unit is used to perform frame segmentation and windowing processing on the user's voice information to obtain multiple voice frames, and to obtain multiple time-domain-frequency-domain energy values ​​based on the multiple voice frames. The feature vector acquisition unit is used to acquire a time-frequency energy vector based on multiple time-frequency energy values ​​and a preset time axis, and to use the time-frequency energy vector as an acoustic feature vector. The recognition and decoding unit is used to input the user's voice information into a pre-trained end-to-end speech recognition model, perform greedy decoding, and generate an initial recognition text string; The text feature extraction unit is used to perform maximum positive matching word segmentation on the initial identified text string to obtain a word segmentation sequence, and to perform part-of-speech tagging on each word in the word segmentation sequence based on a hidden Markov model to obtain a part-of-speech tag sequence. The word segmentation sequence and the part-of-speech tag sequence are used together as the initial text feature vector. The feature fusion unit is used to concatenate the acoustic feature vector and the initial text feature vector to obtain a fused feature vector; The conference domain classification unit is used to input the fused feature vector into a pre-trained support vector machine classification model, calculate the probability that the fused feature vector belongs to each of the preset multiple conference domain categories, and output the category corresponding to the highest probability as the conference domain label. The professional dictionary database retrieval module is used to retrieve the corresponding professional dictionary database from a preset dictionary database set based on the conference field tags. 9.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-8 when the computer program is executed by the processor. When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.