Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

64 results about "Speech transcription" patented technology

Transcription (linguistics), the representations of speech or signing in written form Orthographic transcription, a transcription method that employs the standard spelling system of each target language. Phonetic transcription, the representation of specific speech sounds or sign components.

Training reasoning method based on voice-text-image multi-mode contrast learning

The invention particularly relates to a training reasoning method based on voice-text-image multi-mode comparative learning, and relates to the technical field of artificial intelligence, and the method comprises the steps: constructing a voice-text-GUI screenshot triple, and processing element masks; generating positive and negative alignment pairs of the instruction clauses and the GUI elements; the multiple encoders process corresponding modes, and the Rewriter converts a voice transcription text into a structured sequence; through global / local double fusion and double contrast learning, coordinate loss optimization alignment is carried out. According to the invention, through global and local double-layer fusion and a comparative learning mechanism, fine-grained alignment of voice, text and GUI elements is realized; element masks are generated by means of SAM, post-processing optimization is carried out, accurate element-level visual features are extracted in combination with RoIAlign, and a foundation is laid for alignment; all modal dimensions are unified through linear projection, coarse-grained semantic association is realized through a global fusion layer, and a local fusion layer fuses voice rhythm information to enhance association through cross-modal attention and fine correspondence of a gating mechanism capture sub-step and elements.
Owner:杭州长望智创科技有限公司

Multi-person overlapped voice real-time voiceprint recognition method and system

The invention relates to the field of voiceprint recognition and voice transfer, and discloses a real-time voiceprint recognition method and system for multi-user overlapped voices. The method comprises the following steps: acquiring a multi-source audio stream, and carrying out standardization and framing processing to generate a streaming audio frame sequence; based on the sequence, through circular buffering, noise reduction, endpoint detection and overlap detection model processing, obtaining an overlap interval label; task assembly, speaker separation, track numbering, voiceprint feature extraction and identity judgment are carried out, and a track identity binding structure is generated; and finally, voice transcription, fragment splicing and conflict cutting are executed, and the voiceprint template library is updated. According to the method, real-time separation and identity recognition of overlapped voices of multiple persons are realized, and the accuracy and robustness of voiceprint recognition in a complex scene are effectively improved.
Owner:HUNAN ZHENTONG ZHIYONG ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD

Multilingual full-speech processing method and device based on speech recognition, and medium

The invention discloses a multilingual full-speech processing method and device based on speech recognition and a medium, and relates to the technical field of speech recognition, and the method comprises the steps: calculating a language explicit trajectory based on a multilingual speech feature set, carrying out the association with a speech segment through language preference information in a historical session, and constructing a language implicit trajectory, integrating and generating a language weight track; dividing the language weight trajectory into voice trajectory nodes, recording multi-language voice trajectory composite features, connecting the multi-language voice trajectory composite features into a voice trajectory chain, calculating inter-node continuity indexes, and generating a voice trajectory node structure; constructing a track node continuity credibility field according to the voice track node structure, and adjusting multilingual voice recognition decision parameters to generate a multilingual transcription candidate set; and carrying out time sequence splicing and language mark arrangement on the language transcription candidate set to generate a multi-language full-voice transcription result set. According to the method, self-adaptive decoding of structure perception is realized, language switching is optimized, and the transcription precision is improved.
Owner:CHANGCHUN VOCATIONAL INST OF TECH

Power self-service terminal voice intention recognition system based on background noise optimization

The invention relates to the technical field of voice recognition, in particular to an electric power self-service terminal voice intention recognition system based on background noise optimization. The system comprises a voice interaction noise optimization module, a scene voice transcription analysis module, a scene voice intention recognition module and a recognition confidence iterative optimization module, a voice interaction original data set during voice interaction between the power self-service terminal and a user can be collected, and noise optimization is carried out to generate a target voice interaction data set; performing voice transcription analysis based on the target voice interaction data set, and generating a text transcription result; performing keyword semantic matching calculation on the text transcription result and an intention recognition knowledge base to obtain an intention recognition result of the user; and calculating a comprehensive recognition confidence coefficient based on the intention recognition result, if the comprehensive recognition confidence coefficient is lower than a set threshold value, performing secondary noise optimization on the voice interaction original data set and re-executing the judgment step, and if not, outputting the intention recognition result of the user. According to the invention, the accuracy of terminal speech recognition and intention recognition can be improved.
Owner:STATE GRID TIANJIN ELECTRIC POWER COMPANY

IMPROVING SPEECH RECOGNITION TRANSCRIPTIONS

ActiveDE102021122068B4Speech recognitionAudiometry testCommon word
Computer-implemented method (500) for training a model to improve speech recognition, wherein the computer-implemented method comprises: Receiving (502) an utterance by one or more processors, wherein the receiving is performed by a virtual assistant in a specific node of the virtual assistant, wherein frequently occurring terms have been identified for the specific node over a period of time; Transcription (504) of the utterance into text by the one or more processors; Generating (506) a transcription confidence score based on transcription and audiometry by one or more processors; in response to the transcription confidence score being below a threshold, comparing (510) phonemes in the utterance with phonemes in at least one term from a list of frequently occurring terms by one or more processors; Generating a sound similarity score for phonemes in which at least one term from a list of frequently occurring terms is selected by one or more processors; and Replacing (512) the transcription with at least one term from the list of frequently occurring terms if the sound similarity score is above a threshold, by one or more processors.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Speech-to-speech translation

PendingUS20260154515A1Natural language translationSound input/outputSpeech to speech translationSpeech translation
A speech-to-speech translation method comprises transcribing speech spoken in a source language into transcribed text data in the source language using an on-premises speech recognition model. The transcribed text data is translated into translated text data in a target language using a first on-premises machine translation model. The translated text data is reverse translated into retranslated text data in the source language using a second, different on-premises machine translation model. The transcribed text and the retranslated text are displayed on a screen. The method also involves synthesizing, using an on-premises speech synthesis model, translated speech data in the target language based on the translated text data and play back, in response to a user confirmation, translated speech in the target language based on the translated speech data in the target language.
Owner:MABEL AI AB

Natural language processing systems and methods for intent classification of speech transcription

Aspects of the subject disclosure may include, for example, generating a natural language processing model by training an automatic speech recognition (ASR) encoder with manual transcription. The training is performed by correcting and adjusting relevant factors of the ASR encoder based on determined triplet loss, classification loss and Kullback-Leibler divergence loss. In response to an ASR utterance, the trained natural language processing model generates a predicted intent associated with the ASR utterance with improved accuracy. Other embodiments are disclosed.
Owner:JPMORGAN CHASE BANK NA

Ultrasound report generation method, related equipment and computer program product

The invention discloses an ultrasonic report generation method, related equipment and a computer program product, and relates to the technical field of artificial intelligence. According to the method, streaming recognition is performed on the real-time spoken voice of the doctor in the ultrasonic detection process of the doctor, the current voice transcription text fragment is obtained, the large model is called to extract the ultrasonic detection parameters in the current voice transcription text fragment in a fast thinking mode, and the extracted ultrasonic detection parameters are output and previewed in real time, so that the accuracy of ultrasonic detection is improved. Therefore, real-time interaction with doctors is generated, and the result of the detection process is transparent to the doctors. After ultrasonic detection is finished, a large model is called to generate ultrasonic description and diagnosis based on an original complete voice transfer text in a slow thinking mode, ultrasonic detection parameters, the ultrasonic description and diagnosis are integrated to obtain a high-quality ultrasonic report, the clinical requirement of real-time interaction is met, the clinical working efficiency is improved, and the clinical experience is improved. And the quality of the finally obtained ultrasonic report can be ensured.
Owner:ANHUI IFLYHEALTH CO LTD +1

Method and apparatus for inputting voice into webpage text box of intranet website

The application provides a method and device for voice inputting into a webpage text box of an intranet website, to solve the technical problem that reading microphone audio through the webpage of the intranet website and calling voice recognition function are limited. The method comprises the following steps: receiving a voice input instruction of a webpage text box of an intranet website; using a Web Socket protocol, the webpage sends the voice input instruction to a local microphone application program; under the control of the voice input instruction, the local microphone application program collects voice signals; processing the voice signals to generate a character set; and inputting the character set into the webpage text box of the intranet website. In this way, the intranet user can call the voice recognition function and automatically input the voice recognition text into the webpage text box under the condition of intranet communication, so that the voice transcription of the intranet user is more convenient, and the accuracy and completeness of the input text can be effectively improved, and the work efficiency of the intranet user is greatly improved.
Owner:BEIJING THUNISOFT INFORMATION TECH

Speech recognition with selective use of dynamic language models

A computer-implemented method for transcribing an utterance includes receiving, at a computing system, speech data that characterizes an utterance of a user. A first set of candidate transcriptions of the utterance can be generated using a static class-based language model that includes a plurality of classes that are each populated with class-based terms selected independently of the utterance or the user. The computing system can then determine whether the first set of candidate transcriptions includes class-based terms. Based on whether the first set of candidate transcriptions includes class-based terms, the computing system can determine whether to generate a dynamic class-based language model that includes at least one class that is populated with class-based terms selected based on a context associated with at least one of the utterance and the user.
Owner:GOOGLE LLC

Method and system for speech transcription

A system and method of speech transcription may include applying a machine-learning (ML) based encoder module to an audio data element representing a recording of speech, to obtain one or more encoding vectors, representing said recording in an audio encoding space. Embodiments of the invention may include performing an iterative transcription process on the one or more encoding vectors, to generate a token sequence representing a transcription of the recording. In each iteration, an ML-based multilayered decoder may be inferred on (i) the one or more encoding vectors and (ii) a current version of the token sequence, to a candidate token set that includes two or more candidate tokens, each representing a transcription of a respective word in the recording. The two or more candidate tokens may be appended to the current version of the token sequence, thereby updating the token sequence for a subsequent iteration.
Owner:AIOLA LTD

Captioned telephone service system for user with speech disorder

ActiveUS12592993B2Special service for subscribersSpeech recognitionSpeech ProcessorSpeech disturbances
A captioned telephone service (CTS) system provides a transcription service and a speech-to-text and text-to-speech converting service for the deaf or hard-of-hearing user with a speech disorder during a phone call between the user and the peer. The CTS system transcribes the peer's voice into text to be displayed on the user's device. The CTS system transcribes the user's voice into text based on the database storing user's spoken audios and corresponding texts, and converts the text into a clear and articulate speech to be sent to the peer's device instead of the user's voice in order to help the peer better understand what the user said. The CTS system includes a speech-to-text handler and a text-to-speech handler. The speech-to-text handler transcribes user's voice into text using the database, and the text-to-speech handler converts the text into a clear and articulate speech.
Owner:MEZMO CORP

A portable sound amplification and speech transcription system

PendingCN122266383AAdaptive networkSpeech recognitionSound sourcesFeature extraction speech recognition
The application relates to the technical field of audio signal processing and speech recognition, and discloses a portable sound amplification and speech transcription system which comprises a wind noise sensing and separating module, a beam forming processing module, a dual-path processing module, a speech recognition module, a vocabulary correction module, a power consumption management module and an interactive management module.
Owner:LONGQINGWEI (SHANGHAI) INTELLIGENT TECHNOLOGY CO LTD

Speech transcription method, speech transcription apparatus, and electronic device

The application discloses a speech transcription method, a speech transcription device and an electronic device, and belongs to the technical field of communication. The speech transcription method comprises the following steps: based on a first speech, displaying a waveform graph corresponding to the first speech in a first area of a display screen, and displaying text information obtained by transcribing the first speech in a second area of the display screen; receiving a first input of a user; in response to the first input, determining a decibel range; filtering the first speech according to the decibel range, and updating the display of the waveform graph and the text information.
Owner:VIVO MOBILE COMM CO LTD

Speech processing software graphical user interface for electronic device

1. Name of the product subject to this design: Graphical User Interface for Voice Processing Software of Electronic Device. 2. Intended use of this design: for use in an electronic device. 3. The key design feature of this product is its graphical user interface. 4. The image or photograph that best illustrates the design's key features: the front view. 5. Purpose of the graphical user interface: This graphical user interface is used for speech processing operations and content display, such as speech recognition, speech transcription, and speech content translation. 6. Human-computer interaction method of graphical user interface: The main view is the main interface of the speech processing software. In the main view, after the user clicks the tabs of "Real-time Recording", "File Import", "Simultaneous Interpretation" or "Dialogue Translation" at the top of the interface, they will be redirected to the corresponding function interface. After the user clicks the "Simultaneous Interpretation" tab at the top of the interface, a status graph will be displayed. The status change diagram is the interface for simultaneous interpretation. The interface displays the records of simultaneous interpretation. After the user clicks the "English to Chinese" or "Chinese to English" tab at the top of the interface, the corresponding simultaneous interpretation function will be activated. 7. Other situations requiring explanation: In each view, the letter "X" is used to represent areas of text, numbers, or symbols. The characters represented by the letter "X" are replaceable, and the number of letters "X" does not limit the number of characters that can be replaced.
Owner:HANVON CORP

Method and device for processing voice transcription text, equipment, medium and product

The embodiment of the invention relates to a method and device for processing a voice transcription text, equipment, a medium and a product. The method includes acquiring a video and a first speech transcription text corresponding to the video. The method further includes determining, using a model, modification information for the erroneous content in the first speech transcription based on the at least one video frame in the video and the first speech transcription. The method further includes generating a second speech transcription text, where the second speech transcription text is derived based on the first speech transcription text and the modification information. Through the method, on the premise of not changing the original video content, the wrong content in the voice transcription text can be corrected in a targeted manner, and the readability and reliability of the voice transcription result are enhanced while the content correction accuracy in the voice transcription text is improved.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Intelligent video editor for creating non-linear editing timeline

An automated video editing system facilitates the creation of non-linear editing (NLE) timelines using a single prompt from a user. The system automatically ingests digital media, including video, audio, text, and images, and processes them to generate a proxy version with extracted features such as speech transcription, shot detection, facial recognition, and text recognition. A prompt-driven editing engine interprets user input and generates an edit decision list (EDL) using a large language model, which guides the assembly of an edited video timeline. The system also applies advanced editing features-such as captioning, animated title cards, font and color styling, sound effects, and transitions-based on learned user preferences. Additionally, it enables contextual overlays, chapter cards, and hierarchical timelines, while continuously learning user preferences to personalize editing results. The system may be integrated with social media platforms and existing media libraries for content sourcing, customization, and automated publishing.
Owner:OPEN VIDEO EXPLORATION INC

System

An object of the system according to the embodiment is to accurately convert speech into text and further convert the text into natural speech.SOLUTION: A system according to an embodiment includes a voice acquisition unit, a text generation unit, a text editing unit, and a voice generation unit. The voice acquisition unit acquires voice. The text generation unit converts the voice acquired by the voice acquisition unit into text. The text editing unit edits the text generated by the text generation unit. The voice generation unit converts the text edited by the text editing unit into voice.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

Domain and user intent specific disambiguation of transcribed speech

A system includes a computing platform including processing hardware and a system memory storing software code. The processing hardware executes the software code to receive a transcript of speech by a user, generate a phoneme stream corresponding to the transcript, partition the phoneme stream into words, and aggregate subsets of the words to form candidate sentences. The software code is further executed to determine, using one or both of an entity identified in the transcript and a history of the user, one or both of a user intent and a context of the speech by the user, rank the candidate sentences and the transcript based on one or both of the user intent and the context of the speech by the user, and identify, based on the ranking, one of the candidate sentences or the transcript as the best transcription of the speech by the user.
Owner:DISNEY ENTERPRISES INC

Audio playback progress bar for audio recording graphical user interface for electronic device

ActiveCN309909167SGraphical user interfaceProgress bar
1. The name of the design product: audio playing progress bar of audio recording graphical user interface for electronic device. 2. The use of the design product: for an electronic device. 3. The design points of the design product: the interface content expressed by non-dotted line in graphical user interface. 4. The picture or photo that best shows the design points: design 1 front view. 5. Design 1 is designated as the basic design. 6. The use of graphical user interface: the whole and the claimed part of the graphical user interface of the design are used for audio playing, or audio playing of voice transcription text function. In each design, the user can view the information such as the duration and playing progress of the audio according to the audio playing progress bar in the interface shown in the front view, or drag the playing progress bar to adjust the playing progress of the audio. The playing progress of the audio is matched with the display of the text. 7. Other circumstances that need to be explained: "X" in the interface represents the character area. The claimed design does not include the content expressed by dotted line.
Owner:LENOVO (BEIJING) LTD

Medical interview system

ActiveJP2026069425APatient-specific dataMedical departmentAcoustics
A medical interview tool (system, program, method) that automatically transcribes patient speech into text and automatically collects and inputs interview information. This solves the problem of missing information that occurs in conventional medical interview systems, enabling more efficient interviews. [Solution] The medical interview system 2 automatically converts the patient's speech content from the patient terminal 3 into text using speech recognition and extracts symptom and disease information. Based on medical interview templates for each symptom / disease and medical department, it generates additional questions that automatically complete any missing information that cannot be obtained from the speech content and presents them to the patient. Furthermore, by working in cooperation with the medical system 4 to determine whether it is an initial or follow-up visit, it ensures that the medical interview is conducted appropriately. In addition, an automated identity verification process and the generation of additional questions according to priority enable efficient and accurate collection of medical interview information.
Owner:PRECISION CO LTD

Electronic medical record generation method and device, computer equipment and storage medium

The invention relates to an electronic medical record generation method and device, computer equipment and a storage medium. The method comprises the following steps: in response to completion of loading of a voice medical record AI service and triggering operation on a voice acquisition component in an application interface, acquiring audio stream data dictated by a doctor in real time, and transmitting the audio stream data dictated by the doctor to a voice recognition model of the voice medical record AI service for voice transcription processing; in response to triggering operation performed on a medical record generation component in the application interface and receiving text stream data dictated by a doctor, transmitting the patient information, the doctor information, the treatment record information, the target medical record template, the target cue word and the text stream data dictated by the doctor to a large language model of a voice medical record AI service to perform medical record generation processing; and in response to the received electronic medical record text stream data, generating an electronic medical record according to the electronic medical record text stream data and the medical record version identification number. By adopting the method, the recording efficiency and the electronic medical record accuracy can be improved.
Owner:BEIJING UNITED FAMILY HOSPITAL CO LTD

An adversarial sample restoration method based on multi-version pre-processing sequence fusion

The application discloses a plug-and-play voice confrontation sample defense method, first, the input voice is smoothed by Gaussian noise to generate a smooth voice sequence with subtle differences, then the smooth voice sequence is input into a multi-version preprocessing algorithm module, different implementation compression algorithms and enhancement algorithms are used to process the smooth voice to generate a multi-version voice sequence; then the generated multi-version voice sequence is input into a voice recognition system to obtain a corresponding voice transcription text sequence; finally, a mode voting weight allocation (MVWA) method is used to allocate weights to each text sequence, and then a recognizer output voting error reduction (ROVER) algorithm is used to estimate the benign transcription of the input voice according to the weights. Through the application, the aggressive confrontation sample can be restored to a benign sample, the safety performance of the voice recognition system is improved, and an effective restoration method is provided for the defense of the confrontation sample.
Owner:SOUTHEAST UNIV

Dialect speech recognition method, system and model

The present application belongs to the technical field of speech recognition, and particularly relates to a dialect speech recognition method, system and model. The dialect speech recognition method comprises: S1, generating pseudo labels and screening by a plurality of teacher models; S2, sub-model training; S3, data fusion; S4, student model training, repeating steps S3 to S4 until round fusion training is performed, and student models of each region are obtained after training. The present application combines multi-teacher model knowledge distillation and K nearest neighbor dynamic data fusion training to construct an end-to-end speech transcription method suitable for a multi-dialect scene, and is particularly suitable for a less resourceful dialect environment with a complex tone system, rich speech variants and scarce data.
Owner:GUIZHOU UNIV +1

A Depression Detection Method Based on Instruction Fine-tuning Multimodal Speech-Language Model

PendingCN122090880AOvercoming underutilizationfully excavatedSpeech recognitionSpeech inputSpeech sound
This invention proposes a method for depression detection based on a multimodal speech-language model with instruction fine-tuning. The steps are as follows: First, a multimodal depression instruction dataset is constructed, integrating three types of information—original audio, automatically transcribed text, and emotional descriptions—into structured instruction-response samples. The emotional descriptions are automatically generated using a dedicated large-scale emotional inference model. Next, a low-rank adaptation technique is employed to efficiently fine-tune the parameters of the multimodal speech-language model. Audio features are extracted through an audio encoder and mapped to the text embedding space, then fused with the text and emotional description embeddings in a multimodal manner. The model parameters are optimized using a cross-entropy loss function. Finally, the speech to be detected is input into the fine-tuned model, and the depression recognition result is output. The technical solution provided by this invention enables joint inference based on audio, text, and emotional information, significantly improving the accuracy of depression recognition while offering advantages in parameter efficiency.
Owner:EAST CHINA UNIV OF SCI & TECH

Professional speech-oriented Chinese term real-time voice transcription error correction method and system

The invention designs a professional speech-oriented Chinese term real-time voice transcription error correction method and system. The method comprises the following steps: firstly, segmenting a video stream into continuous time domain intervals based on slide page changes in a speech video; converting the speech segments in each time domain interval into original speech transliteration texts by using a streaming automatic speech recognition model, and extracting current page slide texts from a video stream at the same time; secondly, generating a Chinese term set based on the slide texts of the current page and the historical sliding window interval and a preset domain knowledge graph; and then, coding the original voice transliteration text and the Chinese term set respectively by adopting a coding method based on Chinese character phonetic forms, determining segments highly similar to the Chinese term set in the original voice transliteration text by adopting an improved KMP algorithm, and replacing the segments with corresponding vocabularies in the Chinese term set in real time, thereby finally completing error correction. According to the invention, Chinese term errors generated by a speech recognition system for professional speech can be accurately and efficiently corrected.
Owner:NANJING UNIV OF POSTS & TELECOMM

Risk summary generation method based on multi-modal data

This application discloses a risk summary generation method based on multimodal data, belonging to the field of secure content recognition technology. The method includes: acquiring multimodal data, the data type of which includes at least one of video image data, audio data, and text data; inputting the multimodal data into a risk identification and detection model to obtain multiple first semantic information and first detection results, the risk identification and detection model including a trained video image forgery detection sub-model, an audio forgery detection sub-model, a text forgery detection sub-model, an image-text summarization sub-model, and a speech transcription sub-model; associating the multiple first semantic information and the multiple first detection results to obtain second semantic information; and generating a risk summary based on the second semantic information and a risk knowledge base. This method improves the accuracy and effectiveness of risk content recognition and risk summary generation.
Owner:HUBEI UNIV OF TECH

Intelligent analysis system for students to self-adjust learning strategies

The invention provides an intelligent analysis system for students to self-adjust learning strategies. The system comprises a data acquisition and preprocessing layer, a text processing and enhancement layer, a core analysis layer and an application display layer. The data acquisition and preprocessing layer comprises an audio acquisition module and a voice transcription and structuring module, the text processing and enhancement layer comprises a text cleaning module and a semantic normalization module, and the core analysis layer comprises a concept map analysis module and a learning strategy analysis module. And the application display layer comprises a visual rendering module and an instrument panel integration module. According to the invention, total, objective and real-time collection of classroom discussion data is realized, and subjectivity and hysteresis of manual recording are eliminated; through deep semantic analysis, the knowledge structure of discussion content is disclosed, and core technical support is provided for empirical research of cultivating self-adjustment learning ability of students based on an intelligent learning analysis tool.
Owner:TIANJIN NORMAL UNIVERSITY

Systems and methods for securely captioning video calls

A computer-implemented method for securely captioning video calls may include (i) detecting, by a messenger application executing on a client device, speech captured by a microphone of the client device and video of a speaker of the speech being captured by a camera of the client device, (ii) parsing, by the messenger application on the client device, the speech to create a transcript of the speech, and (iii) transmitting, by the messenger application on the client device, the transcript of the speech to an additional device for display to a user of the messenger application in combination with the video of the speaker. Various other methods, systems, and computer-readable media are also disclosed.
Owner:META PLATFORMS INC

A text automatic word extraction method, device, equipment and medium

This application discloses a method, apparatus, device, and medium for automatic text extraction, comprising: receiving currently input speech data and determining the text length of the speech data; obtaining the position of a matched paragraph; determining multiple candidate segments of the current text based on the position of the matched paragraph and the text length, each candidate segment including multiple text characters; determining the confidence vector of each text character based on a pre-trained speech-to-text transcription model; determining the score corresponding to each candidate segment based on the confidence vector of each text character; and determining suitable candidate segments of the speech data based on the scores corresponding to each candidate segment, so as to extract words from the current text based on the suitable candidate segments.
Owner:HANGZHOU QIUGUOJIHUA TECHNOLOGY CO LTD