Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

143 results about "Speech transcription" patented technology

Transcription (linguistics), the representations of speech or signing in written form Orthographic transcription, a transcription method that employs the standard spelling system of each target language. Phonetic transcription, the representation of specific speech sounds or sign components.

Long video multi-modal understanding and question-answering method and system based on large model and retrieval enhancement generation

The invention discloses a long video multi-modal understanding and question-answering method and system based on large model and retrieval enhancement generation. The method comprises the following steps: 1) a multi-modal feature extraction module; 2) a multi-modal synchronization and alignment mechanism; 3) constructing a structured memory pool; 4) querying a drive generation mechanism; 5) incremental updating and memory compression strategy; and 6) unifying the multi-modal representation space. The invention provides a long video multi-mode understanding method fusing a large language model and retrieval enhancement generation, and aims to break through the limitation of a traditional method in the aspects of single-mode processing and semantic fragmentation. According to the method, video image features are extracted through a visual model (such as YOLO and ViT), voice transcription and environment voice description are obtained in combination with an audio model (such as Whisper and Qwen-Audio), and unified coding of vision, voice and audio in a long video is achieved. Then, a structured memory pool is constructed through semantic consistency segmentation and timestamp alignment technologies to store time slice data of different modalities.
Owner:GUANGZHOU BINGO SOFTWARE +1

Two-process error correction method and device for real-time speech transcription

The invention provides a two-process error correction method and device for real-time speech transcription, and relates to the technical field of speech processing, and the method comprises the steps: extracting the Mel spectrum features of each segment, inputting each Mel spectrum feature into a lightweight end-to-end model, and obtaining a preliminary transcription text; splicing the segments according to a preset number to obtain a plurality of long segments, and inputting each long segment into a speech recognition model to obtain a high-precision transcription text; performing text comparison on the preliminary transcription text and the high-precision transcription text according to the confidence degree set of the preliminary transcription text to obtain all error vocabularies in the preliminary transcription text; and performing corresponding error correction processing on each error vocabulary in the preliminary transcription text according to the type of the error vocabulary and the high-precision transcription text to obtain a final transcription text. According to the method, through a two-process transcription error correction mechanism of the preliminary transcription text and the high-precision transcription text, transcription error accumulation is reduced on the premise that the real-time performance is not affected, and the transcription accuracy in a complex scene is improved.
Owner:NANJING DOLPHIN INTELLIGENT TECH CO LTD

Live broadcast behavior tracking system based on deep learning

The invention relates to the technical field of live broadcast behavior monitoring, in particular to a live broadcast behavior tracking system based on deep learning, which obtains high-quality multi-source information and improves the accuracy of feature analysis by synchronously extracting image and audio data from a live broadcast video stream and combining frame extraction, image enhancement and voice recognition. According to the method, image features are extracted through a pre-trained convolutional neural network, audio features are extracted through a deep learning model, voice transliteration texts are fused, the weight of each modal feature is dynamically adjusted based on an attention mechanism, precise recognition of complex scenes and hidden violation behaviors is achieved, and image camouflage and latent language expression risks are effectively coped with. And the illegal type and confidence are output in real time, once suspected illegal behaviors are detected, alarm, interruption or shielding operation is triggered immediately, and related evidences are uploaded to an auditing database. And efficient, accurate and full-process management and control of the live broadcast violation behaviors are realized.
Owner:GUANGZHOU QUNGE INFORMATION TECHNOLOGY CO LTD

Speech recognition and transcription method and system based on multi-modal fusion and sentiment analysis

The invention relates to a speech recognition and transcription method and system based on multi-modal fusion and sentiment analysis, and relates to the field of speech recognizing.The speech recognition and transcription method comprises the steps that a target speech signal and auxiliary modal information of synchronous visual information and text context information are obtained firstly, and the speech signal is segmented and recognized to obtain speech feature vectors; the method comprises the following steps: extracting text context information to obtain a text auxiliary feature vector, carrying out multi-modal fusion on the text auxiliary feature vector and the text auxiliary feature vector to generate fusion feature representation so as to carry out voice transcription to obtain an initial transcription text, and carrying out sentiment analysis according to visual information and the initial transcription text to generate a sentiment feature tag; and finally, optimizing and correcting the initial transliteration text based on the label to obtain a target transliteration text, thereby solving the technical problems that the speech recognition transliteration is difficult to adapt to dialect diversity and the recognition accuracy and robustness are insufficient due to neglect of emotion information, and improving the recognition accuracy and robustness through fusion of multi-modal information and emotion analysis. The voice content can be recognized more accurately, the transcription text can be optimized, and the accuracy and quality of voice recognition transcription are improved.
Owner:山西益通电网保护自动化有限责任公司

Conference summary method based on AI large model

The invention relates to a conference summary summarization method based on an AI large model, and belongs to the technical field of voice processing, and the technical scheme specifically comprises the following steps: obtaining audio information, recognizing all spokesmen involved in the audio information by using a preset spokesman separation algorithm, and splitting the audio information into a plurality of audio segments according to the spokesmen; analyzing and processing each audio segment by using a preset voice transcription algorithm to generate a transcription text; and based on a preset large language model, extracting and summarizing the transcriptional text, and generating and outputting summarized contents for describing the audio information, so that a user can know the summarized contents. The method has the effect of optimizing the information definition and the text readability of the transcribed text.
Owner:SUZHOU CHUANGLUTIANXIA INFORMATION TECH CO LTD

Film and television play table book extraction method and device, storage medium and computer equipment

According to the movie and television play table book extraction method and device, the storage medium and the computer equipment provided by the invention, after an audio and video file of a movie and television play is split into a video file and an audio file, feature recognition is performed on the video file to obtain a subtitle text, speaker face information and a video understanding text; performing voice understanding on the audio file to obtain a voice transcription text and a voice understanding text; wherein the voice transcription text can be corrected into the standard transcription text with high accuracy through the subtitle text. Therefore, based on the face information of the speaker, the line segment of each speaker in the standard transcriptional text and the audio and video file is aligned, so that speaker information with accurate segmentation and semantic coherence can be obtained; and then, through combination with a character side-writing text generated by side-writing analysis on the speaker based on the video, the voice understanding text and the speaker information, table book information is constructed, and related feature description of the character can be covered on the basis of containing the line content, so that the content and depth of the table book are enriched.
Owner:GUANGZHOU QUWAN NETWORK TECH CO LTD +1

Intelligent nursing record generation method and device

The invention discloses an intelligent nursing record generation method and device, and the method comprises the steps: obtaining nursing voice data, and carrying out the noise reduction and sound enhancement processing to obtain a preprocessed voice stream; a speech recognition technology fusing ECAPA-TDNN and an x vector is adopted to recognize medical terminologies, and a text transcription result is generated; recognizing the voice of the target nurse from the preprocessed voice stream and the transcription result in combination with a speaker-independent model and a speaker condition model to obtain a voice transcription text of the voice; performing voice analysis and structured processing on the text by using a multi-modal self-supervised learning model and a hierarchical CNN-BiLSTM framework to obtain a structured text with clinical semantic features; key clinical information is extracted in combination with the medical ontology knowledge base, and a nursing record element set is obtained; and based on the nursing record element set, automatically generating a nursing record meeting the specification. According to the invention, the automatic generation from the nursing voice to the standardized nursing record is realized, and the efficiency and accuracy of the medical nursing record are improved.
Owner:THE FIRST AFFILIATED HOSPITAL ZHEJIANG UNIV COLLEGE OF MEDICINE

Summary generation for live summaries with user and device customization

Described techniques may be utilized to receive a transcription stream including transcribed text that has been transcribed from speech, and to receive a summary request for a summary to be provided on a display of a device. Extracted text may be identified from the transcribed text and in response to the summary request. The extracted text may be processed using a summarization machine learning (ML) model to obtain a summary of the extracted text, and the summary may be displayed on the display of the device. When an image is captured, an augmented summary may be generated that includes the image together with a visual indication of one or more of an emotion, an entity, or an intent associated with the image, the summary, or the extracted text.
Owner:GOOGLE LLC

Training reasoning method based on voice-text-image multi-mode contrast learning

The invention particularly relates to a training reasoning method based on voice-text-image multi-mode comparative learning, and relates to the technical field of artificial intelligence, and the method comprises the steps: constructing a voice-text-GUI screenshot triple, and processing element masks; generating positive and negative alignment pairs of the instruction clauses and the GUI elements; the multiple encoders process corresponding modes, and the Rewriter converts a voice transcription text into a structured sequence; through global / local double fusion and double contrast learning, coordinate loss optimization alignment is carried out. According to the invention, through global and local double-layer fusion and a comparative learning mechanism, fine-grained alignment of voice, text and GUI elements is realized; element masks are generated by means of SAM, post-processing optimization is carried out, accurate element-level visual features are extracted in combination with RoIAlign, and a foundation is laid for alignment; all modal dimensions are unified through linear projection, coarse-grained semantic association is realized through a global fusion layer, and a local fusion layer fuses voice rhythm information to enhance association through cross-modal attention and fine correspondence of a gating mechanism capture sub-step and elements.
Owner:杭州长望智创科技有限公司

Multi-person overlapped voice real-time voiceprint recognition method and system

The invention relates to the field of voiceprint recognition and voice transfer, and discloses a real-time voiceprint recognition method and system for multi-user overlapped voices. The method comprises the following steps: acquiring a multi-source audio stream, and carrying out standardization and framing processing to generate a streaming audio frame sequence; based on the sequence, through circular buffering, noise reduction, endpoint detection and overlap detection model processing, obtaining an overlap interval label; task assembly, speaker separation, track numbering, voiceprint feature extraction and identity judgment are carried out, and a track identity binding structure is generated; and finally, voice transcription, fragment splicing and conflict cutting are executed, and the voiceprint template library is updated. According to the method, real-time separation and identity recognition of overlapped voices of multiple persons are realized, and the accuracy and robustness of voiceprint recognition in a complex scene are effectively improved.
Owner:HUNAN ZHENTONG ZHIYONG ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD

Intelligent conference voice transcription and summary generation method and system for voiceprint recognition and hot word optimization in power industry

The invention discloses an intelligent conference voice transcription and summary generation method and system for voiceprint recognition and hot word optimization in the power industry. The method comprises the following steps: acquiring voice samples in advance to construct an encrypted voiceprint database; mining terminologies from multi-source data in the power industry and dynamically allocating weights to establish a hot word library; performing unified decoding and adaptive preprocessing on the input audio; after speech features are extracted, speaker separation and recognition are achieved through deep embedding clustering; integrating the hot word bank, and generating a transliteration text with a speaker tag by adopting a transform-based field adaptive model; and carrying out multi-granularity semantic understanding and hierarchical abstract generation by utilizing the pre-training model in the power field, and automatically outputting a structured conference summary. According to the method, the problems that the accuracy of professional term recognition in the power industry is low, separation of multiple speakers is difficult, and summary generation depends on manpower are effectively solved, and the efficiency and the intelligent level of conference recording are remarkably improved.
Owner:CHINA SOUTHERN POWER GRID ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD

Label generation method and device based on emotion recognition, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business scenes such as financial science and technology and medical health, and discloses a label generation method and device based on emotion recognition, equipment and a medium, and the method comprises the steps: carrying out the preprocessing of a target audio, and obtaining the processed audio data; executing voice transcription based on the processed audio data and extracting acoustic features; performing emotion analysis on the acoustic features and the transliterated text; executing intention and appeal recognition based on the transliteration text; performing enhanced matching on the emotion analysis result and the intention and appeal recognition result in a knowledge base; and generating a label set according to a knowledge base enhancement result. According to the method, the acoustic and semantic information of the audio data is jointly processed, and the knowledge base is combined for enhancing matching to realize multi-dimensional tag generation, so that the recognition accuracy and the industry adaptability are improved, and the structured application of the audio data is supported.
Owner:PING AN TECH (SHENZHEN) CO LTD

Interview evaluation method and device, electronic equipment and storage medium

The invention relates to an interview evaluation method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring a target interview video, wherein the target interview video at least comprises video images of interview candidates and voices of the interview candidates and interviewers; generating a voice transcription interview text of the target interview video through a voice transcription service of the first large model, wherein the voice transcription interview text comprises an interview candidate text and an interviewer text; respectively generating corresponding semantic vectors based on the interview candidate text and the interviewer text; generating a plurality of expression vectors corresponding to different specified expressions based on the interview video; and combining the semantic vectors of the two interviews and the plurality of expression vectors into an input vector of an interview evaluation model, inputting the input vector to the interview evaluation model, and outputting an interview score through the interview evaluation model. According to the embodiment of the invention, the comprehensiveness, accuracy and scientificity of interview evaluation can be effectively improved, and the cost is low.
Owner:QIAN JIN NETWORK INFORMATION TECH SHANGHAI LTD

Intelligent quality inspection system and method based on voice transcription and emotion recognition

The invention relates to the technical field of intelligent voice processing and service quality monitoring, and discloses an intelligent quality inspection system and method based on voice transcription and emotion recognition. The method comprises the following steps: firstly, carrying out standardized preprocessing on multi-channel call data; performing voice transcription, speaker separation and verbal skill analysis, and constructing a session context; performing multi-modal emotion recognition and service state labeling by fusing the text and the acoustic features; and performing comprehensive quality inspection analysis in combination with rule matching and model reasoning, and forming closed-loop optimization of a quality inspection strategy. According to the method, the problems of high subjectivity and low coverage rate of a traditional quality inspection method are effectively solved, automatic, multi-dimensional and high-precision quality inspection and risk insight of mass calls are realized, and the efficiency and objectivity of customer service quality inspection are remarkably improved.
Owner:GUANGZHOU XUNHONG NETWORK TECH CO LTD

Automatic depression risk assessment method based on natural language processing

The invention relates to an automatic depression risk assessment method based on a natural language processing technology. The method comprises the following steps: performing selective conversation content interviews with a subject by using a voice conversation function of artificial intelligence; dialogue recording and text data between the patient and artificial intelligence are obtained; performing automatic transcription and speaker separation on the original speech transcription text data, and constructing a structured question and answer dialogue sequence; generating a structured question prompt related to the depressive symptom by using a large language model, and splicing the structured question prompt with the original dialogue content; designing an artificial neural network pre-training language model coding word vector, fusing position coding, a self-attention mechanism, a problem-guided attention mechanism and a bidirectional long short-term memory network, and extracting deep semantic features related to depression risks; finally, continuous depression risk scores are output, correlation verification is carried out on the continuous depression risk scores and standard PHQ-8 scores, and the accuracy, robustness and interpretability of depression risk assessment are effectively improved.
Owner:HARBIN UNIV OF SCI & TECH

Multilingual full-speech processing method and device based on speech recognition, and medium

The invention discloses a multilingual full-speech processing method and device based on speech recognition and a medium, and relates to the technical field of speech recognition, and the method comprises the steps: calculating a language explicit trajectory based on a multilingual speech feature set, carrying out the association with a speech segment through language preference information in a historical session, and constructing a language implicit trajectory, integrating and generating a language weight track; dividing the language weight trajectory into voice trajectory nodes, recording multi-language voice trajectory composite features, connecting the multi-language voice trajectory composite features into a voice trajectory chain, calculating inter-node continuity indexes, and generating a voice trajectory node structure; constructing a track node continuity credibility field according to the voice track node structure, and adjusting multilingual voice recognition decision parameters to generate a multilingual transcription candidate set; and carrying out time sequence splicing and language mark arrangement on the language transcription candidate set to generate a multi-language full-voice transcription result set. According to the method, self-adaptive decoding of structure perception is realized, language switching is optimized, and the transcription precision is improved.
Owner:CHANGCHUN VOCATIONAL INST OF TECH

System and Method for Flowsheet Population

A method, computer program product, and computing system for flowsheet population by voice. Dictation or conversation is transcribed from voice to text. The keys and values relevant to the documentation are extracted from the transcript and used to populate the flowsheet by selecting a subset of rows and examples relevant to the transcription of the information from a plurality of rows in the flowsheet and based on the transcript, each row corresponding to a key and a value for the flowsheet and extracting an instance of information from the transcript of information associated with a key from the subset of rows by processing a prompt including at least a portion of the transcript of information and the subset of rows with a generative artificial intelligence (AI) model using retrieval augmented generation (RAG).
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Privacy-preserving avatar voice transmission

Some implementations relate to methods, systems, and computer-readable media to reproduce a voice stream with removed personally identifiable information (PII). An original voice stream is received from a user by a hardware processor. The original voice stream is converted into text using a machine learning (ML) voice transcription model. Then, the text is used for generating an anonymized voice stream corresponding to the text having similar auditory properties to the original voice stream other than PII using a machine learning (ML) voice synthesis model. The similar properties are determined using metadata corresponding to the original voice stream. The metadata may be obtained by analyzing acoustic characteristics of the original voice stream. The metadata may include volume changes and pitch frequency changes stored as metadata annotations associated with tokens in the text as key-value pairs. The generation may use the metadata annotations to determine the similar audio properties.
Owner:ROBLOX CORP

Bulk commodity transaction market stabilization method based on BSL-CPFS deep learning model

The invention discloses a bulk commodity transaction market stabilization method based on a BSL-CPFS deep learning model. The method comprises the following steps: preprocessing a voice data set into a training set and a test set; introducing a DeepSpeech2 model as a speech transcription model, and training the speech transcription model by using a training set and a text set; constructing a BSL model based on a BERT model and a TextCNN, and inputting the speech transcription text data training set for training to obtain a BSL language processing model; an NER model is introduced, the data set is labeled, and the NER model is input for training to form structured data; constructing a CPFS model based on a BiLSTM network, and inputting structured data for training to obtain a price prediction model; a prediction result and historical price fluctuation are dynamically compared to identify an abnormal fluctuation signal and push the abnormal fluctuation signal in real time, the asymmetry of information is reduced, a supervision department continuously monitors the deviation between the market price and the prediction result and formulates a targeted intervention strategy, then the price correction is realized, the stability of a bulk commodity transaction market is improved, and the market competitiveness is improved. And the market pricing transparency is improved.
Owner:DALIAN MARITIME UNIVERSITY

Power self-service terminal voice intention recognition system based on background noise optimization

The invention relates to the technical field of voice recognition, in particular to an electric power self-service terminal voice intention recognition system based on background noise optimization. The system comprises a voice interaction noise optimization module, a scene voice transcription analysis module, a scene voice intention recognition module and a recognition confidence iterative optimization module, a voice interaction original data set during voice interaction between the power self-service terminal and a user can be collected, and noise optimization is carried out to generate a target voice interaction data set; performing voice transcription analysis based on the target voice interaction data set, and generating a text transcription result; performing keyword semantic matching calculation on the text transcription result and an intention recognition knowledge base to obtain an intention recognition result of the user; and calculating a comprehensive recognition confidence coefficient based on the intention recognition result, if the comprehensive recognition confidence coefficient is lower than a set threshold value, performing secondary noise optimization on the voice interaction original data set and re-executing the judgment step, and if not, outputting the intention recognition result of the user. According to the invention, the accuracy of terminal speech recognition and intention recognition can be improved.
Owner:STATE GRID TIANJIN ELECTRIC POWER COMPANY

IMPROVING SPEECH RECOGNITION TRANSCRIPTIONS

ActiveDE102021122068B4Speech recognitionAudiometry testCommon word
Computer-implemented method (500) for training a model to improve speech recognition, wherein the computer-implemented method comprises: Receiving (502) an utterance by one or more processors, wherein the receiving is performed by a virtual assistant in a specific node of the virtual assistant, wherein frequently occurring terms have been identified for the specific node over a period of time; Transcription (504) of the utterance into text by the one or more processors; Generating (506) a transcription confidence score based on transcription and audiometry by one or more processors; in response to the transcription confidence score being below a threshold, comparing (510) phonemes in the utterance with phonemes in at least one term from a list of frequently occurring terms by one or more processors; Generating a sound similarity score for phonemes in which at least one term from a list of frequently occurring terms is selected by one or more processors; and Replacing (512) the transcription with at least one term from the list of frequently occurring terms if the sound similarity score is above a threshold, by one or more processors.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Method and equipment for assisting in capturing voice keywords and calling evidence in interrogation scene

The invention provides an interrogation scene voice keyword capturing and evidence calling auxiliary method and device. The method comprises the steps that S1, system initialization and parameter configuration are carried out; s2, voice signal acquisition and preprocessing; s3, converting voice into text and enhancing hot words; s4, keyword recognition and semantic understanding; s5, performing evidence library retrieval and multi-algorithm fusion matching; s6, performing multi-evidence fuzzy matching processing; and S7, automatically displaying the evidence. According to the invention, the voice transfer accuracy is improved, and the keyword omission ratio is reduced; the multi-evidence screening time is shortened, the evidence similarity is automatically adjusted after confirmation by a judge, and the accuracy of subsequent similar query is improved; the instruction recognition accuracy of voice control file page turning is improved, and court trial interruption times are reduced.
Owner:辽宁速服达数据科技有限公司

Speech transcription method and device and electronic equipment

The invention provides a voice transcription method and device and electronic equipment. The method comprises the following steps: acquiring audio data in real time; sampling the audio data based on a first rule to obtain first audio data; processing the first audio data through a target model to generate a first text; sampling the audio data based on a second rule to obtain second audio data; processing the second audio data through the target model to generate a second text; wherein the second text replaces the first text; wherein the duration of the first audio data is smaller than that of the second audio data; each second audio data comprises two adjacent first audio data, and two adjacent second audio data comprise the same first audio data.
Owner:LENOVO (BEIJING) LTD

Creation method and system based on speech recognition and text generation model

The invention provides a text creation method and system based on speech recognition and a text generation model, and relates to the technical field of natural language processing and artificial intelligence. Carrying out redundant word filtering on the voice data through acoustic feature analysis, and carrying out dialect phoneme correction by using a phoneme mapping table to obtain structured voice data; converting the structured voice data into a text through an encoder-decoder voice recognition model, and extracting a semantic triple and an emotion tag in combination with emotion analysis and dependency syntactic analysis to form structured information; generating a constraint parameter table according to the target body, inputting the structured information and the constraint parameter table into a text generation model, and generating a to-be-processed text; fusing the user preference information with the to-be-processed text to generate a first target text; and evaluating the text quality based on a resonance index threshold, and generating a final target text through iterative optimization. According to the method, the accuracy of voice transcription of the old people is improved, and the effectiveness of text generation in the aspects of form specification and emotion expression is enhanced.
Owner:HAINAN MIRROR DIGITAL CREATIVE CO LTD

Speech-to-speech translation

PendingUS20260154515A1Natural language translationSound input/outputSpeech to speech translationSpeech translation
A speech-to-speech translation method comprises transcribing speech spoken in a source language into transcribed text data in the source language using an on-premises speech recognition model. The transcribed text data is translated into translated text data in a target language using a first on-premises machine translation model. The translated text data is reverse translated into retranslated text data in the source language using a second, different on-premises machine translation model. The transcribed text and the retranslated text are displayed on a screen. The method also involves synthesizing, using an on-premises speech synthesis model, translated speech data in the target language based on the translated text data and play back, in response to a user confirmation, translated speech in the target language based on the translated speech data in the target language.
Owner:MABEL AI AB

Natural language processing systems and methods for intent classification of speech transcription

Aspects of the subject disclosure may include, for example, generating a natural language processing model by training an automatic speech recognition (ASR) encoder with manual transcription. The training is performed by correcting and adjusting relevant factors of the ASR encoder based on determined triplet loss, classification loss and Kullback-Leibler divergence loss. In response to an ASR utterance, the trained natural language processing model generates a predicted intent associated with the ASR utterance with improved accuracy. Other embodiments are disclosed.
Owner:JPMORGAN CHASE BANK NA

Ultrasound report generation method, related equipment and computer program product

The invention discloses an ultrasonic report generation method, related equipment and a computer program product, and relates to the technical field of artificial intelligence. According to the method, streaming recognition is performed on the real-time spoken voice of the doctor in the ultrasonic detection process of the doctor, the current voice transcription text fragment is obtained, the large model is called to extract the ultrasonic detection parameters in the current voice transcription text fragment in a fast thinking mode, and the extracted ultrasonic detection parameters are output and previewed in real time, so that the accuracy of ultrasonic detection is improved. Therefore, real-time interaction with doctors is generated, and the result of the detection process is transparent to the doctors. After ultrasonic detection is finished, a large model is called to generate ultrasonic description and diagnosis based on an original complete voice transfer text in a slow thinking mode, ultrasonic detection parameters, the ultrasonic description and diagnosis are integrated to obtain a high-quality ultrasonic report, the clinical requirement of real-time interaction is met, the clinical working efficiency is improved, and the clinical experience is improved. And the quality of the finally obtained ultrasonic report can be ensured.
Owner:ANHUI IFLYHEALTH CO LTD +1

Method and apparatus for inputting voice into webpage text box of intranet website

The application provides a method and device for voice inputting into a webpage text box of an intranet website, to solve the technical problem that reading microphone audio through the webpage of the intranet website and calling voice recognition function are limited. The method comprises the following steps: receiving a voice input instruction of a webpage text box of an intranet website; using a Web Socket protocol, the webpage sends the voice input instruction to a local microphone application program; under the control of the voice input instruction, the local microphone application program collects voice signals; processing the voice signals to generate a character set; and inputting the character set into the webpage text box of the intranet website. In this way, the intranet user can call the voice recognition function and automatically input the voice recognition text into the webpage text box under the condition of intranet communication, so that the voice transcription of the intranet user is more convenient, and the accuracy and completeness of the input text can be effectively improved, and the work efficiency of the intranet user is greatly improved.
Owner:BEIJING THUNISOFT INFORMATION TECH

Automatic software testing method and system based on multi-modal AI collaboration

The present application relates to the technical field of automatic testing, more particularly, to an automatic software testing method and system based on multi-modal AI collaboration. The scheme includes encapsulating optical character recognition, speech recognition and natural language processing analysis tools into independent containers, and inputting the analysis results into a multi-modal semantic correlation model in real time; calculating the semantic similarity between image description text and corresponding speech transcription text; constructing a test requirement complexity evaluation model; defining a use case quality index, and establishing an error sample library to store artificially annotated problem use cases and their correction labels; finally determining the intended requirements according to the screening results of the three-level filter; setting a test case template, optimizing the path according to the fusion quality, improving the use case generation efficiency, and synchronously updating the path to the knowledge base associated with the error sample library. The scheme realizes automatic software testing through multi-modal AI collaboration, containerized deployment and model dynamic optimization, thereby improving the efficiency and recognition accuracy.
Owner:RUIJIAN TECHNOLOGY (BEIJING) CO LTD

English video splitting learning method based on large model

The invention relates to the technical field of video splitting, in particular to an English video splitting learning method based on a large model. According to the method, a teaching text and a voice transcription text of a teaching video are obtained through ASR and OCR technologies respectively, the text is segmented into single sentences, vocabularies are extracted and stored in a difficulty vocabulary library, and an integrated video meeting the requirements of an examination outline is generated through integration according to the expected learning level and difficulty quantitative indexes of a user. Segmenting the video length according to the learning duration demand of the user; a user playing record, a repetition rate, a complete playing rate and interaction feedback data are collected, a user learning level and video teaching difficulty are comprehensively evaluated through a big data model, and video content is dynamically adjusted. According to the scheme, the limitation of a traditional static splitting method is filled, the effects of effectively reducing knowledge omission and improving the learning efficiency are achieved through dynamic optimization and personalized customization of the teaching content, meanwhile, accurately matched teaching resources are provided for learners of different levels, and the learning efficiency is remarkably improved.
Owner:读书郎教育科技有限公司