Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

4403 results about "Subvocal recognition" patented technology

Subvocal recognition (SVR) is the process of taking subvocalization and converting the detected results to a digital output, aural or text-based.

Real-time virtual reality scene system based on natural language description using multimodal artificial intelligence

A real-time system for the multimodal generation of virtual reality scenes based on artificial intelligence for the creation of immersive three-dimensional environments from natural language narratives, consisting of: a speech capture module configured to continuously record a user's spoken narrative via one or more directional microphones, preprocesses the captured signal by noise reduction and temporal alignment, and outputs a digital speech stream; A speech-to-text processing unit that is operationally coupled to the speech capture module and configured for real-time speech recognition using a continuous neural transformer model. The unit is trained to transcribe natural language utterances into structured text data while maintaining contextual continuity throughout the evolving narrative. a semantic interpretation processing unit that is communicatively linked to the speech recognition unit and configured to perform natural language understanding techniques to extract contextual entities, spatial references, temporal relationships, and object attributes from the transcribed narrative; the engine includes a large language model that is fine-tuned for spatial reasoning tasks; a scene graph generation module configured to transform the interpreted semantic data into a structured, hierarchical representation that defines nodes for identified entities and edges for corresponding relationships, with each node associated with metadata describing geometry, position, orientation, texture, and linking attributes between objects; a multimodal image-language model processor coupled with the scene graph generation module, wherein the processor is configured to retrieve, adapt, or synthesize appropriate three-dimensional elements from a pre-trained visual-lexical embedding space and align these elements with their semantic and spatial definitions derived from the scene graph; a scene assembly and rendering controller configured to create a cohesive virtual scene from the aligned assets, perform real-time rendering using a GPU-accelerated ray tracing pipeline, and produce a stereoscopic visual output that corresponds to the evolving narrative; A head-mounted virtual reality visualization device connected to the rendering engine and configured to display the generated immersive environment to the user in real time. The device features motion sensors and inside-out tracking cameras to detect head and body movements, dynamically updating viewing angles and perspective within the rendered scene; and a bidirectional feedback module integrated into the head-mounted device and connected to the semantic interpretation processing unit; the module is configured to interpret corrective commands, gestures, or supplementary comments from the user to refine or modify specific scene elements without interrupting the real-time visualization; The system continuously updates the virtual scene as the narrative develops, ensuring temporal synchronization between speech input and rendered output below a defined latency threshold, thus enabling a natural, dialogic construction of complex three-dimensional virtual environments.
Owner:GOUNDER MOHAN SELLAPPA DR BENGALURU +3

Voice intention recognition method and device, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a voice intention recognition method, device and equipment and a medium, and the method comprises the steps: obtaining a to-be-processed voice signal, carrying out the voice activity detection processing of the voice signal, dividing the voice signal into a plurality of voice segments, analyzing semantic contents of the plurality of voice segments, determining semantic correlation information of each voice segment, analyzing sound source attributes of the plurality of voice segments, determining sound field type information of each voice segment, screening out a target voice segment from the plurality of voice segments according to the semantic correlation information and the sound field type information, and executing intention recognition processing based on the target voice segment to generate an intention recognition result. According to the invention, through a dual analysis mechanism of semantic correlation information and sound field type information, effective screening of voice segments is realized before voice recognition, non-target voice or interference segments are effectively prevented from being sent to an intention recognition model, and the accuracy of a recognition result is improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Text prediction-based large-model real-time voice text intention recognition method and system

The invention discloses a large-model real-time voice text intention recognition method and system based on text prediction, and the method comprises the steps: obtaining the real-time voice data of a user, carrying out the real-time voice recognition processing through a streaming voice recognition interface, and obtaining a part of transcriptional text; inputting the partial transcription text into a mask language model for text prediction, and generating a plurality of high-credibility complete sentence candidates; based on the complete sentence candidates, the complete sentence candidates are input into a large language model in parallel for intention recognition, a corresponding intention result is obtained, and a mapping relation between the candidate sentences and the intention recognition result is established; and obtaining a sentence completely expressed by the user, calculating the similarity between the complete actual sentence and a plurality of high-credibility complete sentence candidates through a multi-level text similarity algorithm, selecting the candidate sentence with the highest similarity score, and directly obtaining a corresponding final intention recognition result based on the mapping relationship. The objective of the invention is to solve the technical problem of high response delay of an existing voice intention recognition system.
Owner:BEIJING YULORE INNOVATION TECH

Comprehensive AI-enabled systems for immersive voice, companion, and augmented / virtual reality interaction solutions

A computer-implemented method for operating an artificial intelligence voice agent system includes receiving voice input through communication channels; analyzing converted text through natural language processing (NLP) pipelines implementing intent recognition and sentiment analysis detecting emotional cues using a multimodal large language model (LLM); generating response content using machine learning models trained on domain-specific corpora; converting generated responses to synthetic speech through text-to-speech (TTS) engines; integrating with a customer relationship management (CRM) platforms or an enterprise resource planning (ERP) database; and implementing continuous learning by updating language understanding models using conversation logs, voice recognition parameters based on user feedback, and response generation patterns. One implementation is a computer-implemented system and method that operates a suite of intelligent interactive devices and platforms including an artificial intelligence voice agent, enhanced communication platforms, an intimacy companion system, and augmented / virtual reality eyeglasses. Further, one implementation includes AR / VR eyeglasses that project visual content onto interchangeable lenses or directly onto the user's retina via laser-based retinal projection, provide prescription adjustments, incorporate ear-mounted sensors for monitoring physiological parameters like heart rate, oxygen saturation, and blood pressure, and utilize wireless data transmission, onboard environmental sensing, and remote calibration, all designed to offer dynamically adaptive, secure, and context-aware interactions across communication, personal assistance, health monitoring, and immersive augmented or virtual reality environments.
Owner:TRAN BAO

Voice interaction method and device based on lip language enhancement, equipment and storage medium

The invention discloses a voice interaction method and device based on lip language enhancement, equipment and a storage medium, and the method comprises the steps: extracting lip language features based on an image sequence of a lip region, and carrying out the feature extraction of a voice signal, and obtaining an audio feature; performing cross-modal fusion coding on the lip language features and the audio features to generate mixed features containing audio-visual information; inputting the mixed features into a large language model, understanding the intention of the interaction object and generating a corresponding semantic reply; and finally, synthesizing into voice and / or converting into characters. According to the invention, by introducing the lip features, additional visual clues are provided for speech recognition, and the robustness and accuracy of speech recognition can be significantly improved; effective fusion coding is carried out on the lip language features and the sound features, and semantic information splitting caused by simple and independent recognition is avoided; and the capability of the large model is fully utilized, so that more natural and more intelligent interaction experience is realized.
Owner:SHENZHEN WANRUI INTELLIGENT TECH CO LTD

Voice interaction method and system of AI intelligent robot

The invention relates to the technical field of voice interaction, particularly discloses an AI intelligent robot voice interaction method and system, and aims to solve the problems of low voice interaction accuracy, insufficient reliability and lack of authority control in a complex noise environment. A dynamic noise feature library containing steady-state noise, impact noise and human voice interference features and a pre-stored gesture instruction library are constructed, audio signals are collected in real time, low-frequency-band, middle-frequency-band and high-frequency-band differential noise reduction is executed, Mel-frequency cepstral coefficient features are extracted, noise scenes are matched, corresponding voice recognition models are switched, and voice recognition is achieved. And calculating a confidence value of the voice instruction, outputting multi-modal verification data in combination with a dynamic confidence threshold, and outputting an authority control signal through voiceprint matching, authority verification and instruction consistency judgment. Through multi-modal fusion, dynamic adaptation and authority control, the voice recognition accuracy and interaction safety in a complex noise environment are remarkably improved, and the method is suitable for scenes such as factory intelligent inspection.
Owner:HANGZHOU SOHA TECH CO LTD

Voice interaction optimization method and system based on multi-modal large model

The invention discloses a voice interaction optimization method and system based on a multi-modal large model, and relates to the technical field of artificial intelligence and voice interaction, and the method comprises the steps: carrying out the voice recognition in response to the real-time voice of a user, and obtaining voice text information; according to the voice text information, combining the voice waveform of the real-time voice of the user as the input of a multi-modal recognition model, so as to judge whether the voice dialogue is interrupted and recognize the interruption intention, and obtaining a voice interruption result; obtaining a new intention of the user according to a voice interruption result, and dynamically adjusting a system response strategy to realize interaction optimization; the accuracy and comprehensiveness of interruption detection are improved through multi-modal fusion, so that the interruption intention is accurately recognized to respond to the user intention in real time, the system dialogue interaction efficiency and reliability are remarkably improved, and the defects that existing voice interruption detection is not high in detection accuracy, dynamic response of user interruption behaviors cannot be achieved, and user experience is poor are overcome. Therefore, the problems of low interaction efficiency and poor reliability of the voice dialogue system are solved.
Owner:HANGZHOU YIWISE INTELLIGENT TECH CO LTD +1

Speech recognition authentication method and system based on multi-modal features and dynamic evaluation

The invention discloses a speech recognition and authentication method and system based on multi-modal features and dynamic evaluation in the technical field of speech recognition and authentication, and the method comprises the steps: collecting an original speech signal of a user through a microphone, and carrying out the preprocessing of the original speech signal, and obtaining the preprocessing speech data; and extracting feature data of the preprocessed voice data by adopting a multi-dimensional feature hierarchical extraction technology, and injecting a multi-source noise sample into an acoustic feature space of a voiceprint feature model based on an initial training stage established by the voiceprint feature model to construct an anti-noise mixed voiceprint map. Through integrating voiceprint, semantics, behavior characteristics and an environment adaptation mechanism, an authentication threshold is adjusted in real time according to dynamic risk assessment, meanwhile, a risk scoring model is utilized to calculate a comprehensive risk value, authentication modes of different levels are started according to risk scenes of different degrees, and two-factor authentication is forcibly implemented for high-risk scenes. And the authentication security and reliability can be obviously enhanced.
Owner:JIANGSU VARIABLE SUPERCOMP TECH

Methods and systems of text-conditioned audio-visual speech generation with multi-modal latent diffusion models

Methods, systems, and computer programs are presented for audio-visual speech generation with multi-modal latent diffusion models. One method includes encoding raw audio signals and video frames into respective latent spaces using audio and visual autoencoders. A text transcript is processed into phoneme sequences using a text transcript processor. The audio and visual latent spaces are conditioned using the text transcript and a conditioning variable. Joint distributions of the visual and audio latent spaces, text transcript, and conditioning variable are learned using a multi-modal latent diffusion model. The model adds noise to the latent audio-visual representations and predicts the noise through denoising neural networks. An inverted diffusion process is utilized to generate diverse speech content and speaker characteristics, resulting in realistic audio-visual speech. The technology presented provides a novel approach to conditional speech generation with potential applications in speech synthesis, voice conversion, and speech recognition.
Owner:TENSORTYPE INC

Intelligent conference summary automatic generation method based on voice recognition and large model

The invention discloses an intelligent conference summary automatic generation method based on voice recognition and a large model. The method comprises the following steps: S1, executing voice activity detection operation on an audio data stream; s2, extracting embedding vectors of continuous and effective voice segments, and generating a voice segment set to which a spokesman belongs; s3, inputting the voice fragment set to which the spokesman belongs into an improved Whisper model, fusing a Speaker-Aware attention mechanism and a connection time sequence classification auxiliary path, and outputting a conference transcription text sequence set; s4, inputting the processed structured dialogue format into a GPT-4 large language model, and generating a conference semantic representation sequence; s5, generating a conference summary first draft text according to a preset summary generation template; and S6, performing formatting output operation on the conference summary first draft text. The conference semantic elements can be automatically extracted, the structured summary text can be generated, and the method is suitable for efficient conference recording and task tracking in government affair office, enterprise collaboration, academic discussion and other scenes.
Owner:JIANGSU GUOHUACHENJIAGANG POWER GENERATION CO LTD

Intelligent safety protection management method and system

The invention relates to an intelligent safety protection management method and system. The method comprises the following steps: acquiring behavior data and position data of personnel in an industrial site; analyzing the behavior data and the position data by using a pre-trained personnel behavior model to obtain a behavior recognition result; according to the behavior recognition result and the current operation state of the target equipment, judging a risk level corresponding to the personnel behavior; based on the risk level, a corresponding safety response strategy is matched in the edge computing node, a control instruction is generated according to the safety response strategy, and the control instruction is used for driving the target device to execute a corresponding response action; and the voice interaction module is used for acquiring voice input of an on-site operator, performing semantic recognition on the voice input to obtain a voice recognition result, and executing corresponding emergency control operation to cover a current control instruction when the voice recognition result meets a preset emergency instruction condition. The method has the effect of improving the accuracy of intelligent safety management in the industrial environment.
Owner:SHENZHEN HUAYIXIN ELECTRONICS CO LTD

Voice interaction task execution method and device based on large model, equipment and medium

The invention discloses a voice interaction task execution method and device based on a large model, equipment and a medium, and relates to the field of artificial intelligence, and the method comprises the steps: carrying out the voice recognition of voice features, obtaining a text character sequence, and carrying out the optimization of the text character sequence through a preset large language model, and obtaining a target text character sequence; performing entity recognition on the target text character sequence by using a preset large language model to obtain an entity recognition result, determining a relationship type among entities in the target text character sequence according to the entity recognition result, and constructing a knowledge graph according to the relationship type among the entities; generating an initial triple based on the knowledge graph, and optimizing the initial triple by using a preset large language model to obtain a target triple; and fusing the target triple with the initial knowledge graph, and executing a voice interaction task in the target voice interaction scene based on the updated knowledge graph. According to the method and the device, the efficiency and the accuracy of extracting the structured knowledge from the Chinese speech are improved.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Speech recognition method and related device

ActiveCN114360510AImprove fault tolerancePrecise Syllable Probability DistributionSpeech recognitionSyllableAcoustic model
The embodiment of the invention discloses a speech recognition method and a related device, and at least relates to a speech recognition technology in artificial intelligence, speech data to be recognized are used as input data of a time delay neural network in an acoustic model, and an output layer of the time delay neural network comprises acoustic modeling units corresponding to a plurality of syllables respectively, so that the speech recognition efficiency is improved. And the syllable probability distribution corresponding to the voice frames included in the voice data can be obtained by taking the syllables as the recognition granularity through the time delay neural network. When syllable recognition is carried out through the output layer, auxiliary judgment can be carried out on the syllables to which the voice frames belong on the basis of pronunciation rules in combination with front and back syllable information of the voice frames, so that more accurate syllable probability distribution is output. Moreover, since the syllables are generally composed of one or more phonemes, the method has higher fault-tolerant capability, not only can more accurately determine the speech recognition result based on the probability distribution of the syllables, but also has low requirements for the quality of the speech data to be recognized, and effectively expands the application scenarios of the speech recognition technology.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Electronic medical record LLM generation method based on animal injury

The invention discloses an electronic medical record LLM generation method based on animal injury, which realizes dialogue structuring and timestamp synchronization through multistage speech recognition and role affiliation. Using standardized medical term mapping and coding to align the free text to a standardized medical entity, and constructing a high-confidence medical entity network based on a semantic anchor point pool; according to the method, context-sensitive entity relationship extraction is realized by combining a large language model and a semantic enhancement template, a high-accuracy structured relationship chain is generated through clinical logic rule set verification, and finally, an electronic medical record template under diagnosis and treatment specifications is automatically filled and privacy desensitization processing is completed. The semantic consistency, the structural accuracy and the data security of automatic generation of the electronic medical record are improved, and standardization and intelligent circulation of medical information are effectively promoted.
Owner:GUANGZHOU WUCHUAN ELECTRONIC TECHNOLOGY CO LTD +1

Audio and video recording-based ASR identification enhancement method

The invention discloses an ASR identification enhancement method based on audio and video recording. According to the method, the accuracy and compliance of voice recognition in the financial service interaction process are improved by fusing the audio and environment feature information in the banking business double-recording scene. The method comprises the following steps: firstly, constructing an acoustic model for a bank outlet environment, and extracting audio features and interaction scene information of conversation between a client and a worker; and then, designing a vocabulary recognition module special for the financial field, dynamically adjusting language model parameters according to professional term libraries and utterance modes of different business types, and effectively coping with key links such as financial product introduction, risk prompt and customer confirmation. Compared with a traditional ASR system, the voice recognition accuracy in the banking business handling process is remarkably improved, particularly, key term recognition and important information extraction are prominent, and more reliable technical support is provided for financial service standardized management and double-recording quality inspection.
Owner:GUANGZHOU BAIRUI NETWORK TECH CO LTD

Intelligent Bluetooth voice remote control system based on AI semantic analysis

The invention relates to the technical field of intelligent voice interaction, in particular to an intelligent Bluetooth voice remote control system based on AI semantic analysis. Comprising a voice acquisition unit, a Bluetooth communication unit, a voice recognition unit, an AI semantic analysis unit and a control execution unit, and the Bluetooth communication unit is used for establishing low-power-consumption Bluetooth connection with target equipment and supporting bidirectional data transmission. By combining advanced localized model library, edge calculation optimization, multi-modal data fusion and dynamic semantic map technologies, the recognition accuracy, the real-time response speed and the context understanding ability of the voice instruction are remarkably improved, and meanwhile, the data privacy is guaranteed, so that efficient, personalized, natural and smooth user interaction experience is realized.
Owner:SHENZHEN XINGWEI TECHNOLOGY CO LTD

Robust audio and video speech recognition method and device based on multilayer perception fusion

The invention discloses a robust audio and video speech recognition method and device based on multilayer perception fusion, and belongs to the technical field of audio and video multi-mode semantic modeling and speech recognition. According to the method, audio and visual bimodal input is utilized, a teacher-student structure is introduced in a training stage, and a student model is guided to learn stable semantic representation under various noise conditions through a self-distillation mechanism. In order to enhance the alignment capability and anti-interference performance between audio and video features, a multi-layer suppression and enhancement interaction module is introduced into the joint encoder, layer-by-layer fusion and noise suppression between modes are realized, and a robust multi-mode fusion encoder (RMIE) is constructed. The RMIE models modal alignment and feature enhancement in a multi-level semantic space at the same time, and the semantic offset problem caused by modal difference and noise interference is effectively relieved. Furthermore, a decoder based on an attention mechanism is introduced on the basis of the RMIE, and an audio and video speech recognition model with end-to-end recognition capability is obtained through fine tuning.
Owner:SICHUAN UNIV

Model training and speech recognition method and device, equipment and medium

The invention discloses a training method of a voice recognition model based on noise deconstruction, a voice recognition method, a device, equipment and a medium. According to the method, a staged training strategy that a noise unwrapping module is firstly isolated and trained and then a Conformer-Transducer architecture is finely adjusted and trained is adopted, so that high calculation complexity and training difficulty caused by simultaneous training of a plurality of complex modules are avoided. In the isolation training stage, the performance of the noise unwrapping module can be quickly optimized; in the fine tuning training stage, the trained noise unwrapping module is utilized to concentrate on optimizing the Conformer-Transducer architecture, so that the training efficiency is improved, and the training time and the consumption of computing resources are reduced. In the isolation training and fine tuning training process, the parameters of part of modules are frozen, the number of parameters needing to be optimized is reduced, and therefore the calculation complexity is reduced. Noise and pure voice in a voice signal are deconstructed through the noise unwrapping module, and accurate semantic understanding is carried out in combination with a Conformer-Transducer architecture, so that the whole voice recognition model has higher robustness to noise.
Owner:SHANGHAI NORMAL UNIVERSITY +1

Method and system for preventing pressing plate of transformer substation control screen cabinet from being touched by mistake

The invention discloses a transformer substation control screen cabinet pressing plate mistaken touch prevention method and system, and the method comprises the steps: the system automatically detects the approaching of the hand of an operator through infrared induction and gesture recognition, and judges whether the recognition operation is effective or not; according to environmental changes such as high humidity or temperature, the system automatically adjusts the sensitivity of the touch screen to optimize the operation experience; before key operation, the system starts a multiple confirmation mechanism, and an operator needs to confirm an operation intention through a touch screen button and voice recognition; for high-risk operation, the system pops up an alarm and carries out secondary confirmation; meanwhile, the system dynamically adjusts the authority according to the identity of the operator and the task requirement, and the low-authority operator is limited to access the key component and only can execute the operation within the authority range of the low-authority operator; according to the method and the system for preventing the pressing plate of the transformer substation control screen cabinet from being touched mistakenly, the accuracy and the flexibility of preventing mistakenly touching are effectively improved, and the defects in the prior art are overcome through an intelligent means.
Owner:GUANGZHOU KAJUN MASCH EQUIP CO LTD

Voice recognition processing method, system and equipment based on conference scene and medium

The invention relates to a voice recognition processing method, system and device based on a conference scene and a medium, and belongs to the technical field of voice processing. The voice recognition processing method comprises the following steps: acquiring an original conference audio stream collected by a microphone array; performing signal preprocessing on the original conference audio stream collected by the main channel, and outputting a pure voice signal; generating a sound source orientation thermodynamic diagram based on the original conference audio stream; extracting multi-dimensional voiceprint feature vectors from the pure voice signals, performing dynamic grouping, outputting a voice fragment set marked with voiceprint IDs, and generating an initial transcription text; dynamically correcting the initial transliteration text, and outputting a transliteration text stream with an industry term tag; and performing periodic memory enhancement processing on the transliteration text stream, outputting and analyzing a long text, and generating structured conference summary data. According to the invention, the automation level and accuracy of conference voice processing can be improved.
Owner:CHINA TRANSPORT INFORMATION TECH GRP CO LTD

Two-process error correction method and device for real-time speech transcription

The invention provides a two-process error correction method and device for real-time speech transcription, and relates to the technical field of speech processing, and the method comprises the steps: extracting the Mel spectrum features of each segment, inputting each Mel spectrum feature into a lightweight end-to-end model, and obtaining a preliminary transcription text; splicing the segments according to a preset number to obtain a plurality of long segments, and inputting each long segment into a speech recognition model to obtain a high-precision transcription text; performing text comparison on the preliminary transcription text and the high-precision transcription text according to the confidence degree set of the preliminary transcription text to obtain all error vocabularies in the preliminary transcription text; and performing corresponding error correction processing on each error vocabulary in the preliminary transcription text according to the type of the error vocabulary and the high-precision transcription text to obtain a final transcription text. According to the method, through a two-process transcription error correction mechanism of the preliminary transcription text and the high-precision transcription text, transcription error accumulation is reduced on the premise that the real-time performance is not affected, and the transcription accuracy in a complex scene is improved.
Owner:NANJING DOLPHIN INTELLIGENT TECH CO LTD

End-side collaborative lightweight voice interaction large model optimization method and system

The invention discloses an end-edge collaborative lightweight voice interaction large model optimization method and system, belongs to the technical field of data mining, data analysis and artificial intelligence, and aims to solve the technical problems that the existing webpage control technology is low in integration level, high in resource demand, insufficient in robustness and limited in adaptability. According to the technical scheme, the method comprises the following steps of: analyzing a natural language instruction: capturing the natural language instruction of a user on end-side equipment by utilizing an ASR model, converting a voice signal into text data, and performing semantic recognition and intention recognition on an edge server through a generative large model so as to generate a structured task instruction; scene vocabulary extraction and ASR model fine tuning: constructing a scene exclusive vocabulary library based on a control scene, performing fine tuning on the pre-trained ASR model, and improving the speech recognition accuracy in a specific scene; semantic understanding is optimized based on the RAG model; optimizing instruction generation based on a cue word technology; performing instruction verification and error correction; performing automatic webpage operation; and performing model optimization and reasoning acceleration.
Owner:INSPUR COMM TECH CO LTD

Estimating the accuracy of automatically transcribed speech with pronunciation impairments

PendingUS20250266036A1SensorsDiagnostic recording/measuringSpeech recordingAutomatic speech
Systems and computer-implemented methods for determining an accuracy of automatically transcribed pathological speech comprise recording speech from a person to obtain an original speech recording; combining a perturbation with the original speech recording to obtain a perturbed speech recording; performing automatic speech recognition, ASR, on the original speech recording to obtain a first transcript; performing automatic speech recognition on the perturbed speech recording to obtain a second transcript; comparing the first transcript with the second transcript to quantify a mismatch between the first transcript and the second transcript.
Owner:F HOFFMANN LA ROCHE INC

Intelligent voice recognition and analysis system based on universal smart phone chip

The invention discloses an intelligent voice recognition and analysis system based on a universal smart phone chip, and the system comprises a data collection module which synchronously captures a voice signal and motion sensor data through a built-in microphone array and a motion sensor, and generates an original multi-mode data package with a time sequence stamp; and the noise reduction processing module is used for receiving the original multi-mode data packet, executing environmental noise spectrum analysis, generating an anti-phase sound wave, performing signal enhancement and outputting a pure voice stream. According to the method, through multi-modal noise separation, nonlinear signal enhancement and hierarchical privacy protection, the contradiction between speech recognition precision and privacy security in a complex environment is solved, and meanwhile, by means of dynamic resource scheduling and a lightweight model, the finite computing power of a mobile phone chip is utilized to the maximum extent in the aspect of speech recognition.
Owner:BEIJING ZHIMAI TECHNOLOGY CO LTD

Grid telephone traffic quality inspection intelligent analysis system and method based on large language model

The invention relates to the technical field of intelligent telephone traffic quality inspection, and discloses a grid telephone traffic quality inspection intelligent analysis system and method based on a large language model, and the system comprises the steps: collecting and obtaining a multi-role call audio signal in real time, and preliminarily carrying out the speaker separation and role marking through a voice recognition module and a voiceprint recognition module; forming a preliminary role recognition result; detecting a suspected role identity error region in combination with multi-dimensional features, and when a detection result meets a preset condition, triggering a dynamic correction mechanism, generating auxiliary judgment information in combination with identity declaration keywords, business term matching and dialogue context logic inference, adopting a multi-dimensional weight decision strategy, and re-correcting a role identity tag, so as to obtain a role identity error region. Updating a role recognition result; and setting an accurate evaluation module of a dynamic correction result, feeding back and adjusting a multi-feature weight and a trigger threshold in real time, and forming an iterative optimization mechanism of dynamic correction. The method has the advantage of improving the dynamic correction capability.
Owner:XIANGYANG POWER SUPPLY COMPANY OF STATE GRID HUBEI ELECTRIC POWER

Robot dialogue intelligent early warning system based on voice outbound

The invention relates to the technical field of voice dialogues, and discloses a robot dialogue intelligent early warning system based on voice outbound, which comprises a voice recognition module for receiving a voice signal in a real-time call initiated by an outbound robot and converting the voice signal into a dialogue text; the intention analysis module is used for extracting semantic features and emotional features from the dialogue text, inputting the semantic features and the emotional features into a pre-trained intention classification model and outputting a user intention label and confidence; the emergency degree evaluation module is used for generating a dialogue emergency level based on the user intention label, the confidence coefficient and a preset emergency keyword library; the early warning execution module is used for triggering real-time early warning operation when the conversation emergency level reaches a preset threshold value; the real-time early warning operation comprises at least one of dynamically adjusting a robot dialogue strategy, generating a manual seat transfer instruction and upgrading a dialogue priority. The method can solve the problems that an existing voice outbound robot cannot accurately judge the user intention, and the conversation emergency degree recognition is insufficient, and improves the user experience and the overall efficiency.
Owner:BEIJING XUNYIN TECH CO LTD

Video fact and viewpoint alignment traceability method

The invention discloses a video fact and viewpoint alignment traceability method, which relates to the technical field of information retrieval and verification, and comprises the following steps of: performing frame analysis on a video by using a computer vision technology, extracting scenes, objects, dynamic characteristics and background information in the video, extracting audios from the video by using a voice recognition technology, converting the audios into text information, and storing the text information in a database; utilizing an event extraction technology to extract event elements from the video; the event elements are identified through scenes, objects, dynamic features and background information, and event facts are obtained; and based on the text information, using a natural language processing technology to calculate semantic similarity between the viewpoints and the event facts in the video, and automatically aligning the viewpoints and the event facts in the video according to the semantic similarity. According to the method, high-dimensional semantic modeling is carried out on the text, the similarity is calculated, the internal relation between viewpoints and facts can be accurately recognized, and therefore automatic matching of the viewpoints and the facts is achieved.
Owner:CHONGQING QINGZHI NET EAGLE TECHNOLOGY CO LTD

Hearing aid intelligent noise reduction and human voice enhancement technology based on electroencephalogram signals

The invention relates to a hearing aid intelligent noise reduction and human voice enhancement system based on electroencephalogram signals, and belongs to the field of biomedical engineering and acoustic signal processing. The system comprises an electroencephalogram signal acquisition module, a multi-channel acoustic sensor array, an embedded neural signal processor, an adaptive beam forming module, a dynamic speech enhancement engine and a dual-mode output device, and constructs electroencephalogram-acoustics joint features by extracting an alpha / theta wave power ratio, a P300 component and auditory cortical Gamma phase synchronism. A deep network is driven to separate target voice, a wave beam direction and a frequency response curve are dynamically adjusted based on neural feedback, a closed-loop calibration unit is innovatively adopted, gain is reversely adjusted according to N1-P2 wave amplitude, heart rate variability and eye movement data are fused to optimize decisions, and when the signal-to-noise ratio is-5dB, the voice recognition rate reaches 89%, the auditory fatigue is reduced by 37%, and the decision conflict rate is smaller than 6%. The defects of attention blind area, noise separation failure and physiological adaptation of a traditional hearing aid are overcome. The system is suitable for the fields of hearing impairment rehabilitation, special communication and intelligent cabins.
Owner:MAXSON GLOBAL GROUP INC

Live broadcast behavior tracking system based on deep learning

The invention relates to the technical field of live broadcast behavior monitoring, in particular to a live broadcast behavior tracking system based on deep learning, which obtains high-quality multi-source information and improves the accuracy of feature analysis by synchronously extracting image and audio data from a live broadcast video stream and combining frame extraction, image enhancement and voice recognition. According to the method, image features are extracted through a pre-trained convolutional neural network, audio features are extracted through a deep learning model, voice transliteration texts are fused, the weight of each modal feature is dynamically adjusted based on an attention mechanism, precise recognition of complex scenes and hidden violation behaviors is achieved, and image camouflage and latent language expression risks are effectively coped with. And the illegal type and confidence are output in real time, once suspected illegal behaviors are detected, alarm, interruption or shielding operation is triggered immediately, and related evidences are uploaded to an auditing database. And efficient, accurate and full-process management and control of the live broadcast violation behaviors are realized.
Owner:GUANGZHOU QUNGE INFORMATION TECHNOLOGY CO LTD

Electronic medical record automatic generation method based on voice recognition

InactiveCN120690205ASpeech recognitionPatient-specific dataMedical recordSpeech segmentation
The invention relates to the technical field of electronic medical record generation, and discloses an electronic medical record automatic generation method based on voice recognition, which comprises the following steps: S1, initializing voice input, distributing a unique voice acquisition identifier for a medical session, and completing identifier generation, input, storage, association and identity verification; s2, voice information intelligent recognition: converting voice into a text by using a voice recognition engine, and ensuring semantic consistency through a context verification unit and a semantic analysis unit; and S3, performing multi-speaker processing based on intelligent interference detection and resolution, positioning an interference time period and an interference source through an interference detection unit, and realizing time period distribution and priority ranking of multi-speaker voices by using a voice segmentation protocol and a linear weighting model. And finally, extracting related information from the text generated by voice conversion, and filling the related information into a medical record template of a hospital. The method improves the efficiency and accuracy of electronic medical record generation, solves the problems of multi-speaker interference and semantic logic, and is suitable for medical informatization scenes.
Owner:THE FIRST AFFILIATED HOSPITAL OF GUANGZHOU MEDICAL UNIV (GUANGZHOU RESPIRATORY CENT)