Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

3338 results about "Subvocal recognition" patented technology

Subvocal recognition (SVR) is the process of taking subvocalization and converting the detected results to a digital output, aural or text-based.

Real-time virtual reality scene system based on natural language description using multimodal artificial intelligence

A real-time system for the multimodal generation of virtual reality scenes based on artificial intelligence for the creation of immersive three-dimensional environments from natural language narratives, consisting of: a speech capture module configured to continuously record a user's spoken narrative via one or more directional microphones, preprocesses the captured signal by noise reduction and temporal alignment, and outputs a digital speech stream; A speech-to-text processing unit that is operationally coupled to the speech capture module and configured for real-time speech recognition using a continuous neural transformer model. The unit is trained to transcribe natural language utterances into structured text data while maintaining contextual continuity throughout the evolving narrative. a semantic interpretation processing unit that is communicatively linked to the speech recognition unit and configured to perform natural language understanding techniques to extract contextual entities, spatial references, temporal relationships, and object attributes from the transcribed narrative; the engine includes a large language model that is fine-tuned for spatial reasoning tasks; a scene graph generation module configured to transform the interpreted semantic data into a structured, hierarchical representation that defines nodes for identified entities and edges for corresponding relationships, with each node associated with metadata describing geometry, position, orientation, texture, and linking attributes between objects; a multimodal image-language model processor coupled with the scene graph generation module, wherein the processor is configured to retrieve, adapt, or synthesize appropriate three-dimensional elements from a pre-trained visual-lexical embedding space and align these elements with their semantic and spatial definitions derived from the scene graph; a scene assembly and rendering controller configured to create a cohesive virtual scene from the aligned assets, perform real-time rendering using a GPU-accelerated ray tracing pipeline, and produce a stereoscopic visual output that corresponds to the evolving narrative; A head-mounted virtual reality visualization device connected to the rendering engine and configured to display the generated immersive environment to the user in real time. The device features motion sensors and inside-out tracking cameras to detect head and body movements, dynamically updating viewing angles and perspective within the rendered scene; and a bidirectional feedback module integrated into the head-mounted device and connected to the semantic interpretation processing unit; the module is configured to interpret corrective commands, gestures, or supplementary comments from the user to refine or modify specific scene elements without interrupting the real-time visualization; The system continuously updates the virtual scene as the narrative develops, ensuring temporal synchronization between speech input and rendered output below a defined latency threshold, thus enabling a natural, dialogic construction of complex three-dimensional virtual environments.
Owner:GOUNDER MOHAN SELLAPPA DR BENGALURU +3

Text prediction-based large-model real-time voice text intention recognition method and system

The invention discloses a large-model real-time voice text intention recognition method and system based on text prediction, and the method comprises the steps: obtaining the real-time voice data of a user, carrying out the real-time voice recognition processing through a streaming voice recognition interface, and obtaining a part of transcriptional text; inputting the partial transcription text into a mask language model for text prediction, and generating a plurality of high-credibility complete sentence candidates; based on the complete sentence candidates, the complete sentence candidates are input into a large language model in parallel for intention recognition, a corresponding intention result is obtained, and a mapping relation between the candidate sentences and the intention recognition result is established; and obtaining a sentence completely expressed by the user, calculating the similarity between the complete actual sentence and a plurality of high-credibility complete sentence candidates through a multi-level text similarity algorithm, selecting the candidate sentence with the highest similarity score, and directly obtaining a corresponding final intention recognition result based on the mapping relationship. The objective of the invention is to solve the technical problem of high response delay of an existing voice intention recognition system.
Owner:BEIJING YULORE INNOVATION TECH

Comprehensive AI-enabled systems for immersive voice, companion, and augmented / virtual reality interaction solutions

A computer-implemented method for operating an artificial intelligence voice agent system includes receiving voice input through communication channels; analyzing converted text through natural language processing (NLP) pipelines implementing intent recognition and sentiment analysis detecting emotional cues using a multimodal large language model (LLM); generating response content using machine learning models trained on domain-specific corpora; converting generated responses to synthetic speech through text-to-speech (TTS) engines; integrating with a customer relationship management (CRM) platforms or an enterprise resource planning (ERP) database; and implementing continuous learning by updating language understanding models using conversation logs, voice recognition parameters based on user feedback, and response generation patterns. One implementation is a computer-implemented system and method that operates a suite of intelligent interactive devices and platforms including an artificial intelligence voice agent, enhanced communication platforms, an intimacy companion system, and augmented / virtual reality eyeglasses. Further, one implementation includes AR / VR eyeglasses that project visual content onto interchangeable lenses or directly onto the user's retina via laser-based retinal projection, provide prescription adjustments, incorporate ear-mounted sensors for monitoring physiological parameters like heart rate, oxygen saturation, and blood pressure, and utilize wireless data transmission, onboard environmental sensing, and remote calibration, all designed to offer dynamically adaptive, secure, and context-aware interactions across communication, personal assistance, health monitoring, and immersive augmented or virtual reality environments.
Owner:TRAN BAO

Intelligent conference summary automatic generation method based on voice recognition and large model

The invention discloses an intelligent conference summary automatic generation method based on voice recognition and a large model. The method comprises the following steps: S1, executing voice activity detection operation on an audio data stream; s2, extracting embedding vectors of continuous and effective voice segments, and generating a voice segment set to which a spokesman belongs; s3, inputting the voice fragment set to which the spokesman belongs into an improved Whisper model, fusing a Speaker-Aware attention mechanism and a connection time sequence classification auxiliary path, and outputting a conference transcription text sequence set; s4, inputting the processed structured dialogue format into a GPT-4 large language model, and generating a conference semantic representation sequence; s5, generating a conference summary first draft text according to a preset summary generation template; and S6, performing formatting output operation on the conference summary first draft text. The conference semantic elements can be automatically extracted, the structured summary text can be generated, and the method is suitable for efficient conference recording and task tracking in government affair office, enterprise collaboration, academic discussion and other scenes.
Owner:JIANGSU GUOHUACHENJIAGANG POWER GENERATION CO LTD

Speech recognition method and related device

ActiveCN114360510AImprove fault tolerancePrecise Syllable Probability DistributionSpeech recognitionSyllableAcoustic model
The embodiment of the invention discloses a speech recognition method and a related device, and at least relates to a speech recognition technology in artificial intelligence, speech data to be recognized are used as input data of a time delay neural network in an acoustic model, and an output layer of the time delay neural network comprises acoustic modeling units corresponding to a plurality of syllables respectively, so that the speech recognition efficiency is improved. And the syllable probability distribution corresponding to the voice frames included in the voice data can be obtained by taking the syllables as the recognition granularity through the time delay neural network. When syllable recognition is carried out through the output layer, auxiliary judgment can be carried out on the syllables to which the voice frames belong on the basis of pronunciation rules in combination with front and back syllable information of the voice frames, so that more accurate syllable probability distribution is output. Moreover, since the syllables are generally composed of one or more phonemes, the method has higher fault-tolerant capability, not only can more accurately determine the speech recognition result based on the probability distribution of the syllables, but also has low requirements for the quality of the speech data to be recognized, and effectively expands the application scenarios of the speech recognition technology.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Electronic medical record LLM generation method based on animal injury

The invention discloses an electronic medical record LLM generation method based on animal injury, which realizes dialogue structuring and timestamp synchronization through multistage speech recognition and role affiliation. Using standardized medical term mapping and coding to align the free text to a standardized medical entity, and constructing a high-confidence medical entity network based on a semantic anchor point pool; according to the method, context-sensitive entity relationship extraction is realized by combining a large language model and a semantic enhancement template, a high-accuracy structured relationship chain is generated through clinical logic rule set verification, and finally, an electronic medical record template under diagnosis and treatment specifications is automatically filled and privacy desensitization processing is completed. The semantic consistency, the structural accuracy and the data security of automatic generation of the electronic medical record are improved, and standardization and intelligent circulation of medical information are effectively promoted.
Owner:GUANGZHOU WUCHUAN ELECTRONIC TECHNOLOGY CO LTD +1

Voice recognition processing method, system and equipment based on conference scene and medium

The invention relates to a voice recognition processing method, system and device based on a conference scene and a medium, and belongs to the technical field of voice processing. The voice recognition processing method comprises the following steps: acquiring an original conference audio stream collected by a microphone array; performing signal preprocessing on the original conference audio stream collected by the main channel, and outputting a pure voice signal; generating a sound source orientation thermodynamic diagram based on the original conference audio stream; extracting multi-dimensional voiceprint feature vectors from the pure voice signals, performing dynamic grouping, outputting a voice fragment set marked with voiceprint IDs, and generating an initial transcription text; dynamically correcting the initial transliteration text, and outputting a transliteration text stream with an industry term tag; and performing periodic memory enhancement processing on the transliteration text stream, outputting and analyzing a long text, and generating structured conference summary data. According to the invention, the automation level and accuracy of conference voice processing can be improved.
Owner:CHINA TRANSPORT INFORMATION TECH GRP CO LTD

End-side collaborative lightweight voice interaction large model optimization method and system

The invention discloses an end-edge collaborative lightweight voice interaction large model optimization method and system, belongs to the technical field of data mining, data analysis and artificial intelligence, and aims to solve the technical problems that the existing webpage control technology is low in integration level, high in resource demand, insufficient in robustness and limited in adaptability. According to the technical scheme, the method comprises the following steps of: analyzing a natural language instruction: capturing the natural language instruction of a user on end-side equipment by utilizing an ASR model, converting a voice signal into text data, and performing semantic recognition and intention recognition on an edge server through a generative large model so as to generate a structured task instruction; scene vocabulary extraction and ASR model fine tuning: constructing a scene exclusive vocabulary library based on a control scene, performing fine tuning on the pre-trained ASR model, and improving the speech recognition accuracy in a specific scene; semantic understanding is optimized based on the RAG model; optimizing instruction generation based on a cue word technology; performing instruction verification and error correction; performing automatic webpage operation; and performing model optimization and reasoning acceleration.
Owner:INSPUR COMM TECH CO LTD

Grid telephone traffic quality inspection intelligent analysis system and method based on large language model

The invention relates to the technical field of intelligent telephone traffic quality inspection, and discloses a grid telephone traffic quality inspection intelligent analysis system and method based on a large language model, and the system comprises the steps: collecting and obtaining a multi-role call audio signal in real time, and preliminarily carrying out the speaker separation and role marking through a voice recognition module and a voiceprint recognition module; forming a preliminary role recognition result; detecting a suspected role identity error region in combination with multi-dimensional features, and when a detection result meets a preset condition, triggering a dynamic correction mechanism, generating auxiliary judgment information in combination with identity declaration keywords, business term matching and dialogue context logic inference, adopting a multi-dimensional weight decision strategy, and re-correcting a role identity tag, so as to obtain a role identity error region. Updating a role recognition result; and setting an accurate evaluation module of a dynamic correction result, feeding back and adjusting a multi-feature weight and a trigger threshold in real time, and forming an iterative optimization mechanism of dynamic correction. The method has the advantage of improving the dynamic correction capability.
Owner:XIANGYANG POWER SUPPLY COMPANY OF STATE GRID HUBEI ELECTRIC POWER

Robot dialogue intelligent early warning system based on voice outbound

The invention relates to the technical field of voice dialogues, and discloses a robot dialogue intelligent early warning system based on voice outbound, which comprises a voice recognition module for receiving a voice signal in a real-time call initiated by an outbound robot and converting the voice signal into a dialogue text; the intention analysis module is used for extracting semantic features and emotional features from the dialogue text, inputting the semantic features and the emotional features into a pre-trained intention classification model and outputting a user intention label and confidence; the emergency degree evaluation module is used for generating a dialogue emergency level based on the user intention label, the confidence coefficient and a preset emergency keyword library; the early warning execution module is used for triggering real-time early warning operation when the conversation emergency level reaches a preset threshold value; the real-time early warning operation comprises at least one of dynamically adjusting a robot dialogue strategy, generating a manual seat transfer instruction and upgrading a dialogue priority. The method can solve the problems that an existing voice outbound robot cannot accurately judge the user intention, and the conversation emergency degree recognition is insufficient, and improves the user experience and the overall efficiency.
Owner:BEIJING XUNYIN TECH CO LTD

Hearing aid intelligent noise reduction and human voice enhancement technology based on electroencephalogram signals

The invention relates to a hearing aid intelligent noise reduction and human voice enhancement system based on electroencephalogram signals, and belongs to the field of biomedical engineering and acoustic signal processing. The system comprises an electroencephalogram signal acquisition module, a multi-channel acoustic sensor array, an embedded neural signal processor, an adaptive beam forming module, a dynamic speech enhancement engine and a dual-mode output device, and constructs electroencephalogram-acoustics joint features by extracting an alpha / theta wave power ratio, a P300 component and auditory cortical Gamma phase synchronism. A deep network is driven to separate target voice, a wave beam direction and a frequency response curve are dynamically adjusted based on neural feedback, a closed-loop calibration unit is innovatively adopted, gain is reversely adjusted according to N1-P2 wave amplitude, heart rate variability and eye movement data are fused to optimize decisions, and when the signal-to-noise ratio is-5dB, the voice recognition rate reaches 89%, the auditory fatigue is reduced by 37%, and the decision conflict rate is smaller than 6%. The defects of attention blind area, noise separation failure and physiological adaptation of a traditional hearing aid are overcome. The system is suitable for the fields of hearing impairment rehabilitation, special communication and intelligent cabins.
Owner:MAXSON GLOBAL GROUP INC

Intelligent accompanying robot system

The invention discloses an intelligent accompanying robot system, and the system comprises the following modules: a multi-mode interaction module which is composed of a voice recognition unit, a voice synthesis unit, an emotion recognition unit, and a visual recognition unit, and is used for collecting the voice, facial expression, motion, and environment information of a user; the localized AI decision module is used for deploying a lightweight DeepSeekR1 large language model based on an ESP32-S3 edge computing chip, performing real-time processing on the multi-modal data and generating a social guidance strategy and an emotion intervention instruction; the data management module comprises a user behavior database and a privacy protection unit and is used for desensitizing the data and realizing sensitive data isolation through local storage; according to the method, the single-person intervention cost is reduced, and the time consumption of manual scene simulation is reduced.
Owner:NANTONG UNIV

Voice interaction system and method for customer service based on artificial intelligence

The invention relates to the technical field of voice recognition, in particular to a voice interaction system and method for customer service based on artificial intelligence, and the system comprises a voice input processing module, an intention classification and routing module, a context dynamic adjustment module, a user behavior learning module, a multi-level intention fusion module and a final result module. According to the method, a multi-dimensional feature system is constructed by extracting tone intensity, speech speed frequency and emotional fluctuation amplitude, intention categories, priority weights and confidence scores are generated to realize accurate acquisition of appeals, and dialogue history, context and emotional change dynamic reconstruction path nodes, switching rules and response time sequences are tracked during interaction. Historical behavior mining preference features, habit fusion intention relevance, emergency calculation of an optimal strategy, construction of service steps, resource allocation schemes and execution timelines, adjustment of an interactive interface, a service process and a feedback mechanism according to multi-dimensional analysis, guarantee of differentiated service experience, and improvement of response accuracy and user satisfaction.
Owner:NANJING XIUGUO INTELLIGENT TECH CO LTD

Control method and artificial intelligence experiment system

The invention provides a control method, which is applied to an artificial intelligence experiment system, at least comprises a robot hardware platform, a controller, a wheeled robot chassis and an industrial mechanical arm, and the method comprises the following steps: starting an operation system and executing hardware self-inspection; loading a large language model service and a traditional AI model, wherein the traditional AI model comprises a speech recognition model, a target detection model and a semantic analysis model; a natural language instruction of a user is received, voice is converted into a structured text through a voice recognition model, an instruction intention is analyzed through a large language model and a semantic analysis model, robot control parameters are generated, and a multi-modal interaction control signal is obtained; based on the multi-mode interaction control signal, a camera is called to collect image data, the position and category of a target object to be grabbed by the robot are recognized through a target detection model, the moving path of the robot and the grabbing track of the mechanical arm are calculated in combination with laser radar data, and the wheeled robot chassis and the industrial mechanical arm are controlled to execute coordinated actions.
Owner:BEIJING ETERNAL CREATIVE TECH CO LTD

Multi-modal real-time compliance auditing method, system and device for investment adviser live broadcast and storage medium

The invention relates to a multi-mode real-time compliance auditing method and system for an investment adviser live broadcast industry. According to the system, audio and video frames and texts in a live broadcast stream are analyzed in real time by fusing artificial intelligence technologies such as audio and video processing, voice recognition, natural language processing and image recognition, and potential illegal contents are recognized. The system adopts a sliding window audio slicing mechanism, so that the accuracy of speech recognition and the continuity of semantic analysis are improved; hidden violation contents can be recognized by combining a semantic recognition technology of BERT model fine tuning. In addition, the system supports automatic marking and backtracking of violation contents, and meets the compliance mark leaving requirements in the financial field. According to the method, the compliance auditing efficiency and accuracy of the live broadcast of the investment adviser are remarkably improved, the deployment cost is reduced, the method is suitable for investment adviser scenes such as securities, funds and insurance, and powerful support is provided for real-time risk control and compliance management.
Owner:WUHAN YOUPIN CHUDING TECH CO LTD

Real-time speech recognition method based on Bluetooth audio stream

The invention relates to the technical field of speech recognition, and discloses a real-time speech recognition method based on a Bluetooth audio stream, which comprises the following steps: analyzing bit allocation parameters to calculate quantized bit distribution and generate a frequency domain confidence mask, monitoring a packet loss concealment state flag bit of a decoder, forcibly setting the mask as a blocking threshold when an algorithm is activated, and generating a real-time speech recognition result. According to the method, a cross-level feature purification mechanism based on protocol priori and link states is constructed, acoustic model illusion caused by forged waveforms is blocked through targeted arbitration while deterministic quantization noise is eliminated, and the acoustic model recognition accuracy is improved. And the identification accuracy under a severe channel is ensured.
Owner:SHENZHEN HUIJIEXIN TECH CO LTD

Text processing method and device, electronic equipment, storage medium and program product

The invention provides a text processing method and device, electronic equipment, a storage medium and a program product, and relates to the technical fields of artificial intelligence, natural language processing, large language models, automatic driving, intelligent traffic and the like. The method comprises the following steps: determining an initial voice recognition text and phoneme data of a target voice, and generating target prompt information based on the initial voice recognition text and the phoneme data; obtaining a target recognition text through a large language model based on the target prompt information; the large language model can correct the initial speech recognition text based on the phoneme data; as the accuracy of the phoneme data is far greater than the accuracy of the initial speech recognition text, more and more effective information can be provided for reference for the processing process of the large language model by combining the phoneme data, the large language model is helped to obtain a correct text in the processing process, and then the accuracy of text processing is improved.
Owner:GUANGZHOU TENCENT TECH CO LTD

Pilot earphone hearing protection method and system based on voice recognition compensation

The invention relates to the field of aviation voice signal processing, and discloses a voice recognition compensation pilot earphone hearing protection method and system, and the method comprises the following steps: collecting multi-modal data, and separating a sound source through tensor decomposition; inferring a pilot state by using a dynamic network; performing context recognition and semantic evaluation on the attention target voice; and finally, dynamically modulating the sound field based on deep reinforcement learning, enhancing the voice in a personalized manner, and outputting after noise suppression. The system comprises a multi-mode perception data acquisition module, a sound source decoupling module, a pilot state inference module, a voice processing and semantic evaluation module and a sound field modulation and output module. According to the invention, high-fidelity speech extraction is realized through multi-modal perception and tensor decomposition; evaluating priority key information in combination with attention and semantics; and deep reinforcement learning and model prediction control are adopted to dynamically optimize the sound field, so that the voice recognition accuracy and the pilot information acquisition efficiency are improved.
Owner:FOURTH MILITARY MEDICAL UNIVERSITY

Multi-stage wake-up method and system based on low-power-consumption wake-up protocol, medium and product

The invention provides a multi-level wake-up method and system based on a low-power-consumption wake-up protocol, a medium and a product, and relates to the technical field of intelligent voice recognition, and the method comprises the steps: triggering a first-level wake-up code broadcast module to broadcast a target wake-up signal through a target low-power-consumption wake-up protocol by using a received external trigger signal; obtaining a validity verification result obtained after the to-be-awakened device performs validity verification on the received target awakening signal, and determining whether the to-be-awakened device enters a low-power-consumption standby state or not according to the validity verification result; triggering a second-level voiceprint recognition module to monitor a target user in a preset distance range under the condition that the to-be-awakened device is determined to enter the low-power-consumption standby state, so as to collect a voice signal in the preset distance range; and obtaining a voiceprint feature matching result obtained after the second-stage voiceprint recognition module performs voiceprint feature verification on the collected voice signal, and judging whether to activate the to-be-awakened device or not according to the voiceprint feature matching result.
Owner:HANGZHOU ROOMBANKER TECH CO LTD

Intelligent education robot question answering system based on voice recognition and knowledge graph

The invention discloses an intelligent education robot question answering system based on voice recognition and a knowledge graph, and particularly relates to the technical field of artificial intelligence. The method comprises the following steps: performing feature extraction on a user voice signal to obtain a text sequence and a multi-modal context feature; recognizing subject domain judgment and question answering intentions based on supervised classification and keyword rule fusion, and outputting a subject domain prior probability and a question answering target vector; determining polysemy words in the text sequence by using a context window, generating a semantic item sequence and semantic item confidence, constructing a subject domain sub-graph based on a subject domain prior probability, and obtaining a candidate reasoning path and a path scoring vector; determining a target teaching concept and an optimal reasoning path by combining Bayesian inference and consistency verification; generating a personalized question answering result aiming at the question asked by the user through fact retrieval and knowledge derivation in combination with the question answering target vector; accurate and efficient intelligent teaching question answering can be realized, and question answering accuracy and intelligent interaction capability of the education robot are effectively improved.
Owner:SHANDONG BAIKU EDUCATION TECH CO LTD

Speech recognition large model training method and device, storage medium and equipment

The invention relates to the technical field of artificial intelligence, and provides a training method and device for a voice recognition large model, a medium and equipment, and the method comprises the steps: obtaining a training data set; inputting the training data set into the initial large model, and performing identification processing on the audio sample through a streaming identification branch in the initial large model to obtain a first candidate text set; according to the text consistency between the target candidate text in the first candidate text set and the real labeled text, determining whether to activate a non-streaming recognition branch to generate a second candidate text set; and performing iterative training on the initial large model according to the semantic difference between each second candidate text and the real labeled text and the text length difference between each second candidate text and the target candidate text to obtain a trained united streaming and non-streaming speech recognition large model. According to the method and the device, the problem of subtitle jumping caused by secondary identification length change can be relieved while the identification accuracy is improved.
Owner:JINGDONG CITY BEIJING DIGITS TECH CO LTD

Intelligent control system for old-age care robot based on Internet of Things

The invention relates to the technical field of intelligent control, in particular to an internet-of-things-based old-age care robot intelligent control system, which comprises a multi-mode sensing layer, a safety decision center, an execution driving module, a voice interaction optimization module and an internet-of-things communication module. According to the method, force sense data such as contact force of the mechanical arm, voice data such as voices of old people and view data such as obstacles are collected in real time through a multi-mode sensing layer, a displacement force coupling risk value is calculated through a hyperbolic tangent function coupling model of a safety decision center, and a deviation angle risk value is analyzed and calculated through an improved Lyapunov index spectrum; a total risk value is generated through fusion and is scheduled according to hierarchical intervals, mechanical operation, old people behaviors and environmental risks can be dynamically avoided, and the use safety is remarkably improved; through sub-ambiguity quantitative analysis, comprehensive ambiguity evaluation and intention correction of the voice interaction optimization module, instruction intention is optimized in combination with multi-modal data, the problem of fuzzy voice recognition of old people is effectively solved, and interaction accuracy and response efficiency are greatly improved.
Owner:WEISEN POWER (WUXI) TECH CO LTD

Voice data recognition method and system based on AI voice algorithm

The invention discloses a voice data recognition method and system based on an AI voice algorithm, relates to the technical field of AI voice recognition, and solves the problem that the voice data recognition capability is low. The method comprises the following steps: S1, multi-mode cooperative triggering collection: synchronously collecting lip electromyographic signals and voiceprint features through a multi-mode sensor, an activation instruction is generated through feature fusion, and voice acquisition starting is triggered; s2, AI adaptive noise reduction processing: carrying out noise separation on the original audio signal by adopting a generative adversarial network, separating environmental noise features to generate a dynamic noise reduction mask, and keeping the integrity of human voice features; s3, beam dynamic optimization adjustment: analyzing real-time audio quality based on a reinforcement learning algorithm, dynamically adjusting beam pointing and gain parameters of a microphone array, and focusing a target sound source; and S4, semantic association cache enhancement: carrying out real-time semantic analysis on the collected voice data. According to the invention, the voice data recognition capability of an AI voice algorithm is greatly improved.
Owner:HUAQIAO UNIVERSITY

Intelligent tactile stick interaction method and system based on multi-modal perception and edge calculation

The invention provides an intelligent tactile stick interaction method and system based on multi-modal perception and edge calculation, and relates to the technical field of intelligent auxiliary equipment. The method comprises the following steps: acquiring environment and position data through a camera, a GPS and an inertial measurement unit; performing multi-modal data fusion and deep learning identification by using an edge calculation module to generate a dynamic environment map; performing real-time path planning and obstacle avoidance based on a map and a destination; analyzing a user instruction through voice recognition and interactively adjusting a path; haptic and auditory bimodal feedback is provided through a vibration module and a voice synthesis module according to a navigation result; and when an emergency obstacle is detected or a user actively alarms, a local and remote alarm mechanism is triggered. The system correspondingly comprises an intelligent tactile stick terminal and a cloud server. According to the invention, accurate and real-time perception and intelligent navigation of the environment are realized, visual and natural interaction experience is provided, an effective emergency help channel is established, and the safety, independence and convenience of travel of visually impaired people are significantly improved.
Owner:GUANGXI NORMAL UNIV

Speech recognition accuracy evaluation statistical method

The invention provides a speech recognition accuracy evaluation statistical method, which comprises the following steps: acquiring current underwater depth, helium-oxygen mixed gas density, original speech waveform and respiratory resistance, analyzing to obtain waveform change amplitude and gas density change amplitude, and recognizing main frequency components in the original speech waveform; combining the corrected voice waveform with the helium-oxygen mixed gas density and the respiratory resistance to evaluate the stability degree of the sound production gas flow, and determining audio sample classification according to the stability degree of the sound production gas flow; and classifying and grouping the audio samples according to a preset depth interval so as to adapt to acoustic environment changes of different depths, and identifying a voice signal frequency characteristic offset rule of different underwater depth groups so as to obtain voice identification standards for different underwater depth levels.
Owner:NANJING DINGSHAN INFORMATION TECH CO LTD

Voice privacy leakage detection and desensitization system and method based on end-to-end model

The invention discloses a voice privacy leakage detection and desensitization system and method based on an end-to-end model, and belongs to the field of voice recognition and privacy protection. The problems that semantic desensitization is not thorough, the voiceprint protection effect is poor, and full-process auditing is lacked are solved. According to the technical principle part, a voice recognition module, a semantic desensitization module, a voiceprint desensitization module and a log auditing module are included, original audio is input, and Chinese voice transcription is completed through a pruned and optimized Whisper-medium model; and the semantic desensitization module and the voiceprint desensitization module regenerate a new audio, the voiceprint similarity is ensured to be less than or equal to 0.3 through verification, and the identity traceability of the speaker is cut off. Log auditing records the processing process in real time, and abnormal behavior early warning is conducted. According to the method, a multi-level privacy leakage detection defense line is formed, abnormal behaviors can be found in time, and user privacy safety is guaranteed.
Owner:HUNAN ZHENTONG ZHIYONG ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD

Medical question-answering system based on time sequence knowledge graph and question-answering method thereof

The invention discloses a medical question-answering system based on a time sequence knowledge graph and a question-answering method thereof, and the system comprises a TKG construction module which is used for constructing the time sequence knowledge graph in the medical field and comprises entities, relationships and timestamp information; the hierarchical graph neural network module fused with position coding comprises a sub-graph layer and a global graph layer, the sub-graph layer is used for capturing a structural dependency relationship of concurrent facts under the same timestamp, and the global graph layer is used for capturing time correlation between cross-timestamp entities; the LLM collaborative reasoning module adopts RAG retrieval and combines the reasoning result of the TKG with an external medical knowledge base to generate an answer; the multi-mode interaction module integrates voice recognition and synthesis and supports voice questions and answers; according to the scheme, the position codes are fused into the message propagation process of the relation perception graph convolutional neural network, the distinguishing capacity of the target node for the neighbor nodes is greatly enhanced, and therefore the expression capacity of embedding of the target node is enhanced.
Owner:CHENGDU UNIV OF INFORMATION TECH

Speech recognition method, device and system, electronic equipment and storage medium

The invention provides a voice recognition method, device and system, electronic equipment and a storage medium, and the method is applied to terminal equipment, and comprises the steps: carrying out the acoustic feature extraction of a voice signal based on the language information of the voice signal, obtaining an acoustic feature, decoding the acoustic feature, and obtaining a plurality of initial recognition results of the voice signal; determining a voice recognition result of the voice signal; the speech recognition result is obtained by applying a speech recognition model to perform semantic error correction on the basis of each initial recognition result, the acoustic feature and the language information; a speech recognition model is constructed on the basis of a large-scale language model, the defects that current multi-language speech recognition is low in accuracy and prone to misjudgment are overcome, through a two-step recognition process, a plurality of initial recognition results are generated locally and rapidly, and then the speech recognition accuracy is improved through the strong semantic understanding ability of the large-scale model. And the acoustic features and the language information are fused to carry out multi-modal deep error correction, so that the accuracy of multi-language speech recognition is greatly improved.
Owner:ANHUI IFLYTEK UNIVERSAL LANGUAGE TECH CO LTD

Full-process automatic voice-driven electronic medical record generation system

The invention discloses a full-process automatic voice-driven electronic medical record generation system, and belongs to the technical field of medical information. The invention aims to solve the problems of insufficient medical term precision, low medical record structuring efficiency, lack of clinical knowledge support, data security risk and the like of voice recognition in the prior art. According to the system, a full-automatic process from voice input to structured medical record output is realized by integrating medical level voice acquisition, term enhanced voice recognition, intelligent medical record generation based on a large language model and a knowledge base, multi-modal interaction and system integration and a full-process safety compliance system adopting edge computing and block chain technologies. According to the method, the medical record writing efficiency can be remarkably improved, manual errors are reduced, medical safety and data privacy are guaranteed, and the method has wide clinical application value.
Owner:CHENGDU ZHIXUEYI DIGITAL TECH CO LTD

Speech recognition model training method and device, equipment and medium

The invention relates to a speech recognition technology, discloses a speech recognition model training method and device, equipment and a medium, and aims to improve the speech recognition rate in a noise environment. The method comprises the following steps: firstly, training an automatic speech recognition network to obtain a pre-trained network; a noise reduction module is introduced in front of an output classification layer, the noise reduction module is trained by taking embedded features output by the pre-trained automatic speech recognition network as a reference, and parameters of the pre-trained automatic speech recognition network are fixed during training to form an initial speech recognition model; and finally, noise reduction module parameters are fixed, the pre-trained automatic speech recognition network is retrained, and a final model is obtained. Through staged training and a parameter fixing strategy, training target conflicts among modules are avoided, and the training stability and the convergence speed are improved; the noise reduction module focuses on feature denoising required by recognition, is high in adaptability, has the characteristics of light weight and low delay, and can be widely applied to low-resource real-time scenes such as embedded equipment and edge computing.
Owner:WUXUE GUANGJI DATA TECHNOLOGY CO LTD