Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1468 results about "Speech input" patented technology

Speech input is one of the most innovative browser technologies to appear in recent months. It’s easy to implement and there are several obvious uses: assistive dictation for those with impaired mobility. an alternative input option for mobile phones and tablets, and. any environment where a keyboard or mouse is impractical.

Scene interactive AI rehabilitation assessment training and health monitoring system

The invention discloses a scene interactive AI rehabilitation evaluation training and health monitoring system, and relates to the technical field of rehabilitation medical treatment and artificial intelligence, a semantic perception module is used for collecting and recognizing voice input, facial expressions, action tracks and eye movement paths of a user in a training process, and extracting context parameters; the knowledge-driven training generation module is used for calling a rehabilitation knowledge graph constructed by a graph neural network based on context parameters and individual training history, and generating a multi-path training scheme; training a feedback regulation engine, collecting posture offset, physiological stress and emotion feedback, and dynamically adjusting task difficulty, rhythm and prompt mode based on a dual-channel reinforcement learning model; the prediction module fuses training and monitoring data, and predicts a network identification function degradation risk through degradation driving; the cloud edge fusion platform is used for realizing task quick response and graph strategy iterative updating; according to the invention, the individuation, self-adaption and intelligent prediction capabilities of rehabilitation training are improved, and the rehabilitation effect and the system practicability are obviously optimized.
Owner:WEIFANG MEDICAL UNIV

Real-time virtual reality scene system based on natural language description using multimodal artificial intelligence

A real-time system for the multimodal generation of virtual reality scenes based on artificial intelligence for the creation of immersive three-dimensional environments from natural language narratives, consisting of: a speech capture module configured to continuously record a user's spoken narrative via one or more directional microphones, preprocesses the captured signal by noise reduction and temporal alignment, and outputs a digital speech stream; A speech-to-text processing unit that is operationally coupled to the speech capture module and configured for real-time speech recognition using a continuous neural transformer model. The unit is trained to transcribe natural language utterances into structured text data while maintaining contextual continuity throughout the evolving narrative. a semantic interpretation processing unit that is communicatively linked to the speech recognition unit and configured to perform natural language understanding techniques to extract contextual entities, spatial references, temporal relationships, and object attributes from the transcribed narrative; the engine includes a large language model that is fine-tuned for spatial reasoning tasks; a scene graph generation module configured to transform the interpreted semantic data into a structured, hierarchical representation that defines nodes for identified entities and edges for corresponding relationships, with each node associated with metadata describing geometry, position, orientation, texture, and linking attributes between objects; a multimodal image-language model processor coupled with the scene graph generation module, wherein the processor is configured to retrieve, adapt, or synthesize appropriate three-dimensional elements from a pre-trained visual-lexical embedding space and align these elements with their semantic and spatial definitions derived from the scene graph; a scene assembly and rendering controller configured to create a cohesive virtual scene from the aligned assets, perform real-time rendering using a GPU-accelerated ray tracing pipeline, and produce a stereoscopic visual output that corresponds to the evolving narrative; A head-mounted virtual reality visualization device connected to the rendering engine and configured to display the generated immersive environment to the user in real time. The device features motion sensors and inside-out tracking cameras to detect head and body movements, dynamically updating viewing angles and perspective within the rendered scene; and a bidirectional feedback module integrated into the head-mounted device and connected to the semantic interpretation processing unit; the module is configured to interpret corrective commands, gestures, or supplementary comments from the user to refine or modify specific scene elements without interrupting the real-time visualization; The system continuously updates the virtual scene as the narrative develops, ensuring temporal synchronization between speech input and rendered output below a defined latency threshold, thus enabling a natural, dialogic construction of complex three-dimensional virtual environments.
Owner:GOUNDER MOHAN SELLAPPA DR BENGALURU +3

Comprehensive AI-enabled systems for immersive voice, companion, and augmented / virtual reality interaction solutions

A computer-implemented method for operating an artificial intelligence voice agent system includes receiving voice input through communication channels; analyzing converted text through natural language processing (NLP) pipelines implementing intent recognition and sentiment analysis detecting emotional cues using a multimodal large language model (LLM); generating response content using machine learning models trained on domain-specific corpora; converting generated responses to synthetic speech through text-to-speech (TTS) engines; integrating with a customer relationship management (CRM) platforms or an enterprise resource planning (ERP) database; and implementing continuous learning by updating language understanding models using conversation logs, voice recognition parameters based on user feedback, and response generation patterns. One implementation is a computer-implemented system and method that operates a suite of intelligent interactive devices and platforms including an artificial intelligence voice agent, enhanced communication platforms, an intimacy companion system, and augmented / virtual reality eyeglasses. Further, one implementation includes AR / VR eyeglasses that project visual content onto interchangeable lenses or directly onto the user's retina via laser-based retinal projection, provide prescription adjustments, incorporate ear-mounted sensors for monitoring physiological parameters like heart rate, oxygen saturation, and blood pressure, and utilize wireless data transmission, onboard environmental sensing, and remote calibration, all designed to offer dynamically adaptive, secure, and context-aware interactions across communication, personal assistance, health monitoring, and immersive augmented or virtual reality environments.
Owner:TRAN BAO

Intelligent safety protection management method and system

The invention relates to an intelligent safety protection management method and system. The method comprises the following steps: acquiring behavior data and position data of personnel in an industrial site; analyzing the behavior data and the position data by using a pre-trained personnel behavior model to obtain a behavior recognition result; according to the behavior recognition result and the current operation state of the target equipment, judging a risk level corresponding to the personnel behavior; based on the risk level, a corresponding safety response strategy is matched in the edge computing node, a control instruction is generated according to the safety response strategy, and the control instruction is used for driving the target device to execute a corresponding response action; and the voice interaction module is used for acquiring voice input of an on-site operator, performing semantic recognition on the voice input to obtain a voice recognition result, and executing corresponding emergency control operation to cover a current control instruction when the voice recognition result meets a preset emergency instruction condition. The method has the effect of improving the accuracy of intelligent safety management in the industrial environment.
Owner:SHENZHEN HUAYIXIN ELECTRONICS CO LTD

Intelligent dialogue method and device combining voiceprint recognition and voice synthesis

The invention relates to the technical field of voiceprint recognition, and particularly provides a voiceprint recognition and voice synthesis combined intelligent dialogue method, which comprises the following steps: acquiring a user voice input signal in real time; voiceprint feature extraction is carried out on the voice input signal, and a composite voiceprint vector containing user identity features and rhythm features is obtained; user identity matching is carried out based on the composite voiceprint vector, and a dynamic user portrait database is associated and called; generating personalized semantic feedback content according to the weight-updated dynamic parameters in the dynamic user portrait database; generating a target timbre synthesis parameter set based on the timbre representation vector in the composite voiceprint vector in combination with the spatio-temporal characteristic parameters of the interaction scene; and performing voice synthesis on the personalized semantic feedback content by adopting the target timbre synthesis parameter set, adjusting rhythm expression parameters of the synthesized voice based on real-time rhythm characteristics in the composite voiceprint vector, and finally outputting personalized voice, thereby realizing natural interaction of one person, one voice and one strategy.
Owner:MINAMI ACOUSTICS LTD

System and method for multi-modal ai conversational interface improving website navigation and user interaction

The present invention relates to a system for transforming static websites into artificial intelligence (AI)-enabled interactive multi-modal conversational platforms. The system comprises a computing device having a processor for receiving user queries as text or speech input through an input module cooperating with a speech-to-text module. A natural language processing (NLP) module interprets intent, classifies user context, and retrieves grounded information from multiple webpages. A persona adaptation module dynamically modifies vocabulary, tone, and avatar representation across roles such as sales assistant, recruiter, educator, healthcare professional, etc. A response generator module produces structured natural language output, transmitted to a text-to-speech synthesis module and an avatar generation module to render synchronized lifelike video responses. An output rendering module displays multi-modal responses include text, audio, and video, thereby enabling direct navigation and escalation beyond limitations of conventional static websites.
Owner:NALLAM SREE RAMA CHANDRA MURTY

Lightweight intelligent traditional Chinese medicine inquiry system and construction method thereof

The invention relates to the field of artificial intelligence medical application, and discloses a lightweight intelligent traditional Chinese medicine inquiry system and a construction method thereof, and the system comprises a multi-dialect adaptive speech recognition module, a traditional Chinese medicine intelligent dialogue large language model module, a natural speech synthesis module, and a continuous learning mechanism module. The multi-dialect adaptive speech recognition module is used for converting dialect speech input of a patient into a standard text; the traditional Chinese medicine intelligent dialogue big language model module is the core of the system and is used for carrying out natural language understanding, dialectical reasoning and inquiry dialogue generation, and the natural speech synthesis module is used for converting a text response generated by the system into speech output; and the continuous learning mechanism module realizes continuous optimization of the large language model through incremental learning architecture and clinical feedback integration. According to the method, while the professional traditional Chinese medicine diagnosis capability is maintained, the calculation complexity is remarkably reduced, and the universality and sustainable development capability of system application are improved.
Owner:SUZHOU ANGSHENG NETWORK TECHNOLOGY CO LTD

Large model-based automobile live risk assessment system and method

The invention belongs to the field of automobile evaluation, and provides an automobile live risk evaluation system and method based on a large model, and the method comprises the steps: receiving voice input information of an evaluator in a vehicle on-site inspection process; extracting component state information of the target vehicle based on the structured check item text; acquiring historical maintenance record data of the target vehicle, and performing association fusion on the historical maintenance record data and the component state information; based on a large language model, performing joint reasoning on a currently input inspection label and the historical maintenance record, constructing a causal chain related to an accident risk, and generating at least one subsequent to-be-inspected item; generating voice guide content and feeding back the voice guide content to the evaluator so as to prompt the evaluator to continue to check the target item of the target vehicle; gradually perfecting the risk check path of the target vehicle until the risk chain closed loop is completed; and calling an estimation formula to calculate a residual estimation value of the target vehicle, and outputting a structured risk assessment report.
Owner:GUANGZHOU SUISHENG INFORMATION TECH CO LTD

Enterprise intelligent finance and tax system fusing supervised learning and block chain

The invention provides an enterprise intelligent finance and taxation system fusing supervised learning and a block chain, and solves the core technical problems of insufficient data credibility, transparent contradiction between privacy protection and supervision, AI model training data island and the like of a traditional finance and taxation system. According to the system, an innovative technical fusion scheme is adopted, bank API, OCR recognition, voice input and other multi-source financial data are integrated through a multi-modal data fusion unit, and high-quality data fusion is achieved through confidence evaluation and an intelligent conflict resolution algorithm; an intelligent classification unit based on BERT deep learning integrates a pre-training model with a business rule engine and a gradient boosting decision tree, and accurate and automatic classification of financial transactions is achieved. A complete technical solution is provided for enterprise finance and taxation digital transformation, intelligent, automatic and credible processing of finance and taxation businesses is achieved on the premise that data safety and privacy protection are guaranteed, and the method has wide market application prospects.
Owner:张宏

Electronic medical record automatic generation method based on voice recognition

InactiveCN120690205ASpeech recognitionPatient-specific dataMedical recordSpeech segmentation
The invention relates to the technical field of electronic medical record generation, and discloses an electronic medical record automatic generation method based on voice recognition, which comprises the following steps: S1, initializing voice input, distributing a unique voice acquisition identifier for a medical session, and completing identifier generation, input, storage, association and identity verification; s2, voice information intelligent recognition: converting voice into a text by using a voice recognition engine, and ensuring semantic consistency through a context verification unit and a semantic analysis unit; and S3, performing multi-speaker processing based on intelligent interference detection and resolution, positioning an interference time period and an interference source through an interference detection unit, and realizing time period distribution and priority ranking of multi-speaker voices by using a voice segmentation protocol and a linear weighting model. And finally, extracting related information from the text generated by voice conversion, and filling the related information into a medical record template of a hospital. The method improves the efficiency and accuracy of electronic medical record generation, solves the problems of multi-speaker interference and semantic logic, and is suitable for medical informatization scenes.
Owner:THE FIRST AFFILIATED HOSPITAL OF GUANGZHOU MEDICAL UNIV (GUANGZHOU RESPIRATORY CENT)

Voice interaction system and method for customer service based on artificial intelligence

The invention relates to the technical field of voice recognition, in particular to a voice interaction system and method for customer service based on artificial intelligence, and the system comprises a voice input processing module, an intention classification and routing module, a context dynamic adjustment module, a user behavior learning module, a multi-level intention fusion module and a final result module. According to the method, a multi-dimensional feature system is constructed by extracting tone intensity, speech speed frequency and emotional fluctuation amplitude, intention categories, priority weights and confidence scores are generated to realize accurate acquisition of appeals, and dialogue history, context and emotional change dynamic reconstruction path nodes, switching rules and response time sequences are tracked during interaction. Historical behavior mining preference features, habit fusion intention relevance, emergency calculation of an optimal strategy, construction of service steps, resource allocation schemes and execution timelines, adjustment of an interactive interface, a service process and a feedback mechanism according to multi-dimensional analysis, guarantee of differentiated service experience, and improvement of response accuracy and user satisfaction.
Owner:NANJING XIUGUO INTELLIGENT TECH CO LTD

Wearable electronic devices, extended reality systems including neuromuscular sensors, and methods for generating text from speech input and modifying the generated text based on neuromuscular data

The disclosed system for interacting with objects in an extended reality (XR) environment generated by an XR system may include (1) neuromuscular sensors configured to sense neuromuscular signals from a wrist of a user and (2) at least one computer processor programmed to (a) determine, based at least in part on the sensed neuromuscular signals, information relating to an interaction of the user with an object in the XR environment and (b) instruct the XR system to, based on the determined information relating to the interaction of the user with the object, augment the interaction of the user with the object in the XR environment. Other embodiments of this aspect include corresponding apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods.
Owner:META PLATFORMS TECHNOLOGIES LLC

Method for processing cross-modal question answerning based on large model, apparatus and storage medium

A method for processing cross-modal question answering based on large model, an apparatus, and a storage medium are suggested, which relates to the field of artificial intelligence technologies such as speech interaction processing, large models, machine learning and natural language processing. The specific implementation includes: performing an activity detection on a target speech input by a user; in response to detecting a pause in the inputting of the target speech, obtaining a first text corresponding to a first input speech before the moment of the pause in the target speech; performing a text response processing using a pre-trained speech question answering processing system based on the first text and the first input speech.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Autonomous traveling robot operation system, autonomous traveling robot operation method, and non-transitory computer readable medium

An autonomous traveling robot operation system that improves the ease with which an autonomous traveling robot is operated by using a voice input and a handwritten input is provided. An autonomous traveling robot operation system including: an autonomous traveling robot configured to shoot an environment near the autonomous traveling robot and operates an object to be operated; a handwritten input interface configured to display an image shot by the autonomous traveling robot and receive a handwritten input to the displayed image; and a voice input interface configured to receive a voice input to the object to be operated, in which the autonomous traveling robot operation system operates the autonomous traveling robot in such a way that the autonomous traveling robot operates the object to be operated in accordance with instructions of the handwritten input and the voice input is provided.
Owner:TOYOTA JIDOSHA KK +1

Personal protective equipment for navigation and map generation within a visually obscured environment

A system includes a personal protective equipment (PPE) configured to be worn by an agent. The PPE includes a sensor assembly comprising a radar device configured to generate radar data, a microphone configured to capture speech input from the agent, and an inertial measurement device configured to generate inertial data. The system includes a computing device configured to: process sensor data from the sensor assembly, the sensor data including at least the radar data and the inertial data, to generate pose data of the agent based on the processed sensor data, the pose data including a location and an orientation of the agent as a function of time, to process the speech input to identify an item of interest in an environment in which the PPE is deployed, and to form mapping information for the environment with the item of interest being marked, based on the processed speech input.
Owner:3M INNOVATIVE PROPERTIES CO

Voice interaction method and system based on quantum heuristic algorithm, and computer equipment

The invention relates to an artificial intelligence technology, and discloses a voice interaction method and system based on a quantum heuristic algorithm, and computer equipment, and the method comprises the steps: carrying out the preprocessing of a voice signal after receiving voice input, and extracting voice features; taking different combinations of the voice features as superposition of quantum states, dynamically adjusting the selection probability of the voice features based on quantum revolving door operation, and iterating and screening feature subsets with the probability meeting a preset requirement; inputting the screened feature subset into a character recognition model for training, and generating a character recognition result; performing semantic understanding on a character recognition result based on a large language model; after the text information is obtained, extracting voice feature parameters based on the text information; the voice feature parameters are optimized and adjusted through quantum evolution operation; and based on the optimized voice feature parameters, performing voice output by adopting a voice synthesis algorithm. The invention further discloses a computer readable storage medium. The invention aims to improve the efficiency and accuracy of voice interaction.
Owner:深圳玄源科技有限公司

Dialect speech recognition and conversion method and device

The invention provides a dialect speech recognition and conversion method and device. The method comprises the steps of obtaining dialect speech input data, extracting local features and a global dependency relationship by using a Conformer model, and generating an audio feature sequence. The sequence is input into a shared GRU encoder to generate a hidden state sequence, and the hidden state sequence is transmitted to a CTC decoder of a dialect text and a mandarin text in parallel. And constructing a multi-task learning framework to associate the components, and controlling parameter updating of the components. Through the framework, dialect features are efficiently extracted, and dialect and mandarin texts are generated in parallel. According to the dialect speech recognition and conversion method, the advantages of Conform and CTC-GRU models are combined, and dialect speech recognition and conversion with high accuracy, strong generalization and robustness are realized.
Owner:XINXING JIHUA TECHNOLOGY (TIANJIN) CO LTD

Driving route identification system and method, electronic equipment and computer readable medium

The invention provides a driving route recognition system and method, electronic equipment and a readable medium, and belongs to the field of intelligent driving. According to the processed original image and preset navigation information, navigation and road condition information guidance are provided for a driver or intelligent driving is provided for the driver, and according to the intelligent driving technology provided by the patent, an intelligent driving auxiliary system with the core target of solving driver navigation confusion under complex road conditions is constructed. Through deep fusion of high-precision visual perception and map navigation information, strong road identification judgment logic is utilized, when a driver actively requests or a system pre-judges needs, an intelligent driving module provides accurate vehicle control assistance, and clear navigation prompts are combined to jointly ensure that a vehicle safely and correctly passes through a complex road scene, so that the driving safety of the vehicle is improved. And the driving convenience and safety are obviously improved. Voice input is a key link of man-machine cooperation.
Owner:DONGFENG MOTOR GRP

Full-process automatic voice-driven electronic medical record generation system

The invention discloses a full-process automatic voice-driven electronic medical record generation system, and belongs to the technical field of medical information. The invention aims to solve the problems of insufficient medical term precision, low medical record structuring efficiency, lack of clinical knowledge support, data security risk and the like of voice recognition in the prior art. According to the system, a full-automatic process from voice input to structured medical record output is realized by integrating medical level voice acquisition, term enhanced voice recognition, intelligent medical record generation based on a large language model and a knowledge base, multi-modal interaction and system integration and a full-process safety compliance system adopting edge computing and block chain technologies. According to the method, the medical record writing efficiency can be remarkably improved, manual errors are reduced, medical safety and data privacy are guaranteed, and the method has wide clinical application value.
Owner:CHENGDU ZHIXUEYI DIGITAL TECH CO LTD

Method and system for providing assistance for cognitively impaired users by utilizing artificial intelligence

ActiveUS20250342833A1Natural language translationSpeech recognitionCognitively impairedEngineering
In an embodiment, the disclosure relates to a device for assisting a respondent in a conversation. The device includes a microphone configured to detect a voice input, and a transmitter communicatively coupled to a server and configured to transmit the voice input to the server. The server is to generate vectors associated with the voice input, feed the vectors associated with the voice input to an Artificial Intelligence utilizing a trained Machine Learning (ML) model, and obtain, from the trained ML model, an output corresponding to the vectors. The device further includes a receiver communicatively coupled to the server, and configured to receive from the server, the output generated by the ML model. A speaker is communicatively coupled with the receiver and is configured to generate a voice-based response based on the output, for assisting the respondent in responding to the conversation.
Owner:HORIZON IP TECH LLC

Intelligent voice interaction system and method based on streaming multi-mode fusion and equipment control protocol

PendingCN121260156ASpeech recognitionSpeech synthesisSpeech comprehensionEngineering
The embodiment of the invention discloses an intelligent voice interaction system and method based on streaming multi-mode fusion and an equipment control protocol, the system comprises a voice input processing module, a voice understanding and generating module and a voice synthesis module, the voice input processing module is used for converting an audio signal into a first token sequence, and the first token sequence is used for converting the audio signal into a second token sequence; the voice understanding and generating module is used for determining a response token sequence according to the first token sequence on the basis of a multi-modal Transform architecture so as to realize voice understanding and generation; and the voice synthesis module is used for synthesizing the response token sequence into an output audio so as to carry out at least one of the following adjustments on the converted audio of the response token sequence: emotion parameter adjustment, tone adjustment and rhythm adjustment. By adopting the embodiment of the invention, low-delay and high-naturalness intelligent voice interaction can be realized, multi-modal fusion and equipment control are supported, and the user experience is remarkably improved.
Owner:SHENZHEN HUANZHI TECHNOLOGY CO LTD

Information acquisition system and method suitable for visually impaired people

The invention discloses an information acquisition system and method suitable for visually impaired people. The system comprises a hardware layer, a driving layer, an operating system layer, a middleware layer and an application layer which are sequentially in communication connection through interfaces. The application layer comprises a voice interaction module, a large model integration module and a braille conversion module; the hardware layer is used for collecting a voice input instruction and outputting a braille character sequence and a voice output result; the voice interaction module is used for converting a voice input instruction into a text input instruction and converting a text output result into a voice output result; the large model integration module is used for analyzing the text input instruction and generating a text output result; and the braille conversion module is used for converting the text output result into a braille character sequence, so that a hardware layer outputs the braille character sequence based on a timestamp alignment technology. Therefore, based on powerful semantic understanding and information retrieval of the large model integration module, accurate information acquisition is realized, and a five-layer decoupling architecture is adopted, so that the convenience of later system maintenance is improved.
Owner:ZHEJIANG LAB

Asynchronous anthropomorphic communication method and system based on voice transfer

The invention provides an asynchronous anthropomorphic communication method and system based on voice transfer, which are suitable for various terminals such as wearable equipment, earphones, dolls, bolsters and the like. The method comprises the steps of user voice input, edge or cloud recognition, playback confirmation, content translation and voice synthesis, asynchronous transmission and broadcast and the like. The system does not depend on a specific hardware form, emphasizes a user confirmation mechanism and anthropomorphic voice broadcast and supports multilingual translation and personalized voice styles, the communication process is bound based on a device ID or nickname identity, social account login is not needed, and interaction privacy security is guaranteed. The method is widely applicable to various asynchronous social application scenes such as children, old people, lovers, autism rehabilitation and the like. The system supports nickname binding and friend relationship establishment, users can complete social connection through voice instructions or two-dimensional codes, and controllability and interestingness of communication interaction are enhanced.
Owner:GUANGDONG OPERATOR WIRE INTELLIGENT TECHNOLOGY CO LTD

Suzhou dialect medical voice electronic medical record conversion system and method

The invention provides a Suzhou dialect medical voice electronic medical record conversion system and method, and the system comprises an acoustic feature extraction module which is used for extracting the acoustic features of an input voice signal; the dialect tone recognition module is used for recognizing a multi-tone system of Suzhou dialects; the voiced sound processing module is used for detecting voiced sound initial consonants in Suzhou dialects and performing acoustic feature mapping; the medical term mapping module comprises a corresponding relation library of dialect medical vocabularies and standard medical terms and a context-based ambiguity resolution unit; the speech recognition engine comprises an acoustic model, a pronunciation dictionary and a language model; and the medical record generation module is used for converting the identification result into a structured electronic medical record. According to the method, a sliding window processing strategy is adopted, continuous voice input of a doctor can be effectively processed, a long-time voice input scene is supported, and various requirements in actual clinical application are met.
Owner:NANJING WANGSHI INTELLIGENT TECHNOLOGY CO LTD

Speaker authentication and deep forgery detection cascade voice privacy protection device

The invention relates to a speaker authentication and deep forgery detection cascade voice privacy protection device which comprises a microphone, an ASV module, an ADD module, a control unit and an audio output module, the microphone, the ASV module, the ADD module, the control unit and the audio output module are connected in sequence, the input end of the control unit is connected with the output end of the ASV module, and the output end of the ADD module is connected with the output end of the audio output module. A voice input signal is firstly collected by a microphone and then is input into an ASV module for identity verification, and meanwhile voice data is input into an ADD module for forgery detection; the outputs of the two are transmitted to the control unit, and the control unit determines whether to allow the audio signal to be output according to preset logic; according to the invention, the ASV and ADD modules are cascaded, and dual judgment of identity and authenticity is realized on a physical equipment level for the first time; the control unit performs joint decision through a signal fusion strategy, so that the safety is improved; and the structure is clear, the function partition is reasonable, and modular deployment and replacement can be realized.
Owner:FUDAN UNIVERSITY

Real Time Dynamic Classification and Orchestration of Test Automated Components Leveraging Supervised Learning and Multi-Modal AI

This invention relates to systems and methods for real-time dynamic classification and orchestration of test automation components in distributed DevOps environments. The system features an Auto Identify Automation (AIA) engine that leverages supervised learning, Multi-Modal Artificial Intelligence (AI), and Generative AI technologies. It includes a Smart Scenario Designer interface that allows users to author test scenarios using handwriting and voice inputs, which are processed in real-time by AI-driven handwriting recognition, voice recognition, and Natural Language Processing (NLP). The system dynamically suggests relevant automated components via a smart bubble pane, facilitating rapid scenario creation. The architecture is tool-agnostic and scalable, with a Shared Workbench Engine that supports real-time collaboration and conflict resolution. The system continuously adapts and improves, ensuring that the automation suite remains consistent, up-to-date, and aligned with evolving software requirements, enabling efficient and user-friendly management of complex test automation processes.
Owner:BANK OF AMERICA CORP

Russian pronunciation error correction system based on AI speech recognition

The invention relates to the technical field of Russian pronunciation teaching, and discloses a Russian pronunciation error correction system based on AI speech recognition, and the system comprises a speech input module which is used for collecting a Russian speech signal of a user; the rule engine module is internally provided with a Russian linguistics rule library, comprises an accent shift rule library and a vowel weakening acoustic threshold library, and is used for detecting pronunciation errors based on rules; the deep learning evaluation module comprises a neural network model subjected to adversarial data enhancement training, and is used for extracting phoneme-level and rhythm-level features from the voice of the user and detecting errors; a confrontation data generation module; a display module; and an AI interactive teaching module. According to the method, Russian speech data with preset phonemes and rhythm errors are synthesized through the generative adversarial network, the generator directionally injects error features through time-frequency convolution, the discriminator and the evaluation model share a feature extraction layer, and the detection sensitivity of the model to complex pronunciation errors such as native language migration consonant confusion is improved.
Owner:YELLOW RIVER CONSERVANCY TECHN INST

Service control method based on voice analysis and large language model

The invention discloses a service control method based on voice analysis and a large language model, relates to the technical field of voice interaction, and solves the technical problems that in the prior art, an enterprise service system needs a user to manually fill in form fields item by item, so that operation is tedious and low in efficiency, and a process is interrupted due to lack of an intelligent completion mechanism due to lack of necessary parameters. According to the invention, a full-closed-loop control chain from the voice instruction to the service operation is constructed, the intention of the user voice input is analyzed through the fine-tuning large language model, and the service interface is dynamically matched, so that the system function operation is directly driven by the voice, and a two-way cooperation mechanism of voice input-system operation-result feedback is formed; and meanwhile, a dynamic parameter completion technology is adopted, the interface field needing to be filled is automatically verified in the parameter extraction stage, missing parameters are completed through context completion or voice guidance, the operation bottleneck of traditional form item-by-item filling is eliminated, missing of the field needing to be filled is avoided, and the service execution efficiency and the process robustness are remarkably improved.
Owner:CEEC ANHUI ELECTRICAL POWER CONSTR NO 1 CO

Dynamic intention chain modeling and causal cleaning method and system under multi-round dialogue scene

The invention relates to the technical field of natural language processing, in particular to a dynamic intention chain modeling and causal cleaning method and system under a multi-round dialogue scene. According to the method, through multi-round dialogue data collection and preprocessing, one-hot codes are allocated to each sub-word or word group, intention recognition and sequence modeling are carried out based on the codes, and the intention of the user can be accurately recognized; quick mapping for directly obtaining user intentions from voice input data by skipping text data conversion is established, so that the intention recognition efficiency is improved; according to the method, the causal relationship graph is constructed based on the user intention sequence, the Bayesian network is utilized to perform causal modeling, the causal relationship between intentions is derived, and the probability of reasoning the next intention through the previous intention is calculated, so that the next intention of the user and the interactive operation to be executed are accurately reasoned; and collecting user feedback data to calculate a satisfaction score, and correcting the priority ranking rule of the rule engine based on the score.
Owner:CHENGDU YUNDING INTELLIGENT CONTROL TECH CO LTD

Intelligent agent-based spoken dialogue data processing method

ActiveCN120726996ASpeech recognitionEngineeringSpeech rhythm
The invention discloses a spoken dialogue data processing method based on an agent, and belongs to the technical field of artificial intelligence big data processing, and the method comprises the steps: receiving a voice input file sent by a user side, converting a voice signal into text data, extracting an emotion feature, generating an emotion tag and a voice rhythm feature, text and emotion information are combined to form dialogue anchor point information, a path diagram is constructed, and edge weights are calculated; constructing a propagation matrix according to the similarity between a current dialogue anchor point and a historical path node, performing intelligent path reasoning, calculating a path score, ensuring the influence of emotional fluctuation on task path selection, selecting an optimal path by an intelligent agent according to the path score to perform subsequent task processing, and performing real-time correction after user feedback; anchor point information of the current dialogue is adjusted in real time through the sentiment analysis technology, the reasoning process is adjusted according to feedback, and a closed-loop self-adaptive adjustment mechanism is formed, so that the response quality of the system is optimized, and the trust of a user to the system is enhanced.
Owner:SHANDONG LINGCHAO SOFTWARE TECH CO LTD