Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1942 results about "Speech input" patented technology

Speech input is one of the most innovative browser technologies to appear in recent months. It’s easy to implement and there are several obvious uses: assistive dictation for those with impaired mobility. an alternative input option for mobile phones and tablets, and. any environment where a keyboard or mouse is impractical.

Scene interactive AI rehabilitation assessment training and health monitoring system

The invention discloses a scene interactive AI rehabilitation evaluation training and health monitoring system, and relates to the technical field of rehabilitation medical treatment and artificial intelligence, a semantic perception module is used for collecting and recognizing voice input, facial expressions, action tracks and eye movement paths of a user in a training process, and extracting context parameters; the knowledge-driven training generation module is used for calling a rehabilitation knowledge graph constructed by a graph neural network based on context parameters and individual training history, and generating a multi-path training scheme; training a feedback regulation engine, collecting posture offset, physiological stress and emotion feedback, and dynamically adjusting task difficulty, rhythm and prompt mode based on a dual-channel reinforcement learning model; the prediction module fuses training and monitoring data, and predicts a network identification function degradation risk through degradation driving; the cloud edge fusion platform is used for realizing task quick response and graph strategy iterative updating; according to the invention, the individuation, self-adaption and intelligent prediction capabilities of rehabilitation training are improved, and the rehabilitation effect and the system practicability are obviously optimized.
Owner:WEIFANG MEDICAL UNIV

Real-time virtual reality scene system based on natural language description using multimodal artificial intelligence

A real-time system for the multimodal generation of virtual reality scenes based on artificial intelligence for the creation of immersive three-dimensional environments from natural language narratives, consisting of: a speech capture module configured to continuously record a user's spoken narrative via one or more directional microphones, preprocesses the captured signal by noise reduction and temporal alignment, and outputs a digital speech stream; A speech-to-text processing unit that is operationally coupled to the speech capture module and configured for real-time speech recognition using a continuous neural transformer model. The unit is trained to transcribe natural language utterances into structured text data while maintaining contextual continuity throughout the evolving narrative. a semantic interpretation processing unit that is communicatively linked to the speech recognition unit and configured to perform natural language understanding techniques to extract contextual entities, spatial references, temporal relationships, and object attributes from the transcribed narrative; the engine includes a large language model that is fine-tuned for spatial reasoning tasks; a scene graph generation module configured to transform the interpreted semantic data into a structured, hierarchical representation that defines nodes for identified entities and edges for corresponding relationships, with each node associated with metadata describing geometry, position, orientation, texture, and linking attributes between objects; a multimodal image-language model processor coupled with the scene graph generation module, wherein the processor is configured to retrieve, adapt, or synthesize appropriate three-dimensional elements from a pre-trained visual-lexical embedding space and align these elements with their semantic and spatial definitions derived from the scene graph; a scene assembly and rendering controller configured to create a cohesive virtual scene from the aligned assets, perform real-time rendering using a GPU-accelerated ray tracing pipeline, and produce a stereoscopic visual output that corresponds to the evolving narrative; A head-mounted virtual reality visualization device connected to the rendering engine and configured to display the generated immersive environment to the user in real time. The device features motion sensors and inside-out tracking cameras to detect head and body movements, dynamically updating viewing angles and perspective within the rendered scene; and a bidirectional feedback module integrated into the head-mounted device and connected to the semantic interpretation processing unit; the module is configured to interpret corrective commands, gestures, or supplementary comments from the user to refine or modify specific scene elements without interrupting the real-time visualization; The system continuously updates the virtual scene as the narrative develops, ensuring temporal synchronization between speech input and rendered output below a defined latency threshold, thus enabling a natural, dialogic construction of complex three-dimensional virtual environments.
Owner:GOUNDER MOHAN SELLAPPA DR BENGALURU +3

Cabin active recommendation system and method based on knowledge graph and semantic reasoning

The invention discloses a cockpit active recommendation system and method based on a knowledge graph and semantic reasoning, and relates to the technical field of intelligent cockpits. The system receives natural language voice input of a user, executes voice recognition and semantic analysis, extracts user intention, keywords and slot entities, generates structured semantic information, constructs or calls a knowledge graph structure with semantic relation edges in combination with environment context information, and obtains the knowledge graph structure with the semantic relation edges. Semantic path reasoning is carried out based on the path dependence weight and the semantic similarity, a semantic edge label guided graph attention mechanism is introduced to calculate a path consistency score, a candidate recommendation set is generated, the semantic fitting degree and the path score are fused to sort and output recommendation content, and the graph edge weight and the user portrait are updated based on user feedback. According to the method, semantic understanding precision, recommendation path interpretability and system adaptive capacity are improved, and the method is suitable for personalized voice recommendation, man-machine interaction and scene linkage control tasks in an intelligent cockpit.
Owner:RIVOTEK TECH (JIANGSU) CO LTD

Comprehensive AI-enabled systems for immersive voice, companion, and augmented / virtual reality interaction solutions

A computer-implemented method for operating an artificial intelligence voice agent system includes receiving voice input through communication channels; analyzing converted text through natural language processing (NLP) pipelines implementing intent recognition and sentiment analysis detecting emotional cues using a multimodal large language model (LLM); generating response content using machine learning models trained on domain-specific corpora; converting generated responses to synthetic speech through text-to-speech (TTS) engines; integrating with a customer relationship management (CRM) platforms or an enterprise resource planning (ERP) database; and implementing continuous learning by updating language understanding models using conversation logs, voice recognition parameters based on user feedback, and response generation patterns. One implementation is a computer-implemented system and method that operates a suite of intelligent interactive devices and platforms including an artificial intelligence voice agent, enhanced communication platforms, an intimacy companion system, and augmented / virtual reality eyeglasses. Further, one implementation includes AR / VR eyeglasses that project visual content onto interchangeable lenses or directly onto the user's retina via laser-based retinal projection, provide prescription adjustments, incorporate ear-mounted sensors for monitoring physiological parameters like heart rate, oxygen saturation, and blood pressure, and utilize wireless data transmission, onboard environmental sensing, and remote calibration, all designed to offer dynamically adaptive, secure, and context-aware interactions across communication, personal assistance, health monitoring, and immersive augmented or virtual reality environments.
Owner:TRAN BAO

Humanoid robot multi-mode instruction analysis system

The invention discloses a multi-mode instruction analysis system for a humanoid robot. Comprising a voice input module, a visual input module, a voiceprint feature extraction module, an object recognition and pose estimation module, a multi-modal alignment network based on a space-time attention mechanism, a scene semantic tree construction module, an instruction node mapping module, a confidence evaluation module and a decision module. According to the system, accurate alignment of voice and visual information is realized through a space-time attention mechanism, environment information is represented in combination with a scene semantic tree structure, and the instruction analysis accuracy is improved. And dynamically evaluating the confidence coefficient by adopting a fuzzy instruction backtracking algorithm, and if the confidence coefficient is lower than a threshold value, starting multi-round dialogue clarification to reduce misoperation. According to the method, multi-modal data are fused, the historical interaction learning ability is optimized, the understanding efficiency and interaction robustness of complex instructions are remarkably improved, the method is suitable for scenes such as family service and logistics storage, and the intelligent level of man-machine cooperation is enhanced.
Owner:HUIZHOU BEIJIABAO ROBOT CO LTD

Real-time voice conversation method and system based on LLM technology and applied to operation and maintenance platform

The invention provides an LLM technology-based real-time voice conversation method and system applied to an operation and maintenance platform, and relates to the technical field of data processing, the method comprises the following steps: receiving user voice input, carrying out noise reduction preprocessing, and converting a voice signal into text data; semantic analysis is conducted on the text data through an LLMAgent module, the intention of a user instruction is recognized, and a task execution plan is generated based on the intention; decomposing the task execution plan into a basic operation sequence which can be executed by a Unity platform, and determining the priority of task execution; executing the basic operation sequence in a Unity environment, synchronizing an operation state in real time and feeding back an execution result; user interaction data and system performance data are collected, and LLM model parameters and decision strategies are optimized. The technical problems of poor voice interaction effect, large response delay, low accuracy and the like in the existing Unity platform are solved.
Owner:XIANGXING TECH ENG (GUANGDONG) CO LTD

Bank financing product recommendation method and system

The invention provides a bank financial management product recommendation method and system, and the method comprises the steps: firstly collecting a multi-modal dialogue data flow which is generated by interaction of a user terminal and an intelligent customer service system and contains voice input signals and text input contents, and then carrying out the semantic recognition of the voice input signals, and generating a structured demand description set; performing entity extraction and intention analysis on text input content to generate a multi-dimensional semantic tag set, performing time-space alignment and fusion on the text input content and the multi-dimensional semantic tag set to form a user portrait enhanced feature matrix, and performing multi-level association analysis on the user portrait enhanced feature matrix and a financial product knowledge graph based on a deep matching model to generate an initial recommendation result queue; and finally, according to the real-time environment perception data, queue context perception optimization is carried out, a dynamic adaptive recommendation list is generated, and visual rendering output is carried out through a multi-channel interaction interface, so that accurate and dynamic financial management product recommendation is realized by utilizing multi-modal data, deep analysis and combination with a real-time environment.
Owner:普益智慧云科技(成都)有限公司

Medical image processing method and system based on artificial intelligence

The invention discloses a medical image processing method and system based on artificial intelligence, and the method comprises the steps: carrying out the voice input of focus information through a voice input terminal, carrying out the preprocessing and character conversion output based on the inputted voice information, and dynamically generating a keyword set highly related to a current diagnosis task by combining structured priori knowledge of a medical knowledge graph and medical image-text multi-modal features, then generating a candidate diagnosis set based on the keyword set, calculating each semantic alignment score with the medical image features, generating a final diagnosis opinion through Langevin Dynamics sampling, and obtaining a final diagnosis result. And finally, the user terminal performs re-checking and enhanced correction, dynamically optimizes model parameters, and outputs a final formal diagnosis report based on an optimization correction result, thereby solving the accuracy and quality problems of traditional artificial intelligence on medical images, and realizing high-precision, high-robustness and interpretable artificial intelligence auxiliary diagnosis on the medical images.
Owner:SUZHOU YUNTU HEALTH TECH CO LTD

Method, device, equipment and product for executing task

The invention relates to a method, a device, equipment and a product for executing tasks. The method comprises the following steps: acquiring an operation sequence of a user for a first task on a user interface, wherein the operation sequence comprises operation steps required for completing the first task; the method includes converting the operating steps into a model context protocol service based on the sequence of operations. The method includes obtaining a user input, the user input indicating a second task, the user input including at least one of a text input or a voice input. The method includes determining a model context protocol service to be executed based on the second task, the model context protocol service to be executed being associated with the second task. The method includes executing a second task based on task parameters of the second task and service parameters of a to-be-executed model context protocol service, the second task being composed of the to-be-executed model context protocol service associated with the second task, and the task parameters of the second task determining and displaying an execution result of the second task based on user input.
Owner:BEIJING VOLCANO ENGINE TECH CO LTD

Vehicle state monitoring and early warning method and system

The invention belongs to the technical field of vehicle state detection and early warning, and particularly relates to a vehicle state monitoring and early warning method and system.The monitoring and early warning method comprises the steps that a sensor obtains real-time data, an anomaly detection algorithm is applied, and an abnormal event is marked; calculating an information priority according to the abnormal severity and the driving scene, and distributing the information priority to a high-priority queue; multi-mode early warning is generated, and high-frequency sound and vibration are used during high-speed driving; extracting a voice prompt, and generating voice waveform data through a voice synthesis module; voice input of a driver is recognized, intention is analyzed, and abnormal information feedback is provided; according to the scene and the abnormal state, interactive output is optimized, and detailed information is displayed during low-speed congestion; dynamically adjusting the interface of the instrument panel, and amplifying the key area in case of abnormal severity; and integrating the image data and the voice data, generating a multi-mode signal, and transmitting the multi-mode signal after rendering processing. The vehicle abnormal information can be timely and accurately transmitted to a driver, and the driving safety is effectively improved.
Owner:XIAN HUODA NETWORK TECH CO LTD

Intelligent safety protection management method and system

The invention relates to an intelligent safety protection management method and system. The method comprises the following steps: acquiring behavior data and position data of personnel in an industrial site; analyzing the behavior data and the position data by using a pre-trained personnel behavior model to obtain a behavior recognition result; according to the behavior recognition result and the current operation state of the target equipment, judging a risk level corresponding to the personnel behavior; based on the risk level, a corresponding safety response strategy is matched in the edge computing node, a control instruction is generated according to the safety response strategy, and the control instruction is used for driving the target device to execute a corresponding response action; and the voice interaction module is used for acquiring voice input of an on-site operator, performing semantic recognition on the voice input to obtain a voice recognition result, and executing corresponding emergency control operation to cover a current control instruction when the voice recognition result meets a preset emergency instruction condition. The method has the effect of improving the accuracy of intelligent safety management in the industrial environment.
Owner:SHENZHEN HUAYIXIN ELECTRONICS CO LTD

Multi-modal driven virtual digital human face animation generation method and system

The invention discloses a multi-mode driven virtual digital human face animation generation method and system, and relates to the field of computer graphics, and the method comprises the steps: obtaining voice input and text input, and extracting voice features and text features; dynamically fusing the two features through an attention fusion model to generate facial expression and head posture control parameters, and dynamically adjusting contribution weights of a voice mode and a text mode to the control parameters by adopting a driving strategy of differentiation of facial upper expression and facial lower expression; and performing local deformation on the virtual digital human face image based on the control parameter to generate an initial animation frame, and performing refining processing on the initial animation by using a generative adversarial network to obtain a refined face animation. By means of the technical scheme, the natural and vivid virtual human face animation matched with the voice content and the text semantics can be generated.
Owner:LIANGSHENG DIGITAL CREATIVE DESIGN (HANGZHOU) CO LTD

Real-time digital human generation system and method based on low-calculation-amount voice driving

The invention discloses a real-time digital human generation system and method based on low-calculation-amount voice driving, and relates to the technical field of digital human interaction, and the system comprises an audio processing module which is configured to receive voice input in real time and extract an audio feature vector; the driving and rendering module is used for mapping the audio feature vectors into parameters representing mouth movement, generating a dynamic mouth image based on the preprocessed static face reference data and the parameters, and fusing the dynamic mouth image with the reference image; the synchronous control module is used for ensuring the synchronization of the audio features and the rendering video frames according to a timestamp alignment mechanism and a PID feedback control algorithm; and the dynamic scheduling module is used for monitoring a hardware resource load in real time and realizing dynamic allocation of computing resources through multi-thread parallel and task priority adjustment. According to the technical scheme provided by the invention, the breakthrough of low delay, high fidelity and low energy consumption of the digital human can be realized in mobile terminal, embedded and multi-platform scenes.
Owner:LIANGSHENG DIGITAL ARTIFICIAL INTELLIGENCE (SHENZHEN) CO LTD

Intelligent dialogue method and device combining voiceprint recognition and voice synthesis

The invention relates to the technical field of voiceprint recognition, and particularly provides a voiceprint recognition and voice synthesis combined intelligent dialogue method, which comprises the following steps: acquiring a user voice input signal in real time; voiceprint feature extraction is carried out on the voice input signal, and a composite voiceprint vector containing user identity features and rhythm features is obtained; user identity matching is carried out based on the composite voiceprint vector, and a dynamic user portrait database is associated and called; generating personalized semantic feedback content according to the weight-updated dynamic parameters in the dynamic user portrait database; generating a target timbre synthesis parameter set based on the timbre representation vector in the composite voiceprint vector in combination with the spatio-temporal characteristic parameters of the interaction scene; and performing voice synthesis on the personalized semantic feedback content by adopting the target timbre synthesis parameter set, adjusting rhythm expression parameters of the synthesized voice based on real-time rhythm characteristics in the composite voiceprint vector, and finally outputting personalized voice, thereby realizing natural interaction of one person, one voice and one strategy.
Owner:MINAMI ACOUSTICS LTD

Video generation method and parameter generation model training method

Provided in the embodiments of the present disclosure are a video generation method and a parameter generation model training method. The video generation method comprises: acquiring speech to be processed; inputting an emotion feature of a target object and said speech into a parameter generation model, so as to obtain an expression parameter, wherein the expression parameter is used for describing facial motion information of the target object under the influence of the emotion feature, the parameter generation model is obtained by means of performing training on the basis of a sample emotion feature, sample speech and an expression parameter label that corresponds to the sample speech, and the sample emotion feature and the sample speech are obtained on the basis of a sample video; and inputting an object image of the target object and the expression parameter into a video generation model, so as to obtain a target video of the target object. By means of generating an expression parameter on the basis of an emotion feature and speech to be processed, and further generating a target video on the basis of the expression parameter, diversified emotion information is fused into the target video while ensuring the synchronization between speech and expression in the target video, thereby improving the accuracy and vividness of the target video.
Owner:ALIBABA (CHINA) CO LTD

System and method for multi-modal ai conversational interface improving website navigation and user interaction

The present invention relates to a system for transforming static websites into artificial intelligence (AI)-enabled interactive multi-modal conversational platforms. The system comprises a computing device having a processor for receiving user queries as text or speech input through an input module cooperating with a speech-to-text module. A natural language processing (NLP) module interprets intent, classifies user context, and retrieves grounded information from multiple webpages. A persona adaptation module dynamically modifies vocabulary, tone, and avatar representation across roles such as sales assistant, recruiter, educator, healthcare professional, etc. A response generator module produces structured natural language output, transmitted to a text-to-speech synthesis module and an avatar generation module to render synchronized lifelike video responses. An output rendering module displays multi-modal responses include text, audio, and video, thereby enabling direct navigation and escalation beyond limitations of conventional static websites.
Owner:NALLAM SREE RAMA CHANDRA MURTY

Flexible robot motion control system and method based on visual language action model

The invention discloses a flexible robot motion control system and method based on a visual language action model, and the system comprises visual sensors which are disposed at a plurality of joints of a flexible robot, and are used for observing the multi-joint sensing information of the flexible robot; the remote voice input unit is installed on the flexible robot body and used for inputting a voice instruction of remote operation to the flexible robot; the motion control module based on the VLA framework is mounted in the flexible robot and used for controlling the flexible robot to act based on a visual language action model according to the multi-joint sensing information and the voice instruction; through the design of the intelligent control system, the complexity of the system is remarkably reduced, the control precision and flexibility are optimized, the coordination control problem of complex body segment units is solved, high-performance autonomous motion control is achieved, the system cost is reduced, the adaptability of the robot in diversified environments is enhanced, and the robot can be widely deployed in complex application scenes.
Owner:JIANGSU IND INNOVATION CENT OF INTELLIGENT EQUIP CO LTD

Lightweight intelligent traditional Chinese medicine inquiry system and construction method thereof

The invention relates to the field of artificial intelligence medical application, and discloses a lightweight intelligent traditional Chinese medicine inquiry system and a construction method thereof, and the system comprises a multi-dialect adaptive speech recognition module, a traditional Chinese medicine intelligent dialogue large language model module, a natural speech synthesis module, and a continuous learning mechanism module. The multi-dialect adaptive speech recognition module is used for converting dialect speech input of a patient into a standard text; the traditional Chinese medicine intelligent dialogue big language model module is the core of the system and is used for carrying out natural language understanding, dialectical reasoning and inquiry dialogue generation, and the natural speech synthesis module is used for converting a text response generated by the system into speech output; and the continuous learning mechanism module realizes continuous optimization of the large language model through incremental learning architecture and clinical feedback integration. According to the method, while the professional traditional Chinese medicine diagnosis capability is maintained, the calculation complexity is remarkably reduced, and the universality and sustainable development capability of system application are improved.
Owner:SUZHOU ANGSHENG NETWORK TECHNOLOGY CO LTD

Large model-based automobile live risk assessment system and method

The invention belongs to the field of automobile evaluation, and provides an automobile live risk evaluation system and method based on a large model, and the method comprises the steps: receiving voice input information of an evaluator in a vehicle on-site inspection process; extracting component state information of the target vehicle based on the structured check item text; acquiring historical maintenance record data of the target vehicle, and performing association fusion on the historical maintenance record data and the component state information; based on a large language model, performing joint reasoning on a currently input inspection label and the historical maintenance record, constructing a causal chain related to an accident risk, and generating at least one subsequent to-be-inspected item; generating voice guide content and feeding back the voice guide content to the evaluator so as to prompt the evaluator to continue to check the target item of the target vehicle; gradually perfecting the risk check path of the target vehicle until the risk chain closed loop is completed; and calling an estimation formula to calculate a residual estimation value of the target vehicle, and outputting a structured risk assessment report.
Owner:GUANGZHOU SUISHENG INFORMATION TECH CO LTD

Intelligent form filling method based on voice interaction

The invention discloses an intelligent form filling method based on voice interaction. A form filling system is composed of an intention recognition module, an entity recognition module, a data mapping module, an intelligent error correction module and a dialogue management module. According to the system, the traditional manual form filling operation is replaced by a voice input mode, the complexity and time cost of user input are greatly reduced, and the recognition precision of terminologies is remarkably improved by utilizing a voice recognition and natural language processing model optimized for the financial field, so that the user experience is improved. Specific fields of the form can be accurately recognized according to intention recognition and data mapping functions customized for a specific scene of financial reimbursement, the system is designed based on a multi-round dialogue management module, a user is actively guided to complete complex form filling step by step, the learning cost and the operation difficulty are remarkably reduced, and the user experience is improved. Full-process automation of voice recognition, data mapping and intelligent verification is realized, and manual intervention is reduced.
Owner:YONYOU FINANCIAL INFORMATION TECH

Use of generative artificial intelligence for interactive television recommendations

Method comprising: receiving, by a computing device, an indication to launch a voice-based television assistant; displaying a user interface including a prompt for receiving verbal input; receiving first voice data for a first query related to a media content recommendation; sending, by the computing device and to a server computer, the first voice data for the first query; receiving a response to the first query, the request for the additional information being based on at least one of previous queries, previous responses, or a context for the first query; and in response to receiving the response, receiving second voice data for a second query related to the media content recommendation. Method for providing a personalized description of a media content item and method for providing a trivia game related to a media content item.
Owner:GOOGLE LLC

Inferring intent from pose and speech input

Aspects of the present disclosure involve a system comprising a computer-readable storage medium storing at least one program, and a method for performing operations comprising receiving an image that depicts a person, identifying a set of skeletal joints of the person and identifying a pose of the person depicted in the image based on positioning of the set of skeletal joints. The operations also include receiving speech input comprising a request to perform an AR operation and an ambiguous intent, discerning the ambiguous intent of the speech input based on the pose of the person depicted in the image and in response to receiving the speech input, performing the AR operation based on discerning the ambiguous intent of the speech input based on the pose of the person depicted in the image.
Owner:SNAP INC

Intelligent dialogue method and system based on digital human

The invention relates to the technical field of artificial intelligence and multimedia processing, and particularly provides an intelligent dialogue method and system based on a digital human, and the method comprises the following steps: S1, an input stage: a user inputs a question or an instruction through a text or voice; s2, the processing stage comprises large model generation reply, TTS model conversion and audio driving mouth shape; and S3, the output stage comprises video generation, stream pushing and display. Compared with the prior art, the method has the advantages that the reply can be generated through the large model, and the TTS model and the audio-driven mouth shape model are combined, so that more natural and anthropomorphic dialogue interaction is realized, and the user experience is improved.
Owner:浪潮智慧城市科技有限公司

Movement planning method, system and equipment of disabled-helping robot, storage medium and computer program

The invention discloses a movement planning method, system and device for a disabled assisting robot, a storage medium and a computer program. The method comprises the steps that S1, a language instruction is received and processed; s2, sensing an environment and constructing a map; s3, analyzing data and planning a path; and S4, state estimation and dynamic obstacle avoidance. Through voice input, a laser SLAM technology and a grid map optimization method, a high-precision two-dimensional map of the disabled assisting robot is constructed, and the accuracy of navigation planning is remarkably improved; and a more optimized path planning method is adopted, so that the path planning efficiency is higher and more stable.
Owner:ZHEJIANG UNIV OF TECH

Dynamic health monitoring system integrating artificial intelligence and wearable device and method thereof

The invention provides a dynamic health monitoring system integrating artificial intelligence and wearable equipment and a method thereof, and relates to the field of health monitoring, and the method comprises the steps: collecting physiological parameters and environmental data of a user, eliminating time migration of multi-source data through a timestamp alignment technology, extracting dynamic coupling features of the physiological parameters and the environmental data, and obtaining a dynamic health monitoring result; and generating the health state characterization guided by the environment. The method comprises the following steps: synchronously acquiring user voice input information, and extracting and analyzing acoustic features to generate emotional state change features after performing Mel spectrum conversion; then, cross-modal semantic association between emotions and health states is established through a modal soft constraint mechanism, and personalized guidance suggestions are generated and displayed, so that the dynamic perception ability of health state assessment and the scene adaptability of suggestion generation are improved; therefore, the problems of environmental factor splitting, subjective and objective data isolation and insufficient dynamic nonlinear relation modeling in a traditional method are solved.
Owner:JIANGXI YANGNING TECHNOLOGY CO LTD

Spoken Chinese pronunciation automatic evaluation and correction system based on AI speech recognition

The invention relates to the technical field of speech recognition and artificial intelligence, in particular to an automatic evaluation and correction system for spoken Chinese pronunciation based on AI speech recognition, which comprises a speech input module, a speech preprocessing module, a pronunciation analysis module, a pronunciation scoring module and a pronunciation correction module. Wherein the voice input module is used for receiving a Chinese spoken pronunciation signal of a user; the voice preprocessing module is used for preprocessing the voice data; the pronunciation analysis module is used for analyzing pronunciation of the user in real time by using an AI voice recognition technology; the pronunciation scoring module is used for calculating a score value of pronunciation of the user; and the pronunciation correction module is used for providing a targeted pronunciation correction scheme for the user. According to the invention, through the multi-dimensional pronunciation evaluation and personalized correction scheme based on the AI speech recognition technology, the pronunciation deviation of the user can be accurately recognized, real-time feedback can be provided, and the pronunciation correction accuracy and efficiency can be effectively improved.
Owner:HUBEI UNIV OF EDUCATION

Enterprise intelligent finance and tax system fusing supervised learning and block chain

The invention provides an enterprise intelligent finance and taxation system fusing supervised learning and a block chain, and solves the core technical problems of insufficient data credibility, transparent contradiction between privacy protection and supervision, AI model training data island and the like of a traditional finance and taxation system. According to the system, an innovative technical fusion scheme is adopted, bank API, OCR recognition, voice input and other multi-source financial data are integrated through a multi-modal data fusion unit, and high-quality data fusion is achieved through confidence evaluation and an intelligent conflict resolution algorithm; an intelligent classification unit based on BERT deep learning integrates a pre-training model with a business rule engine and a gradient boosting decision tree, and accurate and automatic classification of financial transactions is achieved. A complete technical solution is provided for enterprise finance and taxation digital transformation, intelligent, automatic and credible processing of finance and taxation businesses is achieved on the premise that data safety and privacy protection are guaranteed, and the method has wide market application prospects.
Owner:张宏

Cross-domain voice classification method and device based on feature decoupling and multi-task learning

PendingCN120452429ASpeech recognitionPhonetic environmentData set
The invention relates to a cross-domain voice classification method and device based on feature decoupling and multi-task learning. The method comprises the following steps: firstly, acquiring a multi-data-domain voice file, and preprocessing the multi-data-domain voice file to obtain a cross-domain voice classification data set; then, constructing a cross-domain voice classification model which comprises a voice feature encoder module, a data domain classification module, a supervised comparative learning module and a multi-task classification module; then, a joint optimization loss function is constructed based on data field classification loss, supervised contrast learning loss and task classification loss, and a gradient descent algorithm is adopted to train and optimize the cross-domain voice classification model based on the cross-domain voice classification data set and the joint optimization loss function; and finally, inputting to-be-classified voice into the trained cross-domain voice classification model to obtain a voice corresponding category. The discrimination ability and generalization ability of the model in a cross-domain scene are significantly improved, so that the model still maintains high classification precision in a complex multi-source voice environment.
Owner:SICHUAN UNIV

Electronic medical record automatic generation method based on voice recognition

InactiveCN120690205ASpeech recognitionPatient-specific dataMedical recordSpeech segmentation
The invention relates to the technical field of electronic medical record generation, and discloses an electronic medical record automatic generation method based on voice recognition, which comprises the following steps: S1, initializing voice input, distributing a unique voice acquisition identifier for a medical session, and completing identifier generation, input, storage, association and identity verification; s2, voice information intelligent recognition: converting voice into a text by using a voice recognition engine, and ensuring semantic consistency through a context verification unit and a semantic analysis unit; and S3, performing multi-speaker processing based on intelligent interference detection and resolution, positioning an interference time period and an interference source through an interference detection unit, and realizing time period distribution and priority ranking of multi-speaker voices by using a voice segmentation protocol and a linear weighting model. And finally, extracting related information from the text generated by voice conversion, and filling the related information into a medical record template of a hospital. The method improves the efficiency and accuracy of electronic medical record generation, solves the problems of multi-speaker interference and semantic logic, and is suitable for medical informatization scenes.
Owner:THE FIRST AFFILIATED HOSPITAL OF GUANGZHOU MEDICAL UNIV (GUANGZHOU RESPIRATORY CENT)