Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

15005results about "Speech recognition" patented technology

Real-time virtual reality scene system based on natural language description using multimodal artificial intelligence

A real-time system for the multimodal generation of virtual reality scenes based on artificial intelligence for the creation of immersive three-dimensional environments from natural language narratives, consisting of: a speech capture module configured to continuously record a user's spoken narrative via one or more directional microphones, preprocesses the captured signal by noise reduction and temporal alignment, and outputs a digital speech stream; A speech-to-text processing unit that is operationally coupled to the speech capture module and configured for real-time speech recognition using a continuous neural transformer model. The unit is trained to transcribe natural language utterances into structured text data while maintaining contextual continuity throughout the evolving narrative. a semantic interpretation processing unit that is communicatively linked to the speech recognition unit and configured to perform natural language understanding techniques to extract contextual entities, spatial references, temporal relationships, and object attributes from the transcribed narrative; the engine includes a large language model that is fine-tuned for spatial reasoning tasks; a scene graph generation module configured to transform the interpreted semantic data into a structured, hierarchical representation that defines nodes for identified entities and edges for corresponding relationships, with each node associated with metadata describing geometry, position, orientation, texture, and linking attributes between objects; a multimodal image-language model processor coupled with the scene graph generation module, wherein the processor is configured to retrieve, adapt, or synthesize appropriate three-dimensional elements from a pre-trained visual-lexical embedding space and align these elements with their semantic and spatial definitions derived from the scene graph; a scene assembly and rendering controller configured to create a cohesive virtual scene from the aligned assets, perform real-time rendering using a GPU-accelerated ray tracing pipeline, and produce a stereoscopic visual output that corresponds to the evolving narrative; A head-mounted virtual reality visualization device connected to the rendering engine and configured to display the generated immersive environment to the user in real time. The device features motion sensors and inside-out tracking cameras to detect head and body movements, dynamically updating viewing angles and perspective within the rendered scene; and a bidirectional feedback module integrated into the head-mounted device and connected to the semantic interpretation processing unit; the module is configured to interpret corrective commands, gestures, or supplementary comments from the user to refine or modify specific scene elements without interrupting the real-time visualization; The system continuously updates the virtual scene as the narrative develops, ensuring temporal synchronization between speech input and rendered output below a defined latency threshold, thus enabling a natural, dialogic construction of complex three-dimensional virtual environments.
Owner:GOUNDER MOHAN SELLAPPA DR BENGALURU +3

Text prediction-based large-model real-time voice text intention recognition method and system

The invention discloses a large-model real-time voice text intention recognition method and system based on text prediction, and the method comprises the steps: obtaining the real-time voice data of a user, carrying out the real-time voice recognition processing through a streaming voice recognition interface, and obtaining a part of transcriptional text; inputting the partial transcription text into a mask language model for text prediction, and generating a plurality of high-credibility complete sentence candidates; based on the complete sentence candidates, the complete sentence candidates are input into a large language model in parallel for intention recognition, a corresponding intention result is obtained, and a mapping relation between the candidate sentences and the intention recognition result is established; and obtaining a sentence completely expressed by the user, calculating the similarity between the complete actual sentence and a plurality of high-credibility complete sentence candidates through a multi-level text similarity algorithm, selecting the candidate sentence with the highest similarity score, and directly obtaining a corresponding final intention recognition result based on the mapping relationship. The objective of the invention is to solve the technical problem of high response delay of an existing voice intention recognition system.
Owner:BEIJING YULORE INNOVATION TECH

Comprehensive AI-enabled systems for immersive voice, companion, and augmented / virtual reality interaction solutions

A computer-implemented method for operating an artificial intelligence voice agent system includes receiving voice input through communication channels; analyzing converted text through natural language processing (NLP) pipelines implementing intent recognition and sentiment analysis detecting emotional cues using a multimodal large language model (LLM); generating response content using machine learning models trained on domain-specific corpora; converting generated responses to synthetic speech through text-to-speech (TTS) engines; integrating with a customer relationship management (CRM) platforms or an enterprise resource planning (ERP) database; and implementing continuous learning by updating language understanding models using conversation logs, voice recognition parameters based on user feedback, and response generation patterns. One implementation is a computer-implemented system and method that operates a suite of intelligent interactive devices and platforms including an artificial intelligence voice agent, enhanced communication platforms, an intimacy companion system, and augmented / virtual reality eyeglasses. Further, one implementation includes AR / VR eyeglasses that project visual content onto interchangeable lenses or directly onto the user's retina via laser-based retinal projection, provide prescription adjustments, incorporate ear-mounted sensors for monitoring physiological parameters like heart rate, oxygen saturation, and blood pressure, and utilize wireless data transmission, onboard environmental sensing, and remote calibration, all designed to offer dynamically adaptive, secure, and context-aware interactions across communication, personal assistance, health monitoring, and immersive augmented or virtual reality environments.
Owner:TRAN BAO

Personalized and dynamic text to speech voice cloning using incompletely trained text to speech models

Systems and methods are provided for machine learning models configured as zero-shot personalized text-to-speech models which comprise a feature extractor, a speaker encoder, and a text-to-speech module. The feature extractor is configured to extract acoustic features and prosodic features from new target reference speech associated with the new target speaker. The speaker encoder is configured to generate a speaker embedding corresponding to the new target speaker based on the acoustic features extracted from the new target reference speech. The text-to-speech module is configured to generate the personalized voice corresponding for the new target speaker based on the speaker embedding and the prosodic features extracted from the new target reference speech without applying the text-to-speech module on new labeled training data associated with the new target speaker.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Instruction understanding and task execution method and device, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to service scenes of pension service, financial science and technology, medical health and the like, and discloses an instruction understanding and task execution method, device, equipment and medium, and the method comprises the steps: receiving a voice instruction and a text instruction, and carrying out the cooperative processing through an instruction understanding model, and generating a structured task description; collecting environment data to construct a real-time environment model; generating a task execution strategy by utilizing a task execution model based on the structured task description and the real-time environment model; controlling the intelligent agent to execute the task according to the task execution strategy, and dynamically adjusting the action in combination with real-time sensor information; task execution data and user feedback information are collected, and the instruction understanding model and the task execution model are updated. According to the method, multi-modal information is fused through structural description, an execution strategy is generated in combination with real-time environment perception, actions are dynamically adjusted, model self-optimization is further achieved through execution data and feedback, and the understanding, decision-making and adaptive capacity of an intelligent agent is improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Operation intention recognition method, system and equipment based on multi-modal fusion and medium

The invention relates to the technical field of data processing, and particularly provides an operation intention recognition method, system and device based on multi-mode fusion and a medium, and the method comprises the steps: synchronously collecting interaction data of at least two modes of a user, the modes comprising at least two of gestures, voice and eye gaze; carrying out alignment processing on the interaction data, wherein the alignment processing comprises time synchronization and space mapping to a unified coordinate system; recognizing structured semantic information from each piece of aligned modal data, wherein the structured semantic information comprises a gesture type, a voice text and a fixation point coordinate; and based on a preset semantic rule and context memory, performing semantic association and anaphora resolution on the structured semantic information to obtain an operation intention. The method effectively overcomes the inherent defects of unnatural single-mode interaction, easy ambiguity and poor fault tolerance.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Data processing method and apparatus, electronic device, computer readable storage medium and computer program product

The present application provides a data processing method and apparatus, an electronic device, a computer readable storage medium and a computer program product. The method comprises: acquiring historical interaction information and a predicted interaction text corresponding to the historical interaction information; extracting a first acoustic feature and a first semantic feature of the historical interaction information, and extracting a second semantic feature of the predicted interaction text; performing fusion mapping on the basis of the first acoustic feature, the first semantic feature and the second semantic feature to obtain a first paralanguage feature; denoising initial noise on the basis of the second semantic feature and the first paralanguage feature to obtain a second acoustic feature of the predicted interaction text; and on the basis of the second acoustic feature, generating a voice signal corresponding to the predicted interaction text.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Natural language processing

Techniques for generating an executable API call for an LLM-generated request, where the executable API call is usable to cause a component to generate a potential response to a user input, are described. In some embodiments, the system receives a user input and uses a language model to generate a request for a component to provide a potential response to the user input. The system uses the request, an API description corresponding to the component, and other information not available to the language model during processing to generate an executable API call corresponding to the request. The system can execute the executable API calls (in a system-determined order or concurrently) to cause the corresponding components to generate potential responses to the user input.
Owner:AMAZON TECH INC

Intelligent voice recognition and natural language interaction method based on quadruped robot

The invention discloses an intelligent voice recognition and natural language interaction method based on a quadruped robot, and the method comprises the following steps: S1, collecting a user voice instruction, generating a standardized voice text, and extracting a semantic keyword set; s2, collecting multi-source sensing data of the quadruped robot and generating a structured state data tensor; s3, constructing a multi-modal collaborative modeling mechanism, and generating a multi-modal joint semantic embedding vector; s4, constructing a semantic map based on semantic embedding and generating an action chain plan structure; s5, executing each sub-action in the action chain, and performing path analysis and execution monitoring; s6, storing the interaction task as a multi-modal semantic behavior memory unit; and S7, performing similarity retrieval based on current semantic input and historical memory to realize behavior migration and action chain multiplexing. The method has the advantages of accurate semantic understanding, intelligent interaction response, high behavior migration capability and the like.
Owner:山东浪潮数据库技术有限公司

Government affair intelligent interaction and information extraction method and device, equipment and medium

The invention discloses a government affair intelligent interaction and information extraction method and device, equipment and a medium, and relates to the technical field of digital government affair, digital cities, intelligent interaction and the like. The traditional manual item-by-item filling is converted into end-to-end automatic processing, so that the operation steps of a user are remarkably reduced, the filling burden of a hundred-item-level form is solved, and the filling efficiency is improved; and an ID voice corrector is adopted to realize three-layer progressive verification, so that the output ID offset of the multi-modal large model is effectively inhibited, the accurate mapping of field identifiers is ensured, and meanwhile, a lightweight knowledge graph is adopted to automatically capture logic rules among fields, interspecific compliance conflicts of government affair scenes are intercepted in real time, the logic error risk is reduced, and the filling accuracy is improved.
Owner:SICHUAN ENRISING INFORMATION TECH CO LTD

Evaluating confidence in a classification performed by a generative language machine learning model

A large language model (LLM) may be used to classify an input into one of a plurality of categories. However, given the machine-learning operation of the LLM, the output of the LLM does not represent a definitive statement, but is based on probability computations of the machine learning model. Therefore, the classification performed by the LLM might not be correct. Classification into the wrong category by the LLM results in downstream technical problems. In some implementations, when an LLM generates a response that classifies an input, one or more probability values associated with a token that forms the basis of the response may be used to determine a confidence value. The confidence value is indicative of confidence in the classification performed by the LLM. An action may be taken based on the confidence value.
Owner:SHOPIFY INC

Precise international communication digital human real-time dialogue method fused with multi-modal technology

The invention discloses an accurate international communication digital human real-time dialogue method fused with a multi-modal technology. The method comprises the following steps: S1, constructing a digital human image and tone; s2, propagation content generation and problem guidance; s3, semantic analysis and intention clarification based on the real-time voice dialogue; s4, geographic preference modeling and path planning; s5, cross-context propagation content generated based on retrieval enhancement is generated; s6, visually displaying the propagation content; and S7, carrying out digital human-driven multi-language propagation content real-time output and feedback closed-loop optimization. Through accurate utterance expression analysis, accurate international propagation problem recommendation is provided, cross-context propagation content generation based on semantic understanding is realized, digital people with voice features and visual images are constructed, real-time dialogue interaction of users is realized, and user experience is improved. The system can carry out geographic modeling according to the region where the accurate problem is located, language preference and propagation object culture characteristics, and differential propagation path planning is achieved.
Owner:HUNAN NORMAL UNIVERSITY

Intelligent conference summary automatic generation method based on voice recognition and large model

The invention discloses an intelligent conference summary automatic generation method based on voice recognition and a large model. The method comprises the following steps: S1, executing voice activity detection operation on an audio data stream; s2, extracting embedding vectors of continuous and effective voice segments, and generating a voice segment set to which a spokesman belongs; s3, inputting the voice fragment set to which the spokesman belongs into an improved Whisper model, fusing a Speaker-Aware attention mechanism and a connection time sequence classification auxiliary path, and outputting a conference transcription text sequence set; s4, inputting the processed structured dialogue format into a GPT-4 large language model, and generating a conference semantic representation sequence; s5, generating a conference summary first draft text according to a preset summary generation template; and S6, performing formatting output operation on the conference summary first draft text. The conference semantic elements can be automatically extracted, the structured summary text can be generated, and the method is suitable for efficient conference recording and task tracking in government affair office, enterprise collaboration, academic discussion and other scenes.
Owner:JIANGSU GUOHUACHENJIAGANG POWER GENERATION CO LTD

Audio and video player control method based on voice instruction

The invention relates to the technical field of audio and video control, and discloses an audio and video player control method based on a voice instruction. The method comprises the steps that an original voice instruction stream of a user is collected, the instruction stream comprises a time domain audio signal sequence, an environment noise spectrum and user pronunciation characteristic parameters, and voice information can be comprehensively captured; multi-modal instruction analysis processing is carried out on the original voice instruction stream, a structured control instruction set containing acoustic control intention identification, semantic operation object description and context correlation parameters is generated, and the analysis precision is improved; then executing player state adaptation based on the set, generating a dynamic control response sequence containing an equipment state adjustment command, a media content positioning parameter and an interface interaction logic identifier, driving a player to execute a multi-dimensional control operation and generating real-time play control effect feedback data; and finally, multi-modal analysis parameters are optimized according to feedback data, a self-adaptive instruction analysis strategy is generated, and the control experience of a user on the audio and video player is optimized.
Owner:ONWAY TECH LTD

Training and speech generation methods and apparatuses for speech generation model, electronic device, computer-readable storage medium, and computer program product

The present application provides training and speech generation methods and apparatuses for a speech generation model, an electronic device, a computer-readable storage medium, and a computer program product. The method comprises: obtaining a first speech generation model; obtaining sample data of a plurality of modalities; on the basis of a prompt image sequence and speech text, respectively calling a plurality of encoders to perform encoding, so as to obtain a multi-modal encoding vector sequence; on the basis of the multi-modal encoding vector sequence, calling a decoder to perform decoding, so as to obtain decoded text; determining a probability distribution for the decoded text and the sample data of the plurality of modalities, and determining a target loss on the basis of the probability distribution; and on the basis of the target loss, updating parameters of the decoder and at least one of the encoders, wherein the updated decoder and the plurality of updated encoders are configured to form a second speech generation model, and the second speech generation model is used to generate target speech text corresponding to a prompt image sequence of a target object.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Natural language understanding systems

Techniques are described for identifying functionalities (i.e., user experiences) that are requested by users but are not supported by natural understanding (NU) processing. Some embodiments may involve identifying functionalities by transforming user inputs to functionality-based representations. The functionality-based representations may be grouped into individual functionalities. The user inputs associated with an individual functionality may be evaluated using an NU component to determine whether the functionality is supported. These techniques may enable discovery at a functionality level, rather than at a user input level, an intent level, or an entity level. These techniques may also be used to group user inputs to determine trending functionalities.
Owner:AMAZON TECH INC

Cooperation between language models

A system may be configured for cooperation between language model agents. An agent may be, for example, a computer system, or a software component executing on a computer system, that can accept text and / or natural language inputs, draw upon an LM to process the inputs and perform a function, and respond via text and / or natural language outputs. An agent may act as a mediator to interact with a user, identify a task requested by the user, and delegate one or more subtasks to another agent or other resource. An agent may act as a delegate to handle tasks or subtasks delegated by a mediator. Agents may communicate with each other using a combination of structured and unstructured language; for example, one or more parameters and a natural language message.
Owner:AMAZON TECH INC

Speech recognition method and related device

ActiveCN114360510AImprove fault tolerancePrecise Syllable Probability DistributionSpeech recognitionSyllableAcoustic model
The embodiment of the invention discloses a speech recognition method and a related device, and at least relates to a speech recognition technology in artificial intelligence, speech data to be recognized are used as input data of a time delay neural network in an acoustic model, and an output layer of the time delay neural network comprises acoustic modeling units corresponding to a plurality of syllables respectively, so that the speech recognition efficiency is improved. And the syllable probability distribution corresponding to the voice frames included in the voice data can be obtained by taking the syllables as the recognition granularity through the time delay neural network. When syllable recognition is carried out through the output layer, auxiliary judgment can be carried out on the syllables to which the voice frames belong on the basis of pronunciation rules in combination with front and back syllable information of the voice frames, so that more accurate syllable probability distribution is output. Moreover, since the syllables are generally composed of one or more phonemes, the method has higher fault-tolerant capability, not only can more accurately determine the speech recognition result based on the probability distribution of the syllables, but also has low requirements for the quality of the speech data to be recognized, and effectively expands the application scenarios of the speech recognition technology.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Industrial Internet of Things anomaly detection method based on time sequence and text joint modeling

The invention relates to an industrial Internet of Things anomaly detection method based on time sequence and text joint modeling, and belongs to the technical field of industrial Internet of Things anomaly detection. The method comprises the following steps: constructing text prompt information based on collected industrial Internet of Things time sequence data, and respectively taking the text prompt information as inputs of a time sequence channel and a text prompt channel; a sensor association graph is constructed by using a multi-hop GCN, and on the basis of the association graph, time feature modeling from local to global is completed by using multi-scale expansion convolution and combining a differential attention mechanism; performing word segmentation processing on the text prompt information through a word segmentation device, and encoding the text prompt information into vector representation; and calculating attention weight between time sequence embedding and text prompt embedding, fusing to obtain joint embedding representation, enhancing the joint embedding representation, inputting the enhanced joint embedding representation into MLP for reconstruction, calculating an abnormal score through a reconstruction error, and carrying out industrial Internet of Things anomaly detection according to the abnormal score. The method is high in anomaly detection accuracy, and can improve the equipment anomaly perception and risk early warning capability.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

System for real-time analysis of emotional feedback during motivational presentations

A system for real-time analysis of emotional feedback during motivational speeches, consisting of: a series of multimodal sensors, including at least one visual sensor configured to capture facial expressions of spectators, at least one directional microphone configured to capture the audio responses of the audience, and optionally one or more physiological sensors configured to capture biometric signals from spectators; an edge-based processing unit that is communicatively coupled to the arrangement of multimodal sensors, wherein the edge-based processing unit comprises the following: (a) a feature extraction module configured to extract visual features from captured facial images, acoustic features from voice responses, and physiological features from biometric signals; (b) an emotion inference machine configured to process the features using a deep learning-based emotion recognition model comprising a convolutional neural network (CNN) for classifying facial expressions, a recurrent neural network (RNN) for classifying voice emotions, and a multimodal late fusion layer configured to compute a composite emotion state vector representing the aggregated emotions of the audience; (c) a timestamp and speech alignment module configured to correlate the calculated composite emotion state vector with segmented portions of a live motivational speech based on real-time speech-to-text transcription and semantic analysis; and (d) a session-based storage unit configured to log time-indexed emotional state vectors and corresponding speech segments for post-event analysis; A speaker feedback interface comprising a portable display or a podium-mounted visualization panel, wherein the interface is configured to display visual indicators of emotional feedback in real time, the indicators being derived from the emotional state vector and including at least emotional trend graphs, threshold alerts, or engagement indices.
Owner:1XL LLC FZ +2

Intelligent crawler generation method and system based on large language model and MCP protocol

The invention discloses an intelligent crawler generation method and system based on a large language model and an MCP protocol, belongs to the technical field of network data collection, and solves the problem that the capability of LLM in dynamic webpage analysis and anti-crawling strategy generation links cannot be fully exerted due to the fact that LLM and browser interaction protocols cannot be effectively integrated in the prior art. The method comprises the steps of analyzing an acquisition demand based on a large language model and generating a standardized demand description document, realizing interaction between the large language model and a browser based on an MCP protocol, analyzing a page complete DOM tree structure through a crawler script generation system, and performing quality verification and intelligent repair on a generated crawler script. According to the method, the complete DOM tree and the dynamic data rendered by the browser are obtained through the MCP, and the large language model can be called to automatically analyze the element positioning strategy, so that the collection script is adaptively generated, and high efficiency, intelligence and automation of webpage data collection are ensured.
Owner:钰兔科技集团有限公司

Consultation method and system based on natural language processing and legal knowledge graph

The invention discloses a consultation method and system based on natural language processing and a legal knowledge graph, and relates to the field of data processing, and the method comprises the steps: receiving a multi-format legal consultation demand of a user, converting the multi-format legal consultation demand into a text, inputting the text into a BERT law NLP model, and analyzing key information through word segmentation, intention recognition and entity extraction; based on a pre-constructed multi-level legal knowledge graph, carrying out accurate and fuzzy retrieval and domain filtering in combination with an analysis result, and obtaining an associated law article, a case and a legal relationship; screening conflict law articles and similar cases, and inputting the conflict law articles and the similar cases into a graph neural network reasoning model to generate a preliminary conclusion; the conclusion is converted into a spoken consultation report through a natural language generation module, and output is customized according to a user scene; and if the user feedback satisfaction degree is less than the threshold value, iteratively optimizing the storage data to the historical library. The method has the advantages that accurate retrieval is realized based on the BERT model and the multi-level knowledge graph in the legal field, the oral personalized conclusion combined with the user scene is generated through GNN reasoning, and iterative optimization is performed through user feedback.
Owner:BEIJING INSTITUTE OF TECHNOLOGY (ZHUHAI)

Understanding user intent and enhancing navigation of data analytics through natural language interfaces

Provided is a technique referred to as Voice to Analytics (“Vox2A”), an approach to produce analytic results from a collection of data by using natural language interrogation based on broad but constrained interpretation of user intent to create and present a set of responses containing the sought-after information. The result may be a faster “time-to-analytical answers” tool featuring a shorter user learning curve and easy to navigate experience, making for faster, more informative results, thus improving user productivity.
Owner:CEREBRI AI INC

Intelligent homework tutoring system and method based on multi-modal interaction and adaptive learning

The invention relates to the technical field of artificial intelligence education, in particular to an intelligent homework tutoring system and method based on multi-modal interaction and adaptive learning, and the system comprises a multi-modal input analysis module, a cognitive state dynamic evaluation module, an intelligent decision explanation engine, and a learning effect visualization closed-loop module. The method has the beneficial effects that a three-dimensional adaptive system of explanation granularity-presentation form-interaction frequency is used, and teaching strategies such as visual derivation, concept metaphor and the like can be automatically matched according to cognitive styles of students. Secondly, a'backtracking reinforcement-lateral expansion 'double-intervention mechanism is innovatively used, and the recurrence rate of similar errors is effectively reduced through error real-time detection and correlation knowledge contrast teaching. And finally, constructing a dynamic knowledge graph and an interactive learning report, realizing visual tracking of a learning path, and helping students to establish a systematic knowledge framework.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Natural language processing

Techniques for generating tasks to be completed in order to perform an action responsive to a user input and, for a given task, shortlisting available components to those that are relevant for the task are described. The system processes a user input to determine tasks to be completed in order to perform an action responsive to the user input. The system determines a priority of the tasks and selects a top-ranked task. The system determines descriptions of processing performable by components that are semantically similar to the current task, and requests a description of the function the corresponding components would perform for the current task. Based on the received descriptions, the system selects one or more components to perform the task. Thereafter, the system causes the action to be performed and outputs a response to the user input.
Owner:AMAZON TECH INC

Electronic medical record LLM generation method based on animal injury

The invention discloses an electronic medical record LLM generation method based on animal injury, which realizes dialogue structuring and timestamp synchronization through multistage speech recognition and role affiliation. Using standardized medical term mapping and coding to align the free text to a standardized medical entity, and constructing a high-confidence medical entity network based on a semantic anchor point pool; according to the method, context-sensitive entity relationship extraction is realized by combining a large language model and a semantic enhancement template, a high-accuracy structured relationship chain is generated through clinical logic rule set verification, and finally, an electronic medical record template under diagnosis and treatment specifications is automatically filled and privacy desensitization processing is completed. The semantic consistency, the structural accuracy and the data security of automatic generation of the electronic medical record are improved, and standardization and intelligent circulation of medical information are effectively promoted.
Owner:GUANGZHOU WUCHUAN ELECTRONIC TECHNOLOGY CO LTD +1

Deepfake detection

Disclosed are systems and methods including software processes executed by a server that detect audio-based synthetic speech (“deepfakes”) in a call conversation. The server applies an NLP engine to transcribe call audio and analyze the text for anomalous patterns to detect synthetic speech. Additionally or alternatively, the server executes a voice “liveness” detection system for detecting machine speech, such as synthetic speech or replayed speech. The system performs phrase repetition detection, background change detection, and passive voice liveness detection in call audio signals to detect liveness of a speech utterance. An automated model update module allows the liveness detection model to adapt to new types of presentation attacks, based on the human provided feedback.
Owner:PINDROP SECURITY INC

Multi-modal emotion fusion analysis method and system

The invention discloses a multi-modal emotion fusion analysis method and system, and the method comprises the steps: carrying out the feature extraction of multi-modal emotion data modal by modal through a feature extraction module, and generating a text original feature, an audio original feature and a visual original feature; performing cross-modal alignment interactive fusion on the original text features, the original audio features and the original visual features based on a unified semantic alignment module, and constructing collaborative fusion features; performing mode and channel double-layer dynamic fusion optimization by adopting a dynamic fusion regulation and control module according to the text original feature, the audio original feature, the visual original feature and the collaborative fusion feature, and determining a unified fusion feature; performing hierarchical residual semantic gating enhancement based on the unified fusion features according to a high-order semantic abstraction module to generate semantic enhancement features; and inputting the semantic enhancement features into an emotion prediction module, and outputting an emotion analysis result. Based on the above scheme, a more stable, accurate and reliable emotion recognition result can be provided.
Owner:GUANGDONG UNIV OF TECH

System and method for contextual analysis and metadata database generation for user-specific speech patterns

A system for contextual analysis and metadata database generation for user-specific speech patterns is disclosed. The system accesses a speech signal of a user and identifies the user based on the voice print associated with the user. The system splits the speech signal into a first set of audio frames, where each audio frame comprises an utterance of one or more words. The system determines a context associated with each word. In response, the system detects a context change between a first text and a second text. The system generates a contextually split set of frames by splitting the speech signal into a second set of audio frames according to the detected context changes.
Owner:BANK OF AMERICA CORP

Bluetooth communication intelligent speech translation method and system based on multi-mode enhancement

The invention relates to the technical field of artificial intelligence, and discloses a Bluetooth communication intelligent speech translation method and system based on multi-mode enhancement, and the method comprises the steps: collecting a multi-channel audio signal through a built-in multi-microphone array of a Bluetooth device, carrying out the dynamic direction self-adaptive beam forming of the multi-channel audio signal, and carrying out the self-adaptive beam forming of the multi-channel audio signal; extracting a Mel spectrogram feature of the direction enhancement signal, identifying lip regions of a plurality of candidate speakers in each frame of real-time speaking video captured by a camera, performing time sequence convolution on the lip regions to obtain a lip movement time sequence embedded vector, calculating a correlation score with the Mel spectrogram feature, separating the direction enhancement signal, and obtaining a lip movement time sequence embedded vector; and performing text transcription and conversion on the high-confidence separation voice to obtain a translation language text, and sending the synthesized target translation voice to a preset mobile terminal through the Bluetooth device to obtain a target translation result. According to the method, the real-time performance and accuracy of speech translation are improved in a multi-person scene, far-field speech, noise interference and accent difference.
Owner:SHENZHEN DIE MICRO SEMICON CO LTD