Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

19615 results about "Speech sound" patented technology

Speech sound. noun. 1 : any one of the smallest recurrent recognizably same constituents of spoken language produced by movement or movement and configuration of a varying number of the organs of speech in an act of ear-directed communication.

Real-time virtual reality scene system based on natural language description using multimodal artificial intelligence

A real-time system for the multimodal generation of virtual reality scenes based on artificial intelligence for the creation of immersive three-dimensional environments from natural language narratives, consisting of: a speech capture module configured to continuously record a user's spoken narrative via one or more directional microphones, preprocesses the captured signal by noise reduction and temporal alignment, and outputs a digital speech stream; A speech-to-text processing unit that is operationally coupled to the speech capture module and configured for real-time speech recognition using a continuous neural transformer model. The unit is trained to transcribe natural language utterances into structured text data while maintaining contextual continuity throughout the evolving narrative. a semantic interpretation processing unit that is communicatively linked to the speech recognition unit and configured to perform natural language understanding techniques to extract contextual entities, spatial references, temporal relationships, and object attributes from the transcribed narrative; the engine includes a large language model that is fine-tuned for spatial reasoning tasks; a scene graph generation module configured to transform the interpreted semantic data into a structured, hierarchical representation that defines nodes for identified entities and edges for corresponding relationships, with each node associated with metadata describing geometry, position, orientation, texture, and linking attributes between objects; a multimodal image-language model processor coupled with the scene graph generation module, wherein the processor is configured to retrieve, adapt, or synthesize appropriate three-dimensional elements from a pre-trained visual-lexical embedding space and align these elements with their semantic and spatial definitions derived from the scene graph; a scene assembly and rendering controller configured to create a cohesive virtual scene from the aligned assets, perform real-time rendering using a GPU-accelerated ray tracing pipeline, and produce a stereoscopic visual output that corresponds to the evolving narrative; A head-mounted virtual reality visualization device connected to the rendering engine and configured to display the generated immersive environment to the user in real time. The device features motion sensors and inside-out tracking cameras to detect head and body movements, dynamically updating viewing angles and perspective within the rendered scene; and a bidirectional feedback module integrated into the head-mounted device and connected to the semantic interpretation processing unit; the module is configured to interpret corrective commands, gestures, or supplementary comments from the user to refine or modify specific scene elements without interrupting the real-time visualization; The system continuously updates the virtual scene as the narrative develops, ensuring temporal synchronization between speech input and rendered output below a defined latency threshold, thus enabling a natural, dialogic construction of complex three-dimensional virtual environments.
Owner:GOUNDER MOHAN SELLAPPA DR BENGALURU +3

Text prediction-based large-model real-time voice text intention recognition method and system

The invention discloses a large-model real-time voice text intention recognition method and system based on text prediction, and the method comprises the steps: obtaining the real-time voice data of a user, carrying out the real-time voice recognition processing through a streaming voice recognition interface, and obtaining a part of transcriptional text; inputting the partial transcription text into a mask language model for text prediction, and generating a plurality of high-credibility complete sentence candidates; based on the complete sentence candidates, the complete sentence candidates are input into a large language model in parallel for intention recognition, a corresponding intention result is obtained, and a mapping relation between the candidate sentences and the intention recognition result is established; and obtaining a sentence completely expressed by the user, calculating the similarity between the complete actual sentence and a plurality of high-credibility complete sentence candidates through a multi-level text similarity algorithm, selecting the candidate sentence with the highest similarity score, and directly obtaining a corresponding final intention recognition result based on the mapping relationship. The objective of the invention is to solve the technical problem of high response delay of an existing voice intention recognition system.
Owner:BEIJING YULORE INNOVATION TECH

Decision-making method and device based on multi-modal semantic alignment, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a decision-making method, device, equipment and medium based on multi-modal semantic alignment. Executing cross-modal alignment by taking the voice semantic map as a reference to generate associated information, fusing the voice features, the visual features, the action features and the associated information to generate a fusion feature vector, inputting a decision network to generate a decision feature vector and generate a task execution instruction, obtaining execution feedback information of the task execution instruction, and updating the decision network. According to the method, input is dominated by voice instructions, visual features, action features and semantic map depth alignment and fusion are combined, input naturalness and multi-modal data analysis and decision-making efficiency are improved, and interaction adaptability and decision-making accuracy of the model in a complex scene are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Face-translator: end-to-end system for speech-translated lip-synchronized and voice preserving video generation

A neural end-to-end system is provided for the face and voice preserving translation of videos. The system is a pipeline of multiple models that produces a video of the original speaker speaking in the target language with modified lip movement to match the target speech, while preserving emphases and prosody of the original speech, and voice characteristics of the original speaker. The pipeline starts with automatic speech recognition including emphasis detection, followed by the translation model. The translated text is then synthesized by a Text-to-Speech model that recreates the original emphases in the target sentence. The resulting synthetic speech is then converted back to the original speakers' voice using a voice conversion model. Finally, to synchronize the lips of the speaker with the translated audio, a generative model generates frames of adapted lip movements which are combined with the audio to produce the final output. The disclosure further describes several use-cases and configurations that apply these techniques to video conferencing, dubbing, low-bandwidth transmission, speech enhancement and assistive technology for the hearing impaired.
Owner:WAIBEL ALEXANDER

Comprehensive AI-enabled systems for immersive voice, companion, and augmented / virtual reality interaction solutions

A computer-implemented method for operating an artificial intelligence voice agent system includes receiving voice input through communication channels; analyzing converted text through natural language processing (NLP) pipelines implementing intent recognition and sentiment analysis detecting emotional cues using a multimodal large language model (LLM); generating response content using machine learning models trained on domain-specific corpora; converting generated responses to synthetic speech through text-to-speech (TTS) engines; integrating with a customer relationship management (CRM) platforms or an enterprise resource planning (ERP) database; and implementing continuous learning by updating language understanding models using conversation logs, voice recognition parameters based on user feedback, and response generation patterns. One implementation is a computer-implemented system and method that operates a suite of intelligent interactive devices and platforms including an artificial intelligence voice agent, enhanced communication platforms, an intimacy companion system, and augmented / virtual reality eyeglasses. Further, one implementation includes AR / VR eyeglasses that project visual content onto interchangeable lenses or directly onto the user's retina via laser-based retinal projection, provide prescription adjustments, incorporate ear-mounted sensors for monitoring physiological parameters like heart rate, oxygen saturation, and blood pressure, and utilize wireless data transmission, onboard environmental sensing, and remote calibration, all designed to offer dynamically adaptive, secure, and context-aware interactions across communication, personal assistance, health monitoring, and immersive augmented or virtual reality environments.
Owner:TRAN BAO

Personalized and dynamic text to speech voice cloning using incompletely trained text to speech models

Systems and methods are provided for machine learning models configured as zero-shot personalized text-to-speech models which comprise a feature extractor, a speaker encoder, and a text-to-speech module. The feature extractor is configured to extract acoustic features and prosodic features from new target reference speech associated with the new target speaker. The speaker encoder is configured to generate a speaker embedding corresponding to the new target speaker based on the acoustic features extracted from the new target reference speech. The text-to-speech module is configured to generate the personalized voice corresponding for the new target speaker based on the speaker embedding and the prosodic features extracted from the new target reference speech without applying the text-to-speech module on new labeled training data associated with the new target speaker.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

CAD automatic generation system and method based on intelligent model selection and application

The invention discloses a CAD automatic generation system and method based on intelligent model selection and application, and aims at achieving automatic modeling under the multi-modal design requirement. The system comprises a user interaction module for receiving multi-modal input such as natural language, sketch and voice; the intelligent demand analysis module is used for combining an industrial large language model and a product knowledge graph, combining semantic analysis and generating a structured demand; the intelligent model selection calculation module is used for matching the optimal parameter combination and the component list based on a multi-objective optimization algorithm; the CAD automatic generation module calls a parametric modeling engine to generate an editable three-dimensional model; the constraint solving module is used for processing hard constraints and soft constraints in real time and dynamically adjusting model parameters; and the model output and interaction module feeds back a design state and supports user iteration. The system realizes full-process automation from the design intention to the CAD model, improves the design efficiency and accuracy, and is suitable for the fields of mechanical design, intelligent manufacturing and the like.
Owner:HOFMANN (BEIJING) ENG TECH CO LTD

Intelligent glasses AI voice interaction method

The invention provides a smart glasses AI voice interaction method, which comprises the following steps: simultaneously acquiring a user gesture image and a voice signal, acquiring a user historical interaction record, extracting key point coordinates and pointing direction information of the gesture image, acquiring hand distance information, generating a gesture motion path, simultaneously extracting frequency spectrum information and an intonation peak value of the voice signal, and generating a gesture motion path; forming a voice beat sequence; extracting a spatial semantic mode of the gesture and voice fusion data, and recognizing a core object of a user pointing instruction according to pointing coordinates of a gesture motion path and an intonation peak value of a voice beat sequence; after the core object pointing to the instruction is recognized, the user intention is determined in combination with the spatial semantic mode and the historical interaction record of the user; and the complete intention analysis result is output to the intelligent glasses display module to execute corresponding operation, and is fed back to the acquisition module to adjust the next capture parameters including the acquisition frequency and the recognition sensitivity, so that the response speed and the accuracy are improved.
Owner:SHENZHEN YAWELL LNTELLIGENT TECH CO LTD

Operation intention recognition method, system and equipment based on multi-modal fusion and medium

The invention relates to the technical field of data processing, and particularly provides an operation intention recognition method, system and device based on multi-mode fusion and a medium, and the method comprises the steps: synchronously collecting interaction data of at least two modes of a user, the modes comprising at least two of gestures, voice and eye gaze; carrying out alignment processing on the interaction data, wherein the alignment processing comprises time synchronization and space mapping to a unified coordinate system; recognizing structured semantic information from each piece of aligned modal data, wherein the structured semantic information comprises a gesture type, a voice text and a fixation point coordinate; and based on a preset semantic rule and context memory, performing semantic association and anaphora resolution on the structured semantic information to obtain an operation intention. The method effectively overcomes the inherent defects of unnatural single-mode interaction, easy ambiguity and poor fault tolerance.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Multi-unmanned aerial vehicle task scheduling method and system with dependence perception and feedback mechanism

The invention discloses a multi-UAV (unmanned aerial vehicle) task scheduling method and system with a dependency perception and feedback mechanism, and the method comprises the steps: enabling a commander to input a task demand in a voice or text form, inputting the task into a large language model based on a Python prompt template in combination with environment information and UAV capability configuration, and enabling the large language model to perform task scheduling; and completing subtask disassembly and dependency modeling of the natural language instruction. The method comprises the following steps: establishing a sub-task dependency graph, and determining a sequential relationship and execution logic between tasks; in the aspect of task scheduling, capability vector modeling is carried out on all online unmanned aerial vehicles, and an optimal unmanned aerial vehicle is selected or a multi-vehicle alliance is automatically constructed to execute a task based on a vector matching degree between task skill requirements and unmanned aerial vehicle capabilities. In the task execution process, task state information is collected in real time, and all feedback information is uploaded to the cloud control center for state judgment and abnormity recognition. When the system detects an abnormal condition, task reconstruction, alliance recombination and scheduling graph repair are automatically carried out, and closed-loop adjustment of the task is completed.
Owner:HOHAI UNIV +1

Long video multi-modal understanding and question-answering method and system based on large model and retrieval enhancement generation

The invention discloses a long video multi-modal understanding and question-answering method and system based on large model and retrieval enhancement generation. The method comprises the following steps: 1) a multi-modal feature extraction module; 2) a multi-modal synchronization and alignment mechanism; 3) constructing a structured memory pool; 4) querying a drive generation mechanism; 5) incremental updating and memory compression strategy; and 6) unifying the multi-modal representation space. The invention provides a long video multi-mode understanding method fusing a large language model and retrieval enhancement generation, and aims to break through the limitation of a traditional method in the aspects of single-mode processing and semantic fragmentation. According to the method, video image features are extracted through a visual model (such as YOLO and ViT), voice transcription and environment voice description are obtained in combination with an audio model (such as Whisper and Qwen-Audio), and unified coding of vision, voice and audio in a long video is achieved. Then, a structured memory pool is constructed through semantic consistency segmentation and timestamp alignment technologies to store time slice data of different modalities.
Owner:GUANGZHOU BINGO SOFTWARE +1

Data processing method and apparatus, electronic device, computer readable storage medium and computer program product

The present application provides a data processing method and apparatus, an electronic device, a computer readable storage medium and a computer program product. The method comprises: acquiring historical interaction information and a predicted interaction text corresponding to the historical interaction information; extracting a first acoustic feature and a first semantic feature of the historical interaction information, and extracting a second semantic feature of the predicted interaction text; performing fusion mapping on the basis of the first acoustic feature, the first semantic feature and the second semantic feature to obtain a first paralanguage feature; denoising initial noise on the basis of the second semantic feature and the first paralanguage feature to obtain a second acoustic feature of the predicted interaction text; and on the basis of the second acoustic feature, generating a voice signal corresponding to the predicted interaction text.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Auditory user interfaces and associated systems, methods, devices, and non-transitory computer-readable media

An auditory operating system designed to facilitate context-aware, audio-based user interactions, particularly with artificial intelligence agents or applications. An auditory operating system shell serves as the primary interface, managing and coordinating multiple specialized agents that handle specific domains like music streaming, scheduling, or home automation. Using natural language processing, the auditory operating system shell identifies the appropriate agent or application for a user's command or query and ensures task execution and context preservation across interactions. The auditory operating system shell enforces privacy and stability by controlling agents' and applications' access to data and system privileges. The auditory operating system shell also supports dynamic context management, enabling seamless handoffs between agents when user requests span multiple domains. This auditory operating system shell may reduce the need for users to memorize specific wake words or commands, as the auditory operating system shell may determine the user's intent from speech or contextual cues.
Owner:IYO INC

Intelligent voice recognition and natural language interaction method based on quadruped robot

The invention discloses an intelligent voice recognition and natural language interaction method based on a quadruped robot, and the method comprises the following steps: S1, collecting a user voice instruction, generating a standardized voice text, and extracting a semantic keyword set; s2, collecting multi-source sensing data of the quadruped robot and generating a structured state data tensor; s3, constructing a multi-modal collaborative modeling mechanism, and generating a multi-modal joint semantic embedding vector; s4, constructing a semantic map based on semantic embedding and generating an action chain plan structure; s5, executing each sub-action in the action chain, and performing path analysis and execution monitoring; s6, storing the interaction task as a multi-modal semantic behavior memory unit; and S7, performing similarity retrieval based on current semantic input and historical memory to realize behavior migration and action chain multiplexing. The method has the advantages of accurate semantic understanding, intelligent interaction response, high behavior migration capability and the like.
Owner:山东浪潮数据库技术有限公司

Multi-agent cooperation strategy generation method and device, equipment and medium

The invention relates to a multi-agent cooperation strategy generation method, device and equipment and a medium, and the method comprises the steps: separating acoustic spectrum features and text semantic features of conference voice through environment perception processing, solving a cross-modal information conflict problem, and generating an accurate semantic understanding result; identifying the essence of the problem based on task analysis, associating the responsibility field, and constructing a classifiable problem point set; calling an agent capability library to dynamically match problem requirements, and generating a candidate agent list; quantifying a problem influence range and a decision time limit through weighted emergency scores, and generating a priority-sorted agent sequence; screening and confirming a core problem point and a primary agent; and finally, generating an executable cooperation scheme through multi-agent collaborative optimization. According to the method, the problems of incomplete feature extraction, task allocation delay and resource conflict in the prior art are solved, and the operability and decision-making efficiency of a cooperation strategy are remarkably improved.
Owner:SHAOGUAN XINGCHENG NETWORK TECH CO LTD

Cross-cultural customer service dialogue quality automatic evaluation method in combination with sentiment analysis

The invention discloses a cross-cultural customer service dialogue quality automatic evaluation method in combination with sentiment analysis, and relates to the technical field of natural language processing, and the method comprises the steps: carrying out the alignment of voice and text based on a transmission matrix in real time, extracting a speech, a metaphor and polarity, and generating a speech tag; constructing an emotion channel and a polite channel, and fusing expression and shielding intensity through sharing attention; comparing and aligning with the same language prototype in a regional culture baseline library to obtain a calibration representation and updating a language offset record table; the potential upgrading probability is represented and recurred according to round aggregation calibration, and a risk vector and a high-risk position are formed; fusing risk and business indexes by a capacity integral kernel, outputting a comprehensive quality score, and giving factors and round attributions; sample recovery is triggered according to score and feedback difference, a micro-weight training data set is constructed, gradient increment training is carried out under low-rank adaptation, and cross-language consistency, early recognition of upgrading risks and interpretable evaluation are achieved through a closed loop.
Owner:LANZHOU INST OF TECH

Real-time anti-fraud monitoring system and method based on behavior reasoning and sentiment analysis

The invention relates to the technical field of artificial intelligence, in particular to a real-time anti-fraud monitoring system and method based on behavior reasoning and sentiment analysis, and the system comprises a multi-modal data collection unit, an edge preprocessing unit, a feature fusion and behavior reasoning unit, a large language model context reasoning unit, a risk assessment and decision unit, and an intervention execution unit. A log recording and federal incremental learning unit; the method has the beneficial effects that the traditional isolated single-mode detection is evolved into an emotion and behavior dual-channel collaborative multi-mode recognition system through millisecond-level coaxial alignment of voice, video and user operation logs; the robustness of dialect, noise and expression shielding is greatly improved through the multi-modal fusion model, so that the cross-scene recognition accuracy is improved by nearly three percent compared with that of a traditional single-voice scheme, and high-sensitivity capture of hidden and emotion control type fraud is truly achieved.
Owner:INSPUR TIANYUAN COMM INFORMATION SYST CO LTD

Precise international communication digital human real-time dialogue method fused with multi-modal technology

The invention discloses an accurate international communication digital human real-time dialogue method fused with a multi-modal technology. The method comprises the following steps: S1, constructing a digital human image and tone; s2, propagation content generation and problem guidance; s3, semantic analysis and intention clarification based on the real-time voice dialogue; s4, geographic preference modeling and path planning; s5, cross-context propagation content generated based on retrieval enhancement is generated; s6, visually displaying the propagation content; and S7, carrying out digital human-driven multi-language propagation content real-time output and feedback closed-loop optimization. Through accurate utterance expression analysis, accurate international propagation problem recommendation is provided, cross-context propagation content generation based on semantic understanding is realized, digital people with voice features and visual images are constructed, real-time dialogue interaction of users is realized, and user experience is improved. The system can carry out geographic modeling according to the region where the accurate problem is located, language preference and propagation object culture characteristics, and differential propagation path planning is achieved.
Owner:HUNAN NORMAL UNIVERSITY

Intelligent conference summary automatic generation method based on voice recognition and large model

The invention discloses an intelligent conference summary automatic generation method based on voice recognition and a large model. The method comprises the following steps: S1, executing voice activity detection operation on an audio data stream; s2, extracting embedding vectors of continuous and effective voice segments, and generating a voice segment set to which a spokesman belongs; s3, inputting the voice fragment set to which the spokesman belongs into an improved Whisper model, fusing a Speaker-Aware attention mechanism and a connection time sequence classification auxiliary path, and outputting a conference transcription text sequence set; s4, inputting the processed structured dialogue format into a GPT-4 large language model, and generating a conference semantic representation sequence; s5, generating a conference summary first draft text according to a preset summary generation template; and S6, performing formatting output operation on the conference summary first draft text. The conference semantic elements can be automatically extracted, the structured summary text can be generated, and the method is suitable for efficient conference recording and task tracking in government affair office, enterprise collaboration, academic discussion and other scenes.
Owner:JIANGSU GUOHUACHENJIAGANG POWER GENERATION CO LTD

Audio and video player control method based on voice instruction

The invention relates to the technical field of audio and video control, and discloses an audio and video player control method based on a voice instruction. The method comprises the steps that an original voice instruction stream of a user is collected, the instruction stream comprises a time domain audio signal sequence, an environment noise spectrum and user pronunciation characteristic parameters, and voice information can be comprehensively captured; multi-modal instruction analysis processing is carried out on the original voice instruction stream, a structured control instruction set containing acoustic control intention identification, semantic operation object description and context correlation parameters is generated, and the analysis precision is improved; then executing player state adaptation based on the set, generating a dynamic control response sequence containing an equipment state adjustment command, a media content positioning parameter and an interface interaction logic identifier, driving a player to execute a multi-dimensional control operation and generating real-time play control effect feedback data; and finally, multi-modal analysis parameters are optimized according to feedback data, a self-adaptive instruction analysis strategy is generated, and the control experience of a user on the audio and video player is optimized.
Owner:ONWAY TECH LTD

Short video intelligent editing method and system based on multi-modal analysis

The invention discloses a short video intelligent editing method and system based on multi-modal analysis, and relates to the technical field of video editing. The method is used for improving editing efficiency and visual experience and comprises the following steps: extracting lip motion features of a character, visual saliency features of a commodity and a voice emotion intensity value from a target short video stream to form multi-modal time sequence data; afterwards, the voice stream is recorded, a product keyword timestamp is extracted, the alignment degree is calculated through dynamic time warping in combination with a visual saliency peak value, and a preliminary editing point set is generated through weighted evaluation in combination with an emotional intensity value; constructing an editing decision optimization model based on deep reinforcement learning, taking the multi-modal features as state input, adjusting the retention probability of editing points through a joint reward function, and selecting an optimal transition mode; and the lip movement and voice synchronization error before and after the editing point and the emotional and visual continuity of the transition section are analyzed, the discontinuous region is smoothed, and the edited finished product is output, so that precise short video intelligent editing is realized.
Owner:ANHUI XINGBANG DIGITAL TECHNOLOGY GROUP CO LTD

Training and speech generation methods and apparatuses for speech generation model, electronic device, computer-readable storage medium, and computer program product

The present application provides training and speech generation methods and apparatuses for a speech generation model, an electronic device, a computer-readable storage medium, and a computer program product. The method comprises: obtaining a first speech generation model; obtaining sample data of a plurality of modalities; on the basis of a prompt image sequence and speech text, respectively calling a plurality of encoders to perform encoding, so as to obtain a multi-modal encoding vector sequence; on the basis of the multi-modal encoding vector sequence, calling a decoder to perform decoding, so as to obtain decoded text; determining a probability distribution for the decoded text and the sample data of the plurality of modalities, and determining a target loss on the basis of the probability distribution; and on the basis of the target loss, updating parameters of the decoder and at least one of the encoders, wherein the updated decoder and the plurality of updated encoders are configured to form a second speech generation model, and the second speech generation model is used to generate target speech text corresponding to a prompt image sequence of a target object.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Beamforming using image data

A device capable of using image data for purposes of determining a location of a user and audio beam selection to isolate audio in the direction of the user. The beamforming / beam-steering may occur after determining the user's location in order to conserve computing resources that would otherwise have been spent determining beams for non-desired directions. The beamformed audio may be used for speech processing, a communication session involving the device, or other purposes.
Owner:AMAZON TECH INC

Intelligent hardware dynamic interaction system based on voice semantic fusion and multi-mode perception

The invention relates to the field of intelligent interaction, and discloses an intelligent hardware dynamic interaction system based on voice semantic fusion and multi-modal perception, which comprises the following steps of: constructing a context model of continuous operation by collecting continuous voice instructions, gesture actions and expression information of a user; semantic analysis and feature fusion are carried out on currently collected voice, gesture and expression features, meanwhile, credibility indexes of all modes are calculated through a weighting or deep learning model, weighting correction is carried out on a fusion result, a real-time feedback algorithm is adopted for weight adjustment for continuous optimization, the next operation intention of a user is predicted through deep learning, and the user experience is improved. And in combination with historical interaction data, online feedback and prediction errors, context management, modal weight and intention prediction strategies are adaptively optimized, and the updated strategies are used for next-round context acquisition and multi-modal fusion. The method has the advantage of improving the recognition accuracy in the continuous interaction scene.
Owner:华欧同惠(苏州)科技有限公司

Speech recognition method and related device

ActiveCN114360510AImprove fault tolerancePrecise Syllable Probability DistributionSpeech recognitionSyllableAcoustic model
The embodiment of the invention discloses a speech recognition method and a related device, and at least relates to a speech recognition technology in artificial intelligence, speech data to be recognized are used as input data of a time delay neural network in an acoustic model, and an output layer of the time delay neural network comprises acoustic modeling units corresponding to a plurality of syllables respectively, so that the speech recognition efficiency is improved. And the syllable probability distribution corresponding to the voice frames included in the voice data can be obtained by taking the syllables as the recognition granularity through the time delay neural network. When syllable recognition is carried out through the output layer, auxiliary judgment can be carried out on the syllables to which the voice frames belong on the basis of pronunciation rules in combination with front and back syllable information of the voice frames, so that more accurate syllable probability distribution is output. Moreover, since the syllables are generally composed of one or more phonemes, the method has higher fault-tolerant capability, not only can more accurately determine the speech recognition result based on the probability distribution of the syllables, but also has low requirements for the quality of the speech data to be recognized, and effectively expands the application scenarios of the speech recognition technology.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Plug-and-play wireless high-definition audio and video transmission method and system

The invention relates to the technical field of wireless high-definition audio and video transmission, and discloses a plug-and-play wireless high-definition audio and video transmission method and system.The method comprises the steps that foreground space texture features and audio and voice segments are obtained, and an initial feature set containing space and semantic features is obtained through semantic correlation analysis; a unified representation vector is output through multi-modal feature fusion, and a dynamic importance score is obtained by determining time sequence consistency, calculating an importance weight and performing normalization; and when the score exceeds a threshold value, segmenting the video frame in real time to determine an attention focus area and optimize a boundary, thereby generating an attention weight matrix, preferentially allocating bandwidth and forming a partition differentiation compression result. And in combination with a network bandwidth state, protection is enhanced for a key stream, the priority and the bit rate are dynamically adjusted, and high-quality audios and videos are output through decoding and recombination. According to the invention, the bandwidth allocation and compression strategy can be dynamically optimized, the transmission quality of a key area is guaranteed when the bandwidth fluctuates, and the audio and video transmission efficiency and experience are improved.
Owner:深圳市翼联网络通讯有限公司

Virtual historical character dialogue method and system with role knowledge and context awareness

The invention discloses a virtual historical character dialogue method and system with role knowledge and context awareness, and relates to the technical field of man-machine interaction, and the method comprises the steps: constructing a multi-level role depth model; when a question of a current user is received, identifying information of a virtual scene where the current user is located, analyzing micro-expressions of the face of the user and voice rhythm characteristics of speech of the user, analyzing an emotional state and an interaction intention of the user based on a multi-modal fusion algorithm, and generating a user state vector; executing a dynamic Prompt construction program, extracting related information from the multi-level role depth model and the user state vector, and generating a structured Prompt; and inputting the structured Prompt into a large language model, generating a reply text conforming to role features based on questions of the current user, and driving a virtual character model. The method solves the problem that in the prior art, virtual historical figures cannot provide real immersion and credible interactive experience with emotional connection.
Owner:BEIJING GROWLIB TECH CO LTD

System for real-time analysis of emotional feedback during motivational presentations

A system for real-time analysis of emotional feedback during motivational speeches, consisting of: a series of multimodal sensors, including at least one visual sensor configured to capture facial expressions of spectators, at least one directional microphone configured to capture the audio responses of the audience, and optionally one or more physiological sensors configured to capture biometric signals from spectators; an edge-based processing unit that is communicatively coupled to the arrangement of multimodal sensors, wherein the edge-based processing unit comprises the following: (a) a feature extraction module configured to extract visual features from captured facial images, acoustic features from voice responses, and physiological features from biometric signals; (b) an emotion inference machine configured to process the features using a deep learning-based emotion recognition model comprising a convolutional neural network (CNN) for classifying facial expressions, a recurrent neural network (RNN) for classifying voice emotions, and a multimodal late fusion layer configured to compute a composite emotion state vector representing the aggregated emotions of the audience; (c) a timestamp and speech alignment module configured to correlate the calculated composite emotion state vector with segmented portions of a live motivational speech based on real-time speech-to-text transcription and semantic analysis; and (d) a session-based storage unit configured to log time-indexed emotional state vectors and corresponding speech segments for post-event analysis; A speaker feedback interface comprising a portable display or a podium-mounted visualization panel, wherein the interface is configured to display visual indicators of emotional feedback in real time, the indicators being derived from the emotional state vector and including at least emotional trend graphs, threshold alerts, or engagement indices.
Owner:1XL LLC FZ +2

System and Method for Real-Time Identity-Free Personalization Using Fluid Emotional Trait Vectors, Modular Engine Mesh Architecture, Context-Aware Engagement Logic, and Adaptive Goal Mutation

A system and method for real-time, identity-free personalization using deformable emotional trait vectors to dynamically adapt digital and voice-based experiences. Each user session is modeled as a behavioral object known as a Vectra, composed of fluidic traits—such as mass, viscosity, temperature, volatility, and texture—that evolve continuously in response to live behavioral, contextual, environmental, and voice-derived signals. These Vectras traverse a dynamically warped emotional space, the Vectraverse, influenced by ambient conditions including time of day, noise level, inventory urgency, and engagement rhythm. Gravitational pull toward predefined emotional goal attractors modulates system behavior, while a goal mutation engine reclassifies session intent when confidence decays or friction spikes. Outputs include tone modulation, content pacing, offer framing, and gamified reward logic—all executed without storing identity, login credentials, or historical profiles. The system supports modular deployment across voice, screen, signage, mobile, and in-room environments, and integrates with large language models, AI agents, and third-party personalization stacks via privacy-safe APIs and federated learning. Designed for zero-ID personalization, the platform enables emotionally intelligent, context-aware engagement across any surface or session.
Owner:GINSBERG JUSTIN

Digital human interaction system and method based on multi-modal emotion recognition

ActiveCN121116129ASemantic analysisSpeech analysisInteractive modelingData stream
The embodiment of the invention provides a digital human interaction system and method based on multi-modal emotion recognition, and belongs to the technical field of digital human interaction. The system comprises a multi-modal sensing module used for collecting multi-modal data and preprocessing the multi-modal data to generate a standardized data stream; the cross-modal fusion and emotion recognition module is used for carrying out interactive modeling on the multi-modal features and outputting a current emotion label and emotion intensity; the reaction planning module is used for generating a composite reaction strategy; and the digital human rendering module is used for mapping the composite reaction strategy into control signals corresponding to the voice, the facial expression and the action respectively, and driving a digital human to execute corresponding voice output, facial expression change and limb action through the control signals so as to realize interaction. According to the method, multi-modal data are deeply fused through the cross-modal graph neural network and comparative learning, the weight is dynamically adjusted in combination with the modal confidence, and the emotion recognition accuracy and robustness are improved.
Owner:XIAODUO INTELLIGENT TECH (BEIJING) CO LTD