Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

25018 results about "Speech sound" patented technology

Speech sound. noun. 1 : any one of the smallest recurrent recognizably same constituents of spoken language produced by movement or movement and configuration of a varying number of the organs of speech in an act of ear-directed communication.

Transform-based cross-modal fusion multi-modal emotion recognition method

The invention discloses a Transform-based cross-modal fusion multi-modal emotion recognition method and device, which are used for solving the problems of modal isomerism, difficulty in time alignment and insufficient dynamic emotion modeling in a multi-modal emotion recognition task, and the method takes the accuracy and robustness of emotion recognition as performance evaluation indexes. Firstly, feature information of three modes of vision, voice and text is obtained, feature extraction is performed on each mode through a deep learning model, then features of different modes are fused by using a cross-mode Transform module, and a complex dependency relationship between the modes is dynamically modeled through a multi-head self-attention mechanism, so that more accurate emotion recognition is realized, and the emotion recognition efficiency is improved. And finally, performing emotion prediction on the fused features based on time sequence modeling and an emotion classification module. According to the method, the problems of modal isomerism, difficulty in time alignment and insufficient dynamic emotion modeling in multi-modal emotion recognition can be effectively solved.
Owner:SOUTHEAST UNIV

Customer data processing and insight system based on large language model

The invention belongs to the technical field of artificial intelligence and big data, and discloses a customer data processing and insight system based on a big language model. The system is composed of a multi-source data access module, a data preprocessing and label fusion module, a large language model semantic understanding module, a knowledge enhancement and semantic linkage module, an insight generation and visualization module, an intelligent strategy output module and a feedback learning and self-optimization module. According to the method, multi-source heterogeneous data such as texts, voices and structured behaviors are integrated, and the deep semantic analysis capability of a large language model is combined, so that global modeling of customer behaviors and intentions is realized; a multi-modal synchronous acquisition and standardization mechanism eliminates data format barriers, and a dynamic label mechanism adapts to context changes, so that the system can capture deep semantic association in customer expression, and compared with a traditional keyword matching method, the semantic understanding accuracy is improved by more than 40%, and a more complete data base is provided for insight generation.
Owner:SICHUAN JUFUREN TECHNOLOGY CO LTD

Vision generation method and device based on semantic association modeling, equipment and medium

The invention relates to the technical field of voice semantics, can be applied to business scenes of financial science and technology, medical health, poster design and the like, and discloses a visual sense generation method and device based on semantic association modeling, equipment and a medium. Generating a demand text containing theme and style parameters; semantic features in the demand text are extracted, semantic association weights are constructed, and element layout coordinates are optimized in combination with spatial distribution constraints; and encoding the layout information into a control matrix, fusing the control matrix with the initial noise, adjusting a noise reduction process through an encoding and decoding network, and generating target visual content highly matched with the semantic meaning of the user instruction. According to the method, the layout optimization function is constructed, the diffusion model is guided to focus the semantic salient region in space, language model output and the visual generation process are closely combined, structured response and space mapping of user semantic requirements are achieved, and the expression consistency and personalized adaptation capacity of visual content generation are improved.
Owner:SHENZHEN PINGAN COMM TECH CO LTD

Intelligent sound box voice processing method and system based on artificial intelligence

The invention provides an intelligent sound box voice processing method and system based on artificial intelligence, and the method comprises the steps: obtaining audio data and mouth shape video data, carrying out the processing of the audio data and the mouth shape video data, and carrying out the multi-modal feature fusion, and obtaining a fusion feature; performing bimodal voice activity detection on the fusion features to obtain effective voice data; performing context sensing recognition of audio and video fusion on the effective voice data to obtain a first text; constructing a user feature model, and performing semantic understanding on the text based on the model to obtain an understanding result; performing intention recognition and slot filling based on the understanding result to obtain user intention and key information; generating a response strategy in combination with the user intention, the key information and the environment perception data; generating response voice according to the response strategy; and monitoring feedback information of the user to the response voice in real time, and updating the user feature model and the response strategy evaluation model based on feedback. According to the scheme, the voice can be recognized more accurately, and the safety and robustness of the system are enhanced.
Owner:SHENZHEN ZHANDIAN SMART TECH CO LTD

Voice data interaction feedback control processing method based on large language model

The invention relates to the technical field of large language models, and discloses a voice data interaction feedback control processing method based on a large language model. The method comprises the following steps: acquiring an industrial voice instruction through an AMR main controller, and obtaining a standardized vector through industrial lexicon matching and intention classification; inputting an industrial large language model for reasoning processing, and generating an AMR execution scheme; carrying out distributed coordination and task allocation on the AMR cluster to form a control instruction sequence; and the motion controller executes monitoring, processes exceptions and outputs a feedback strategy. Accurate classification and semantic understanding of complex industrial instructions are achieved, the processing capacity of professional knowledge in the industrial field is improved, and meanwhile high-precision real-time state monitoring is achieved.
Owner:TIANJIN HONGHUANG TECH CO LTD

Real-time virtual reality scene system based on natural language description using multimodal artificial intelligence

A real-time system for the multimodal generation of virtual reality scenes based on artificial intelligence for the creation of immersive three-dimensional environments from natural language narratives, consisting of: a speech capture module configured to continuously record a user's spoken narrative via one or more directional microphones, preprocesses the captured signal by noise reduction and temporal alignment, and outputs a digital speech stream; A speech-to-text processing unit that is operationally coupled to the speech capture module and configured for real-time speech recognition using a continuous neural transformer model. The unit is trained to transcribe natural language utterances into structured text data while maintaining contextual continuity throughout the evolving narrative. a semantic interpretation processing unit that is communicatively linked to the speech recognition unit and configured to perform natural language understanding techniques to extract contextual entities, spatial references, temporal relationships, and object attributes from the transcribed narrative; the engine includes a large language model that is fine-tuned for spatial reasoning tasks; a scene graph generation module configured to transform the interpreted semantic data into a structured, hierarchical representation that defines nodes for identified entities and edges for corresponding relationships, with each node associated with metadata describing geometry, position, orientation, texture, and linking attributes between objects; a multimodal image-language model processor coupled with the scene graph generation module, wherein the processor is configured to retrieve, adapt, or synthesize appropriate three-dimensional elements from a pre-trained visual-lexical embedding space and align these elements with their semantic and spatial definitions derived from the scene graph; a scene assembly and rendering controller configured to create a cohesive virtual scene from the aligned assets, perform real-time rendering using a GPU-accelerated ray tracing pipeline, and produce a stereoscopic visual output that corresponds to the evolving narrative; A head-mounted virtual reality visualization device connected to the rendering engine and configured to display the generated immersive environment to the user in real time. The device features motion sensors and inside-out tracking cameras to detect head and body movements, dynamically updating viewing angles and perspective within the rendered scene; and a bidirectional feedback module integrated into the head-mounted device and connected to the semantic interpretation processing unit; the module is configured to interpret corrective commands, gestures, or supplementary comments from the user to refine or modify specific scene elements without interrupting the real-time visualization; The system continuously updates the virtual scene as the narrative develops, ensuring temporal synchronization between speech input and rendered output below a defined latency threshold, thus enabling a natural, dialogic construction of complex three-dimensional virtual environments.
Owner:GOUNDER MOHAN SELLAPPA DR BENGALURU +3

Marketing video auditing method based on AI

The invention provides an AI-based marketing video auditing method, and relates to the technical field of AI marketing video auditing, and the method comprises the steps: obtaining a multi-modal data original structure set, and extracting image semantic features, voice expression features, text semantic features and scene label information, and obtaining an image semantic feature set, a visual rhythm feature set, a voice expression feature set, a voice and picture synchronous association vector structure, a text semantic feature set and a subtitle semantic and image main body linkage relation graph. By constructing an image semantic feature set, a voice expression feature set, a text semantic feature set and a visual rhythm feature set and fusing the image semantic feature set, the voice expression feature set, the text semantic feature set and the visual rhythm feature set into a multi-modal content fusion feature tensor, unified modeling of an AI marketing video at visual, auditory and semantic levels can be realized; and subsequent microscopic consistency detection, compliance knowledge graph and emotion semantic conflict identification are effectively performed, so that full-link risk perception and accurate auditing of video contents are realized.
Owner:SHANGHAI WANGMAI INFORMATION TECH GRP CO LTD

Voice intention recognition method and device, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a voice intention recognition method, device and equipment and a medium, and the method comprises the steps: obtaining a to-be-processed voice signal, carrying out the voice activity detection processing of the voice signal, dividing the voice signal into a plurality of voice segments, analyzing semantic contents of the plurality of voice segments, determining semantic correlation information of each voice segment, analyzing sound source attributes of the plurality of voice segments, determining sound field type information of each voice segment, screening out a target voice segment from the plurality of voice segments according to the semantic correlation information and the sound field type information, and executing intention recognition processing based on the target voice segment to generate an intention recognition result. According to the invention, through a dual analysis mechanism of semantic correlation information and sound field type information, effective screening of voice segments is realized before voice recognition, non-target voice or interference segments are effectively prevented from being sent to an intention recognition model, and the accuracy of a recognition result is improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Multimodal chatbots based on characters

Techniques are disclosed for enabling creators to create multimodal chatbots that are based on or simulate / model characters. The characters may be from audiovisual (AV) media such as films and TV shows or real people. The application leverages a combination of Visual Interpretation AI, Retrieval-Augmented Generation (RAG), Low-Rank Adaptation (LoRA), and function calling to provide rich, interactive experiences. The inferencing performed by an instant multimodal chatbot utilizes base weights, character weights, relationship weights, experience weights as we as environmental inputs. An instant chatbot uses a number of AI models including a video to text model, an image to text model, a sensory to text model, a large language model, a text to video model, a text to image model and a text to voice model. A user can interact with the chatbot in a variety of ways including text, audio and video.
Owner:IGNITE CHANNEL INC

Cabin active recommendation system and method based on knowledge graph and semantic reasoning

The invention discloses a cockpit active recommendation system and method based on a knowledge graph and semantic reasoning, and relates to the technical field of intelligent cockpits. The system receives natural language voice input of a user, executes voice recognition and semantic analysis, extracts user intention, keywords and slot entities, generates structured semantic information, constructs or calls a knowledge graph structure with semantic relation edges in combination with environment context information, and obtains the knowledge graph structure with the semantic relation edges. Semantic path reasoning is carried out based on the path dependence weight and the semantic similarity, a semantic edge label guided graph attention mechanism is introduced to calculate a path consistency score, a candidate recommendation set is generated, the semantic fitting degree and the path score are fused to sort and output recommendation content, and the graph edge weight and the user portrait are updated based on user feedback. According to the method, semantic understanding precision, recommendation path interpretability and system adaptive capacity are improved, and the method is suitable for personalized voice recommendation, man-machine interaction and scene linkage control tasks in an intelligent cockpit.
Owner:RIVOTEK TECH (JIANGSU) CO LTD

Text prediction-based large-model real-time voice text intention recognition method and system

The invention discloses a large-model real-time voice text intention recognition method and system based on text prediction, and the method comprises the steps: obtaining the real-time voice data of a user, carrying out the real-time voice recognition processing through a streaming voice recognition interface, and obtaining a part of transcriptional text; inputting the partial transcription text into a mask language model for text prediction, and generating a plurality of high-credibility complete sentence candidates; based on the complete sentence candidates, the complete sentence candidates are input into a large language model in parallel for intention recognition, a corresponding intention result is obtained, and a mapping relation between the candidate sentences and the intention recognition result is established; and obtaining a sentence completely expressed by the user, calculating the similarity between the complete actual sentence and a plurality of high-credibility complete sentence candidates through a multi-level text similarity algorithm, selecting the candidate sentence with the highest similarity score, and directly obtaining a corresponding final intention recognition result based on the mapping relationship. The objective of the invention is to solve the technical problem of high response delay of an existing voice intention recognition system.
Owner:BEIJING YULORE INNOVATION TECH

Robot anthropomorphic interaction method based on multi-modal emotion recognition and customized portrait generation

The invention discloses a robot anthropomorphic interaction method based on multi-modal emotion recognition and customized portrait generation. The method comprises the following steps: S1, dynamically fusing multi-modal emotions; the method comprises the following steps: S1, synchronously acquiring voice, visual and text signals through a multi-source heterogeneous sensor, capturing a user voice stream by a high-fidelity microphone array, and extracting acoustic characteristics such as intonation and speed, S2, performing cross-modal reasoning; s3, synchronously generating contents; step S4: style migration; step S5, anthropomorphic voice and expression generation; according to the method, man-machine interaction emotion is analyzed and generated by utilizing a large language model and multi-modal information fusion, the singleness of interaction emotion and the deficiency of emotional sharing ability are avoided, a strong emotion interaction characteristic is achieved, the image of the robot is obtained through a generative technology and can be migrated to any image, the limitation that a specific image is independently made is broken through, and the interaction effect of the robot is improved. The advantage that one robot can be suitable for different scenes is achieved.
Owner:JIANGSU YUNMU ZHIZAO TECH CO LTD

Decision-making method and device based on multi-modal semantic alignment, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a decision-making method, device, equipment and medium based on multi-modal semantic alignment. Executing cross-modal alignment by taking the voice semantic map as a reference to generate associated information, fusing the voice features, the visual features, the action features and the associated information to generate a fusion feature vector, inputting a decision network to generate a decision feature vector and generate a task execution instruction, obtaining execution feedback information of the task execution instruction, and updating the decision network. According to the method, input is dominated by voice instructions, visual features, action features and semantic map depth alignment and fusion are combined, input naturalness and multi-modal data analysis and decision-making efficiency are improved, and interaction adaptability and decision-making accuracy of the model in a complex scene are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

AI interaction intelligent module based on hybrid architecture

The invention relates to the technical field of artificial intelligence and Internet of Things, and discloses an AI interaction intelligent module based on a hybrid architecture, comprising a user interaction unit which supports voice, text and image multi-modal input and integrates intention recognition and context understanding algorithms; the data processing unit is used for carrying out structured processing on the electric appliance specification and the historical fault data and constructing a dynamically updated knowledge graph; the hybrid architecture core unit comprises a deep learning subunit for realizing natural language understanding and generation based on a Transform model, and a knowledge reasoning subunit; a fault diagnosis unit; and a feedback optimization unit. According to the method, seamless cooperation of deep learning and symbol logic is realized through a dynamic routing strategy, a high-confidence-coefficient scene generates a response through a Transform model, a medium-confidence-coefficient scene calls a knowledge graph rule for verification, and a low-confidence-coefficient scene supplements information through multiple rounds of interaction, so that the effect of improving balance efficiency and safety is achieved.
Owner:CHENYANG JINYE ZAITIAN TECHNOLOGY CO LTD

Dynamic interaction method based on multi-modal dynamic fusion large model and intelligent agent collaboration

The invention discloses a dynamic interaction method based on cooperation of a multi-modal dynamic fusion large model and an intelligent agent. The method comprises the following steps: performing feature extraction on user voice information to obtain a voice coding vector, a text semantic vector and an emotion feature vector; performing dynamic weight feature fusion on the voice coding vector, the text semantic vector and the emotion feature vector through a multi-modal dynamic fusion large model to obtain a fusion feature vector; inputting the fusion feature vector into an intention-scene coupling network, and identifying to obtain a user intention label; and identifying according to the user behavior log to obtain a user portrait tag, inputting the user intention tag and the user portrait tag into an autonomous decision-making agent, generating a target decision-making action through a lightweight policy network, and then interacting with the user according to the target decision-making action. The intelligent interaction efficiency and accuracy of the customer service system are improved, the interaction experience of the user is also improved, and the method can be widely applied to the technical field of artificial intelligence.
Owner:E SURFING IOT CO LTD

Techniques for determining conversational intent

The present disclosure relates to systems and methods for enhancing the interaction between users and automated agents, such as digital assistants, by employing Large Language Models (LLMs) to infer the intent of spoken language. The invention involves continuously monitoring ambient audio, converting speech to text, and utilizing LLMs to determine whether spoken language is intended for the automated agent. A structured prompt, including the converted text and specific instructions, is sent to the LLM, which is fine-tuned to process domain-specific prompts. The LLM provides a structured output in a standardized format, indicating the user's intent. The system may involve multiple prompts to perform separate tasks, such as identifying intent and generating additional context-specific data. This approach facilitates a more natural and intuitive user experience by eliminating the need for wake words and allowing seamless conversational interaction with virtual assistants across various platforms and devices.
Owner:SNAP INC

Power grid monitoring bionic robot with multi-mode sensing and voice interaction functions

The invention relates to a power grid monitoring bionic robot with multi-mode sensing and voice interaction functions. The acquisition unit acquires sound, images, temperature and power grid equipment surface micro-vibration information, completes data preprocessing and time and space calibration, and generates a sensing data packet. And the identification unit performs cross judgment based on the abrupt change characteristics and correlation of different types of data, identifies electrical, structural or environmental anomalies, and generates abnormal region description information. The construction unit analyzes whether a multi-factor coupling risk exists in combination with a power grid environment type and a historical rule, and generates task description information including processing suggestions and development path prediction. The analysis unit supports speech enhancement and tone extraction in a high-interference environment, and extracts a user interaction intention in combination with task description information. And the control unit jointly judges a behavior correction strategy according to the task description information and the interaction instruction information, generates a corresponding motion control instruction and voice feedback, and realizes autonomous monitoring and interaction response in a complex power grid environment.
Owner:XIANGYANG POWER SUPPLY COMPANY OF STATE GRID HUBEI ELECTRIC POWER

Audio signal authenticity verification method and device, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses an audio signal authenticity verification method, device, equipment and medium, and the method comprises the steps: constructing an original audio text data set, and generating an adversarial sample set, inputting the original audio text data set and the adversarial sample set into an audio detection model for joint training to obtain an audio detection model subjected to adversarial training; the method comprises the steps of obtaining a to-be-detected audio signal and extracting an acoustic feature of the to-be-detected audio signal, obtaining a non-acoustic feature associated with the to-be-detected audio signal, constructing a multi-dimensional feature vector according to the acoustic feature and the non-acoustic feature, inputting the multi-dimensional feature vector into an audio detection model to generate an abnormal index, and executing a hierarchical response operation based on the abnormal index. According to the method, the robustness of the model is enhanced by introducing adversarial sample training, and the multi-dimensional feature vector is constructed by fusing the multi-modal features, so that accurate recognition and hierarchical response to the voice cloning attack are realized.
Owner:PING AN TECH (SHENZHEN) CO LTD

Face-translator: end-to-end system for speech-translated lip-synchronized and voice preserving video generation

A neural end-to-end system is provided for the face and voice preserving translation of videos. The system is a pipeline of multiple models that produces a video of the original speaker speaking in the target language with modified lip movement to match the target speech, while preserving emphases and prosody of the original speech, and voice characteristics of the original speaker. The pipeline starts with automatic speech recognition including emphasis detection, followed by the translation model. The translated text is then synthesized by a Text-to-Speech model that recreates the original emphases in the target sentence. The resulting synthetic speech is then converted back to the original speakers' voice using a voice conversion model. Finally, to synchronize the lips of the speaker with the translated audio, a generative model generates frames of adapted lip movements which are combined with the audio to produce the final output. The disclosure further describes several use-cases and configurations that apply these techniques to video conferencing, dubbing, low-bandwidth transmission, speech enhancement and assistive technology for the hearing impaired.
Owner:WAIBEL ALEXANDER

Comprehensive AI-enabled systems for immersive voice, companion, and augmented / virtual reality interaction solutions

A computer-implemented method for operating an artificial intelligence voice agent system includes receiving voice input through communication channels; analyzing converted text through natural language processing (NLP) pipelines implementing intent recognition and sentiment analysis detecting emotional cues using a multimodal large language model (LLM); generating response content using machine learning models trained on domain-specific corpora; converting generated responses to synthetic speech through text-to-speech (TTS) engines; integrating with a customer relationship management (CRM) platforms or an enterprise resource planning (ERP) database; and implementing continuous learning by updating language understanding models using conversation logs, voice recognition parameters based on user feedback, and response generation patterns. One implementation is a computer-implemented system and method that operates a suite of intelligent interactive devices and platforms including an artificial intelligence voice agent, enhanced communication platforms, an intimacy companion system, and augmented / virtual reality eyeglasses. Further, one implementation includes AR / VR eyeglasses that project visual content onto interchangeable lenses or directly onto the user's retina via laser-based retinal projection, provide prescription adjustments, incorporate ear-mounted sensors for monitoring physiological parameters like heart rate, oxygen saturation, and blood pressure, and utilize wireless data transmission, onboard environmental sensing, and remote calibration, all designed to offer dynamically adaptive, secure, and context-aware interactions across communication, personal assistance, health monitoring, and immersive augmented or virtual reality environments.
Owner:TRAN BAO

Personalized and dynamic text to speech voice cloning using incompletely trained text to speech models

Systems and methods are provided for machine learning models configured as zero-shot personalized text-to-speech models which comprise a feature extractor, a speaker encoder, and a text-to-speech module. The feature extractor is configured to extract acoustic features and prosodic features from new target reference speech associated with the new target speaker. The speaker encoder is configured to generate a speaker embedding corresponding to the new target speaker based on the acoustic features extracted from the new target reference speech. The text-to-speech module is configured to generate the personalized voice corresponding for the new target speaker based on the speaker embedding and the prosodic features extracted from the new target reference speech without applying the text-to-speech module on new labeled training data associated with the new target speaker.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Voice interaction method and device based on lip language enhancement, equipment and storage medium

The invention discloses a voice interaction method and device based on lip language enhancement, equipment and a storage medium, and the method comprises the steps: extracting lip language features based on an image sequence of a lip region, and carrying out the feature extraction of a voice signal, and obtaining an audio feature; performing cross-modal fusion coding on the lip language features and the audio features to generate mixed features containing audio-visual information; inputting the mixed features into a large language model, understanding the intention of the interaction object and generating a corresponding semantic reply; and finally, synthesizing into voice and / or converting into characters. According to the invention, by introducing the lip features, additional visual clues are provided for speech recognition, and the robustness and accuracy of speech recognition can be significantly improved; effective fusion coding is carried out on the lip language features and the sound features, and semantic information splitting caused by simple and independent recognition is avoided; and the capability of the large model is fully utilized, so that more natural and more intelligent interaction experience is realized.
Owner:SHENZHEN WANRUI INTELLIGENT TECH CO LTD

CAD automatic generation system and method based on intelligent model selection and application

The invention discloses a CAD automatic generation system and method based on intelligent model selection and application, and aims at achieving automatic modeling under the multi-modal design requirement. The system comprises a user interaction module for receiving multi-modal input such as natural language, sketch and voice; the intelligent demand analysis module is used for combining an industrial large language model and a product knowledge graph, combining semantic analysis and generating a structured demand; the intelligent model selection calculation module is used for matching the optimal parameter combination and the component list based on a multi-objective optimization algorithm; the CAD automatic generation module calls a parametric modeling engine to generate an editable three-dimensional model; the constraint solving module is used for processing hard constraints and soft constraints in real time and dynamically adjusting model parameters; and the model output and interaction module feeds back a design state and supports user iteration. The system realizes full-process automation from the design intention to the CAD model, improves the design efficiency and accuracy, and is suitable for the fields of mechanical design, intelligent manufacturing and the like.
Owner:HOFMANN (BEIJING) ENG TECH CO LTD

Humanoid robot multi-mode instruction analysis system

The invention discloses a multi-mode instruction analysis system for a humanoid robot. Comprising a voice input module, a visual input module, a voiceprint feature extraction module, an object recognition and pose estimation module, a multi-modal alignment network based on a space-time attention mechanism, a scene semantic tree construction module, an instruction node mapping module, a confidence evaluation module and a decision module. According to the system, accurate alignment of voice and visual information is realized through a space-time attention mechanism, environment information is represented in combination with a scene semantic tree structure, and the instruction analysis accuracy is improved. And dynamically evaluating the confidence coefficient by adopting a fuzzy instruction backtracking algorithm, and if the confidence coefficient is lower than a threshold value, starting multi-round dialogue clarification to reduce misoperation. According to the method, multi-modal data are fused, the historical interaction learning ability is optimized, the understanding efficiency and interaction robustness of complex instructions are remarkably improved, the method is suitable for scenes such as family service and logistics storage, and the intelligent level of man-machine cooperation is enhanced.
Owner:HUIZHOU BEIJIABAO ROBOT CO LTD

English teaching training system and method fusing semantic matching and cognitive evaluation

The invention relates to the technical field of artificial intelligence, and discloses an English teaching training system and method fusing semantic matching and cognitive assessment, and the method comprises the steps: synchronously capturing a text response, a voice intonation, an eye movement track, a facial micro-expression and a touch rhythm generated in a learning process; semantic deviation deconstruction and cognitive intention quantization processing are carried out on the learning interaction original sequence, and a bidirectional deep semantic matching network is adopted to carry out context alignment on student answers and target corpora; based on the word meaning divergence point set and the cognitive load multi-scale vector, extracting a nonlinear diffusion trajectory of a learning state by using a time gating multi-layer recursive trajectory evolution algorithm; the knowledge point nodes, the deviation type nodes and the emotion triggering nodes associated with the emotion instability candidate segments are fused to construct a local learning map; and forming an emotion cognition feedback result driven by learning interest based on the local learning map and the self-adaptive error correction intervention sequence. The method has the advantage of improving the learning interest of students.
Owner:GUILIN INST OF INFORMATION TECH

Payment scene-oriented interaction intention recognition and error correction system

The invention, which relates to the technical field of payment security, discloses a payment-scene-oriented interaction intention identification and error correction system comprising an input analysis module, an intention simulation module, a dynamic decision module, a biological verification module, an audit evidence storage module, and a cross-scene knowledge migration module. According to the method, multi-modal data such as voice, texts, images and touch tracks are integrated, structured feature vectors are generated through a cross-modal attention network, the problem of incomplete single-modal coverage is solved, cross-modal data consistency verification is achieved based on a unified semantic tag system, and the reliability of input sources is graded by combining equipment fingerprints and geographic positions, so that the reliability of the input sources is improved. A high-risk transaction protection capability is enhanced, a generative adversarial network is utilized to construct a virtual attack sample library, attacks such as tampering with characters similar in shape and AI faking voiceprints are simulated, unknown threats are actively defended through cosine similarity matching, a user historical behavior statistical model is integrated, and known risks such as high-frequency small-amount transfer are passively intercepted. And a closed-loop incremental learning continuous optimization model is supported.
Owner:QUANZHOU NORMAL UNIV

Video generation method and interaction method based on digital human, and device, storage medium and program product

Provided in the embodiments of the present application are a video generation method and interaction method based on a digital human, and a device, a storage medium and a program product. In the embodiments of the present application, text-to-speech processing is performed on the basis of voice features of a user and an emotion label, speech-to-expression processing is performed on the basis of a mapping relationship between the voice features of the user and expression coefficients, and a digital human model is rendered on the basis of speech signals and the expression coefficients, so as to obtain video data of the digital human model. Thus, voice features of a user are accurately simulated, so as to ensure that a speech output of a digital human sounds natural and is also highly personalized, thereby realizing personalized driving of the digital human, and improving the realism of the digital human in terms of voice and dynamic images. Thus, the user experience is improved, and the interactivity of the digital human and the authenticity and immersion are enhanced.
Owner:TAOBAO CHINA SOFTWARE

Intelligent glasses AI voice interaction method

The invention provides a smart glasses AI voice interaction method, which comprises the following steps: simultaneously acquiring a user gesture image and a voice signal, acquiring a user historical interaction record, extracting key point coordinates and pointing direction information of the gesture image, acquiring hand distance information, generating a gesture motion path, simultaneously extracting frequency spectrum information and an intonation peak value of the voice signal, and generating a gesture motion path; forming a voice beat sequence; extracting a spatial semantic mode of the gesture and voice fusion data, and recognizing a core object of a user pointing instruction according to pointing coordinates of a gesture motion path and an intonation peak value of a voice beat sequence; after the core object pointing to the instruction is recognized, the user intention is determined in combination with the spatial semantic mode and the historical interaction record of the user; and the complete intention analysis result is output to the intelligent glasses display module to execute corresponding operation, and is fed back to the acquisition module to adjust the next capture parameters including the acquisition frequency and the recognition sensitivity, so that the response speed and the accuracy are improved.
Owner:SHENZHEN YAWELL LNTELLIGENT TECH CO LTD

Multi-modal interaction method and system of digital human intelligent agent

The invention relates to the field of multi-modal interaction analysis, in particular to a multi-modal interaction method and system of a digital human agent. The method comprises the following steps: acquiring a real-time face image and a voice signal input stream of an interactive user based on an intelligent agent; performing real-time micro-expression recognition and deep emotion analysis based on the real-time facial image to obtain real-time emotion features of the user; performing time sequence evolution analysis on the real-time emotion characteristics of the user, performing holographic user emotion deep mining, and constructing a user emotion holographic characteristic spectrum; carrying out adaptive acoustic gain processing on the voice signal input stream, and carrying out voice-emotion association analysis based on the user emotion holographic characteristic spectrum to generate a voice-emotion linkage mapping spectrum; and carrying out eyeball fixation point migration tracking based on the user emotion holographic feature map and the real-time face image, and generating a user interaction depth intention signal. Through the real-time deep semantic understanding and emotion perception ability, the intelligent agent interaction intelligence and response accuracy are improved.
Owner:GUANGDONG HUITONG INFORMATION TECH CO LTD

Dynamic calibration method and system of vehicle-mounted emotion recognition system

The invention provides a dynamic calibration method and system for a vehicle-mounted emotion recognition system, and the method comprises the steps: S1, obtaining multi-source data which comprises a facial image, a voice signal and a physiological signal; the obtained multi-source data are preprocessed, and preprocessed multi-source data are obtained; s2, performing feature extraction based on the preprocessed facial image, the voice signal and the physiological signal to obtain a facial expression feature vector, an audio feature vector and a physiological state feature vector; s3, evaluating the current environment credibility based on an environment credibility evaluation function; s4, dynamically distributing the weight of the multi-source data according to the credibility of the current environment and the real-time scene; and S5, constructing a multi-modal fusion vector based on the dynamically distributed weight of the multi-source data, the facial expression feature vector, the audio feature vector and the physiological state feature vector, and performing emotion recognition by using the constructed emotion recognition model based on the multi-modal fusion vector.
Owner:SHANGHAI PUFAFEN ELECTRONIC TECH CO LTD