Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

5079 results about "Speech model" patented technology

Serial Processing Models. Serial models of speech production present the process as a series of sequential stages or modules, with earlier stages comprising of the large units (i.e. sentences and phrases), and later stage comprising of their smaller unit constituents (i.e. distinct features like voicing, phonemes, morphemes, syllables).

Customer data processing and insight system based on large language model

The invention belongs to the technical field of artificial intelligence and big data, and discloses a customer data processing and insight system based on a big language model. The system is composed of a multi-source data access module, a data preprocessing and label fusion module, a large language model semantic understanding module, a knowledge enhancement and semantic linkage module, an insight generation and visualization module, an intelligent strategy output module and a feedback learning and self-optimization module. According to the method, multi-source heterogeneous data such as texts, voices and structured behaviors are integrated, and the deep semantic analysis capability of a large language model is combined, so that global modeling of customer behaviors and intentions is realized; a multi-modal synchronous acquisition and standardization mechanism eliminates data format barriers, and a dynamic label mechanism adapts to context changes, so that the system can capture deep semantic association in customer expression, and compared with a traditional keyword matching method, the semantic understanding accuracy is improved by more than 40%, and a more complete data base is provided for insight generation.
Owner:SICHUAN JUFUREN TECHNOLOGY CO LTD

Vision generation method and device based on semantic association modeling, equipment and medium

The invention relates to the technical field of voice semantics, can be applied to business scenes of financial science and technology, medical health, poster design and the like, and discloses a visual sense generation method and device based on semantic association modeling, equipment and a medium. Generating a demand text containing theme and style parameters; semantic features in the demand text are extracted, semantic association weights are constructed, and element layout coordinates are optimized in combination with spatial distribution constraints; and encoding the layout information into a control matrix, fusing the control matrix with the initial noise, adjusting a noise reduction process through an encoding and decoding network, and generating target visual content highly matched with the semantic meaning of the user instruction. According to the method, the layout optimization function is constructed, the diffusion model is guided to focus the semantic salient region in space, language model output and the visual generation process are closely combined, structured response and space mapping of user semantic requirements are achieved, and the expression consistency and personalized adaptation capacity of visual content generation are improved.
Owner:SHENZHEN PINGAN COMM TECH CO LTD

Intelligent operation and maintenance method for power grid equipment based on large language model and knowledge graph

The invention discloses a power grid equipment intelligent operation and maintenance method based on a large language model and a knowledge graph, and relates to the technical field of intelligent power grid operation and maintenance, and the method comprises the steps: carrying out the operation and maintenance of power grid equipment through a power grid operation and maintenance knowledge graph constructed through a large language model and a knowledge federation technology, the power grid equipment operation and maintenance comprises one or more of health state evaluation, fault risk prediction and early warning, intelligent operation and maintenance strategy generation and health degree dialogue query of the power grid equipment. According to the invention, a power grid operation and maintenance mode can be effectively promoted to be transformed and upgraded from a traditional manual experience type and a passive maintenance type to a data-driven, intelligent and active preventive maintenance mode. Key intelligent operation and maintenance technical support is provided for building a novel electric power system with new energy as a main body, the novel electric power system is assisted to achieve the development goals of being safer, more efficient, cleaner and lower in carbon, and important industry strategic significance and social contribution are achieved.
Owner:GANSU ZHENGPENG ELECTRIC POWER TECHNOLOGY CO LTD

Large language model (LLM) for enterprise applications developed by codeless platform

The present invention provides a large language model-based system and method for data processing in application developed by codeless platform. The invention includes identification of intent of a user to process procurement, supply chain, application integration, application restructuring or development scenarios.
Owner:NB VENTURES INC DBA GEP

Electric power work order intelligent processing method with RPA fused with multi-mode large model

The invention relates to the technical field of intelligent operation and maintenance and artificial intelligence crossing of a power system, in particular to an intelligent power work order processing method of an RPA fused multi-modal large model, which analyzes multi-modal work order data such as texts, voices, images and the like through a domain adaptation large language model, and realizes fault key information extraction and conflict resolution in combination with a dynamic knowledge graph; performing work order priority scoring and resource allocation by using space-time constraint reinforcement learning; an analysis result is converted into an automatic execution script through an RPA engine, and a whole-process closed loop of order sending, processing and feedback is achieved; meanwhile, a feedback optimization and conflict resolution cooperation mechanism is constructed, and the knowledge graph and the model precision are continuously iterated. The method improves work order processing efficiency and analysis precision, enhances decision scientificity, and is suitable for an intelligent operation and maintenance scene of a power system.
Owner:FUJIAN ZEYUAN INFORMATION TECHNOLOGY CO LTD

Multi-modal interview automatic quality analysis and evaluation method and system based on large model

The invention discloses a multi-modal interview automatic quality analysis and evaluation method and system based on a large model, and the method comprises the steps: collecting and storing multi-modal data, such as texts, audios, videos and behavior interaction, and carrying out the preprocessing of the multi-modal data to form a standardized data set; utilizing a preset interview structure and a large model to dynamically guide the process, adjusting the topic rhythm according to real-time feedback, and recording stage conversion information to form logic trajectory data for process coherence management; automatically coding text data through a large language model, extracting features such as keywords and performing topic clustering, performing cross validation and semantic fusion in combination with data analysis results of each modal, and generating deep analysis results such as psychological states; and generating a comprehensive assessment report containing qualitative description, quantitative score and psychological abnormality or cognitive disorder risk prompts based on a deep analysis result, thereby providing a basis for psychological health assessment and cognitive competence evaluation. According to the method, automatic analysis of multi-modal data is realized, and evaluation scientificity and efficiency are improved.
Owner:BEIJING NORMAL UNIVERSITY +1

Real-time virtual reality scene system based on natural language description using multimodal artificial intelligence

A real-time system for the multimodal generation of virtual reality scenes based on artificial intelligence for the creation of immersive three-dimensional environments from natural language narratives, consisting of: a speech capture module configured to continuously record a user's spoken narrative via one or more directional microphones, preprocesses the captured signal by noise reduction and temporal alignment, and outputs a digital speech stream; A speech-to-text processing unit that is operationally coupled to the speech capture module and configured for real-time speech recognition using a continuous neural transformer model. The unit is trained to transcribe natural language utterances into structured text data while maintaining contextual continuity throughout the evolving narrative. a semantic interpretation processing unit that is communicatively linked to the speech recognition unit and configured to perform natural language understanding techniques to extract contextual entities, spatial references, temporal relationships, and object attributes from the transcribed narrative; the engine includes a large language model that is fine-tuned for spatial reasoning tasks; a scene graph generation module configured to transform the interpreted semantic data into a structured, hierarchical representation that defines nodes for identified entities and edges for corresponding relationships, with each node associated with metadata describing geometry, position, orientation, texture, and linking attributes between objects; a multimodal image-language model processor coupled with the scene graph generation module, wherein the processor is configured to retrieve, adapt, or synthesize appropriate three-dimensional elements from a pre-trained visual-lexical embedding space and align these elements with their semantic and spatial definitions derived from the scene graph; a scene assembly and rendering controller configured to create a cohesive virtual scene from the aligned assets, perform real-time rendering using a GPU-accelerated ray tracing pipeline, and produce a stereoscopic visual output that corresponds to the evolving narrative; A head-mounted virtual reality visualization device connected to the rendering engine and configured to display the generated immersive environment to the user in real time. The device features motion sensors and inside-out tracking cameras to detect head and body movements, dynamically updating viewing angles and perspective within the rendered scene; and a bidirectional feedback module integrated into the head-mounted device and connected to the semantic interpretation processing unit; the module is configured to interpret corrective commands, gestures, or supplementary comments from the user to refine or modify specific scene elements without interrupting the real-time visualization; The system continuously updates the virtual scene as the narrative develops, ensuring temporal synchronization between speech input and rendered output below a defined latency threshold, thus enabling a natural, dialogic construction of complex three-dimensional virtual environments.
Owner:GOUNDER MOHAN SELLAPPA DR BENGALURU +3

Industrial agent decision-making method and device, electronic equipment and storage medium

The invention provides an industrial agent decision-making method and device, electronic equipment and a storage medium. The method comprises the following steps: loading multi-source industrial data to a knowledge graph, and constructing a domain ontology and an MCP ontology; performing parameter fine tuning on a preset base large language model; receiving an industrial process planning demand input by a user, and performing semantic analysis on the industrial process planning demand; processing the query result by utilizing an MCP context strategy to generate a prompt context; inputting the prompt context and the industrial process planning demand into a fine tuning large language model to generate a first process planning scheme; performing rule verification on the first process planning scheme based on the knowledge graph, and judging whether the first process planning scheme meets a preset rule or not according to a verification result; and converting the final process planning scheme into a control script or an interface calling sequence which can be analyzed by an industrial execution system, and outputting the control script or the interface calling sequence. According to the method, the cross-software system compliance process planning with unified knowledge, automatic generation and rule verification can be realized.
Owner:HANGZHOU HOLLYSYS AUTOMATION

Text prediction-based large-model real-time voice text intention recognition method and system

The invention discloses a large-model real-time voice text intention recognition method and system based on text prediction, and the method comprises the steps: obtaining the real-time voice data of a user, carrying out the real-time voice recognition processing through a streaming voice recognition interface, and obtaining a part of transcriptional text; inputting the partial transcription text into a mask language model for text prediction, and generating a plurality of high-credibility complete sentence candidates; based on the complete sentence candidates, the complete sentence candidates are input into a large language model in parallel for intention recognition, a corresponding intention result is obtained, and a mapping relation between the candidate sentences and the intention recognition result is established; and obtaining a sentence completely expressed by the user, calculating the similarity between the complete actual sentence and a plurality of high-credibility complete sentence candidates through a multi-level text similarity algorithm, selecting the candidate sentence with the highest similarity score, and directly obtaining a corresponding final intention recognition result based on the mapping relationship. The objective of the invention is to solve the technical problem of high response delay of an existing voice intention recognition system.
Owner:BEIJING YULORE INNOVATION TECH

Robot anthropomorphic interaction method based on multi-modal emotion recognition and customized portrait generation

The invention discloses a robot anthropomorphic interaction method based on multi-modal emotion recognition and customized portrait generation. The method comprises the following steps: S1, dynamically fusing multi-modal emotions; the method comprises the following steps: S1, synchronously acquiring voice, visual and text signals through a multi-source heterogeneous sensor, capturing a user voice stream by a high-fidelity microphone array, and extracting acoustic characteristics such as intonation and speed, S2, performing cross-modal reasoning; s3, synchronously generating contents; step S4: style migration; step S5, anthropomorphic voice and expression generation; according to the method, man-machine interaction emotion is analyzed and generated by utilizing a large language model and multi-modal information fusion, the singleness of interaction emotion and the deficiency of emotional sharing ability are avoided, a strong emotion interaction characteristic is achieved, the image of the robot is obtained through a generative technology and can be migrated to any image, the limitation that a specific image is independently made is broken through, and the interaction effect of the robot is improved. The advantage that one robot can be suitable for different scenes is achieved.
Owner:JIANGSU YUNMU ZHIZAO TECH CO LTD

Automated error troubleshooting via generative AI software development assistant

Techniques for leveraging a large language model (LLM) in software development are described. A description of an error message associated with a software system hosted by the multi-tenant provider network is received. An LLM is prompted with an analysis prompt to generate an analysis of the error message, the analysis prompt including the description of the error message, and the analysis of the error message is received from the LLM. The LLM is prompted with a suggested resolution prompt to generate a suggested resolution to a cause of the error message, the suggested resolution prompt including the description of the error message and the analysis of the error message, and the suggested resolution is received from the LLM. A change to resolve the cause of the error message is sent to an originator of the received description, the change based at least in part on the suggested resolution.
Owner:AMAZON TECH INC

AI generation content detection and review method and device, equipment and storage medium

The invention discloses an AI generation content detection and review method, device and equipment and a storage medium, and the method comprises the steps: receiving multi-modal input data which comprises text data, image data and video data; calling a large language model to carry out compliance analysis on the text data to obtain an analysis result; when the analysis result is compliance, calling a multi-modal model to identify and detect the image data and the video data, and respectively obtaining an image detection result and a video detection result; and aggregating the analysis result, the image detection result and the video detection result to obtain a content detection result. According to the method, the processing paths are automatically allocated according to the content types, redundant calculation is avoided, hardware consumption is remarkably reduced, the multi-modal detection results are aggregated into structured output, the large-batch content processing efficiency is remarkably improved, calculation resource occupation is greatly reduced, and seamless integration of the detection results and a downstream service system is achieved.
Owner:深圳市维卓数字营销有限公司

Robot control method, system and equipment based on multi-modal large model and medium

The invention relates to the technical field of robot control, and discloses a robot control method, system, equipment and medium based on a multi-modal large model, and the method comprises the steps: collecting the multi-source modal data of a scene where an operation task is located, and carrying out the processing through a machine learning model, obtaining a multi-modal feature, and carrying out the position coding and Transform fusion processing, multi-modal fusion features are obtained, the multi-modal fusion features and the constructed job task knowledge base are input into a large language model to decompose a target job task, a human-in-the-loop mechanism is introduced to optimize a decomposition result, and a sub-task sequence is obtained; according to a subtask type in the subtask sequence, processing the subtask sequence through a visual language action model or a reinforcement learning model, and generating a motion instruction to enable the robot to start an execution process of the target operation task; live-line work tasks are processed through the multi-modal large models LLM, VLA and the like, and the work efficiency of the autonomous distribution network live-line work robot is improved.
Owner:WENZHOU ELECTRIC POWER BUREAU +2

Face-translator: end-to-end system for speech-translated lip-synchronized and voice preserving video generation

A neural end-to-end system is provided for the face and voice preserving translation of videos. The system is a pipeline of multiple models that produces a video of the original speaker speaking in the target language with modified lip movement to match the target speech, while preserving emphases and prosody of the original speech, and voice characteristics of the original speaker. The pipeline starts with automatic speech recognition including emphasis detection, followed by the translation model. The translated text is then synthesized by a Text-to-Speech model that recreates the original emphases in the target sentence. The resulting synthetic speech is then converted back to the original speakers' voice using a voice conversion model. Finally, to synchronize the lips of the speaker with the translated audio, a generative model generates frames of adapted lip movements which are combined with the audio to produce the final output. The disclosure further describes several use-cases and configurations that apply these techniques to video conferencing, dubbing, low-bandwidth transmission, speech enhancement and assistive technology for the hearing impaired.
Owner:WAIBEL ALEXANDER

Comprehensive AI-enabled systems for immersive voice, companion, and augmented / virtual reality interaction solutions

A computer-implemented method for operating an artificial intelligence voice agent system includes receiving voice input through communication channels; analyzing converted text through natural language processing (NLP) pipelines implementing intent recognition and sentiment analysis detecting emotional cues using a multimodal large language model (LLM); generating response content using machine learning models trained on domain-specific corpora; converting generated responses to synthetic speech through text-to-speech (TTS) engines; integrating with a customer relationship management (CRM) platforms or an enterprise resource planning (ERP) database; and implementing continuous learning by updating language understanding models using conversation logs, voice recognition parameters based on user feedback, and response generation patterns. One implementation is a computer-implemented system and method that operates a suite of intelligent interactive devices and platforms including an artificial intelligence voice agent, enhanced communication platforms, an intimacy companion system, and augmented / virtual reality eyeglasses. Further, one implementation includes AR / VR eyeglasses that project visual content onto interchangeable lenses or directly onto the user's retina via laser-based retinal projection, provide prescription adjustments, incorporate ear-mounted sensors for monitoring physiological parameters like heart rate, oxygen saturation, and blood pressure, and utilize wireless data transmission, onboard environmental sensing, and remote calibration, all designed to offer dynamically adaptive, secure, and context-aware interactions across communication, personal assistance, health monitoring, and immersive augmented or virtual reality environments.
Owner:TRAN BAO

Personalized and dynamic text to speech voice cloning using incompletely trained text to speech models

Systems and methods are provided for machine learning models configured as zero-shot personalized text-to-speech models which comprise a feature extractor, a speaker encoder, and a text-to-speech module. The feature extractor is configured to extract acoustic features and prosodic features from new target reference speech associated with the new target speaker. The speaker encoder is configured to generate a speaker embedding corresponding to the new target speaker based on the acoustic features extracted from the new target reference speech. The text-to-speech module is configured to generate the personalized voice corresponding for the new target speaker based on the speaker embedding and the prosodic features extracted from the new target reference speech without applying the text-to-speech module on new labeled training data associated with the new target speaker.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Voice interaction method and device based on lip language enhancement, equipment and storage medium

The invention discloses a voice interaction method and device based on lip language enhancement, equipment and a storage medium, and the method comprises the steps: extracting lip language features based on an image sequence of a lip region, and carrying out the feature extraction of a voice signal, and obtaining an audio feature; performing cross-modal fusion coding on the lip language features and the audio features to generate mixed features containing audio-visual information; inputting the mixed features into a large language model, understanding the intention of the interaction object and generating a corresponding semantic reply; and finally, synthesizing into voice and / or converting into characters. According to the invention, by introducing the lip features, additional visual clues are provided for speech recognition, and the robustness and accuracy of speech recognition can be significantly improved; effective fusion coding is carried out on the lip language features and the sound features, and semantic information splitting caused by simple and independent recognition is avoided; and the capability of the large model is fully utilized, so that more natural and more intelligent interaction experience is realized.
Owner:SHENZHEN WANRUI INTELLIGENT TECH CO LTD

CAD automatic generation system and method based on intelligent model selection and application

The invention discloses a CAD automatic generation system and method based on intelligent model selection and application, and aims at achieving automatic modeling under the multi-modal design requirement. The system comprises a user interaction module for receiving multi-modal input such as natural language, sketch and voice; the intelligent demand analysis module is used for combining an industrial large language model and a product knowledge graph, combining semantic analysis and generating a structured demand; the intelligent model selection calculation module is used for matching the optimal parameter combination and the component list based on a multi-objective optimization algorithm; the CAD automatic generation module calls a parametric modeling engine to generate an editable three-dimensional model; the constraint solving module is used for processing hard constraints and soft constraints in real time and dynamically adjusting model parameters; and the model output and interaction module feeds back a design state and supports user iteration. The system realizes full-process automation from the design intention to the CAD model, improves the design efficiency and accuracy, and is suitable for the fields of mechanical design, intelligent manufacturing and the like.
Owner:HOFMANN (BEIJING) ENG TECH CO LTD

Robot multi-modal fusion autonomous decision-making method and system based on large language model

The invention relates to the technical field of robot decision making, and provides a robot multi-modal fusion autonomous decision making method and system based on a large language model.The method comprises the steps that a robot obtains multi-modal environment information through a visual sensor, a touch sensor, an auditory sensor and a laser radar which are carried by the robot; performing preliminary filtering and noise reduction processing on the original sensor data, and synchronously recording all the sensor data by timestamps; performing space-time semantic alignment on the preprocessed multi-modal data, mapping pixel coordinates of a target in a visual target coordinate quantization original image to a robot coordinate system, performing uncertainty evaluation on a multi-modal signal through a dynamic Bayesian network, and taking entropy or variance as an uncertainty quantitative evaluation index. According to the method, the information quality is improved from a data fusion source, accurate and reliable basic support is provided for subsequent decision making, and decision making errors caused by data deviation are greatly reduced.
Owner:ANHUI UNIV +1

Question answering system construction method and system based on large language model

The invention provides a question and answer system construction method and system based on a large language model, and the method comprises the steps: obtaining multi-modal data, constructing a question and answer knowledge base and a knowledge graph, and carrying out the dynamic updating of the question and answer knowledge base; obtaining a query text, and respectively carrying out vectorization processing on the query text and the multi-modal data to generate a corresponding query semantic vector and a multi-modal vector; an entity in the query text is extracted by using the recognition model, a triple associated with the entity is extracted from the knowledge graph, the query text and the triple are spliced and vectorized, and a query semantic enhancement vector is generated; according to the method, through a dynamic knowledge base incremental updating mechanism, a context-aware hybrid retrieval strategy, a cross-modal semantic enhancement technology and a user feedback-driven continuous optimization method, real-time processing requirements of various modal data such as texts, images and voices can be effectively met, and accurate semantic understanding and answer generation of complex queries are achieved.
Owner:HUBEI ZHONGKE NETWORK ENG

Cross-modal knowledge optimization system for improving localization adaptability of large language model

PendingCN120542521ASemantic analysis2D-image generationDeep knowledgeEngineering
The invention relates to the field of natural language processing. The invention discloses a cross-modal knowledge optimization system for improving localization adaptability of a large language model, and the system comprises a modal data collection module which collects cross-modal localization data of characters, audios and videos; the data preprocessing module is connected with the data acquisition module and preprocesses acquired data; the cross-modal knowledge fusion module is connected with the preprocessing module, fuses data and large language model general knowledge, and performs mining association by means of a cross-modal learning algorithm to form a localized cross-modal knowledge graph; the model fine-tuning module is connected with the fusion module and uses the map to finely tune the large language model; the evaluation feedback module is connected with the fine tuning module and evaluates the localization performance of the model. Through multi-modal data integration, dynamic preprocessing, deep knowledge fusion, efficient fine adjustment and intelligent feedback, various problems in the large language model localization process are systematically solved, and the expression of the model in dialect understanding, cultural questions and answers and cross-modal tasks is remarkably improved.
Owner:HANGZHOU LANGSHI VIDEO TECH CO LTD

Government information consultation system based on large language model

PendingCN120653787ASemantic analysisKnowledge representationConsultation systemEngineering
The invention belongs to the technical field of artificial intelligence, and discloses a government information consultation system based on a large language model, which comprises a user interaction module, a large language model core engine, a knowledge base integration module, a multi-level authority management module, a feedback optimization mechanism and a risk control module. The comprehensive intelligent government affair service system is constructed through the six core modules, remarkable advantages are shown in government affair service digital transformation, the system innovatively adopts multi-mode interactive design, multiple input and output modes of voice, text and images are supported, an intelligent authority management mechanism is matched, and the intelligent authority management mechanism is matched with the intelligent authority management mechanism. According to the technical scheme, the accessibility and convenience of government affair services are greatly improved, precise services for different user groups are achieved, it is guaranteed that sensitive data are safe and controllable while wide spreading of government affair information is guaranteed, and a large language model of a system core is subjected to professional government affair scene optimization training and is combined with a dynamically-updated knowledge graph technology.
Owner:JIANGXI YUANREN ENTERPRISE MANAGEMENT CO LTD

Prompt management for large language model

Systems and methods for a prompt generation and analysis service for generating and identifying a preferred prompt for performing a function of a large language model (LLM) are provided. The prompt generation and analysis service may generate a set of training prompts for performing a function of an LLM. The prompt generation and analysis service may then query the LLM with the generated set of prompts and characterize the output of the LLM for each prompt. Using the characterization of the output and corresponding prompt, the prompt generation and analysis service can train a classifier model to classify the prompts. The prompt generation and analysis service may generate a set of target prompts for performing a function of an LLM, characterize the target prompts using the training classifier model, and identify a preferred prompt for performing the function based on the classifier model's classification.
Owner:AMAZON TECH INC

Multi-unmanned aerial vehicle task scheduling method and system with dependence perception and feedback mechanism

The invention discloses a multi-UAV (unmanned aerial vehicle) task scheduling method and system with a dependency perception and feedback mechanism, and the method comprises the steps: enabling a commander to input a task demand in a voice or text form, inputting the task into a large language model based on a Python prompt template in combination with environment information and UAV capability configuration, and enabling the large language model to perform task scheduling; and completing subtask disassembly and dependency modeling of the natural language instruction. The method comprises the following steps: establishing a sub-task dependency graph, and determining a sequential relationship and execution logic between tasks; in the aspect of task scheduling, capability vector modeling is carried out on all online unmanned aerial vehicles, and an optimal unmanned aerial vehicle is selected or a multi-vehicle alliance is automatically constructed to execute a task based on a vector matching degree between task skill requirements and unmanned aerial vehicle capabilities. In the task execution process, task state information is collected in real time, and all feedback information is uploaded to the cloud control center for state judgment and abnormity recognition. When the system detects an abnormal condition, task reconstruction, alliance recombination and scheduling graph repair are automatically carried out, and closed-loop adjustment of the task is completed.
Owner:HOHAI UNIV +1

Cross task large language model fine-tuning

Aspects of the disclosure include an architecture for cross task large language model fine-tuning based on a shared context and methods of using the same. An exemplary method includes receiving a pre-trained large language model and receiving a set of fine-tuning tasks for the pre-trained large language model. The set of fine-tuning tasks includes at least a first fine-tuning task and a second fine-tuning task. The method includes generating, from the set of fine-tuning tasks, a first task combination including a subset of the set of fine-tuning tasks, identifying a shared subspace within the subset of the set of fine-tuning tasks, and responsive to identifying the shared subspace, fine-tuning the pre-trained large language model jointly over the first task combination.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Long video multi-modal understanding and question-answering method and system based on large model and retrieval enhancement generation

The invention discloses a long video multi-modal understanding and question-answering method and system based on large model and retrieval enhancement generation. The method comprises the following steps: 1) a multi-modal feature extraction module; 2) a multi-modal synchronization and alignment mechanism; 3) constructing a structured memory pool; 4) querying a drive generation mechanism; 5) incremental updating and memory compression strategy; and 6) unifying the multi-modal representation space. The invention provides a long video multi-mode understanding method fusing a large language model and retrieval enhancement generation, and aims to break through the limitation of a traditional method in the aspects of single-mode processing and semantic fragmentation. According to the method, video image features are extracted through a visual model (such as YOLO and ViT), voice transcription and environment voice description are obtained in combination with an audio model (such as Whisper and Qwen-Audio), and unified coding of vision, voice and audio in a long video is achieved. Then, a structured memory pool is constructed through semantic consistency segmentation and timestamp alignment technologies to store time slice data of different modalities.
Owner:GUANGZHOU BINGO SOFTWARE +1

Mechanical arm natural language instruction control system and method based on large language model

The invention discloses a mechanical arm natural language instruction control system and method based on a large language model, and belongs to the field of intelligent manufacturing. Aiming at the limitation that traditional mechanical arm control depends on pre-programming and a static rule library, a dynamic mapping mode from a natural language instruction to an atomic action sequence is designed, an atomic skill library including detection, grabbing, moving, placement and other operations is constructed, and semantic analysis and a multi-mode cooperation technology are combined, so that the atomic action sequence is obtained. And support is provided for man-machine cooperation of a flexible assembly task. The method specifically comprises the steps that a DeepSeek-Distil-Llam-8B large model and a LoRA fine tuning technology are adopted, and a natural language instruction is converted into an executable atomic action sequence; based on a transfer learning optimized YOLOv8 target detection technology and a binocular vision positioning technology, a sensing module adaptive to an assembly scene is constructed and is fused with a mechanical arm motion planning module, and positioning grabbing of parts and tools is achieved. And an interactive interface is built by combining a voice-to-text large model and a Gradio front-end framework, so that the convenience of man-machine interaction is improved. By optimizing large model reasoning and motion planning cooperation efficiency, response delay from instructions to execution is reduced, and an efficient and extensible solution is provided for man-machine cooperation in intelligent manufacturing.
Owner:BEIJING INST OF TECH

Generating a response for a communication session based on previous conversation content using a large language model

An example operation may include one or more of receiving interaction content from a communication session between a source device and a service provider device of a service provider, identifying a search criteria from the interaction content, retrieving a subset of vectors from a plurality of vectors stored in a vector database based on the search criteria of the interaction content, wherein the subset of vectors includes previous interaction content with the service provider, generating a response for the communication session based on execution of a large language model (LLM) on the subset of vectors, and outputting the response to at least one of the source device and the service provider device during the communication session.
Owner:THE TORONTO DOMINION BANK

Requirements discovery for generative ai software development assistant

Techniques for leveraging a large language model (LLM) in software development are described. A description of a software development task is received from a user. Data associated with the software system is obtained from a data source. An LLM is prompted to identify at least one aspect of the task which requires clarification from the user, at least partly by providing the obtained data to the LLM and asking the LLM to identify a question for the user which remains unanswered by the obtained data. The question is presented to the user. An answer to the question is received from the user. The LLM is prompted to respond to propose an implementation of the task at least partly based on the data associated with the software system and the answer received from the user. The proposed implementation is received from the LLM and caused to be displayed to the user.
Owner:AMAZON TECH INC

Real-time anti-fraud monitoring system and method based on behavior reasoning and sentiment analysis

The invention relates to the technical field of artificial intelligence, in particular to a real-time anti-fraud monitoring system and method based on behavior reasoning and sentiment analysis, and the system comprises a multi-modal data collection unit, an edge preprocessing unit, a feature fusion and behavior reasoning unit, a large language model context reasoning unit, a risk assessment and decision unit, and an intervention execution unit. A log recording and federal incremental learning unit; the method has the beneficial effects that the traditional isolated single-mode detection is evolved into an emotion and behavior dual-channel collaborative multi-mode recognition system through millisecond-level coaxial alignment of voice, video and user operation logs; the robustness of dialect, noise and expression shielding is greatly improved through the multi-modal fusion model, so that the cross-scene recognition accuracy is improved by nearly three percent compared with that of a traditional single-voice scheme, and high-sensitivity capture of hidden and emotion control type fraud is truly achieved.
Owner:INSPUR TIANYUAN COMM INFORMATION SYST CO LTD