Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

9 results about "Visual vocabularies" patented technology

Visual vocabulary consists of images or pictures that stand for words and their meanings. In the same way that individual words make written language possible, individual images make a visual language possible. The term also applies to a theory of visual communication that says pictures and images can be “read” in the same way that words can.

Real-time virtual reality scene system based on natural language description using multimodal artificial intelligence

A real-time system for the multimodal generation of virtual reality scenes based on artificial intelligence for the creation of immersive three-dimensional environments from natural language narratives, consisting of: a speech capture module configured to continuously record a user's spoken narrative via one or more directional microphones, preprocesses the captured signal by noise reduction and temporal alignment, and outputs a digital speech stream; A speech-to-text processing unit that is operationally coupled to the speech capture module and configured for real-time speech recognition using a continuous neural transformer model. The unit is trained to transcribe natural language utterances into structured text data while maintaining contextual continuity throughout the evolving narrative. a semantic interpretation processing unit that is communicatively linked to the speech recognition unit and configured to perform natural language understanding techniques to extract contextual entities, spatial references, temporal relationships, and object attributes from the transcribed narrative; the engine includes a large language model that is fine-tuned for spatial reasoning tasks; a scene graph generation module configured to transform the interpreted semantic data into a structured, hierarchical representation that defines nodes for identified entities and edges for corresponding relationships, with each node associated with metadata describing geometry, position, orientation, texture, and linking attributes between objects; a multimodal image-language model processor coupled with the scene graph generation module, wherein the processor is configured to retrieve, adapt, or synthesize appropriate three-dimensional elements from a pre-trained visual-lexical embedding space and align these elements with their semantic and spatial definitions derived from the scene graph; a scene assembly and rendering controller configured to create a cohesive virtual scene from the aligned assets, perform real-time rendering using a GPU-accelerated ray tracing pipeline, and produce a stereoscopic visual output that corresponds to the evolving narrative; A head-mounted virtual reality visualization device connected to the rendering engine and configured to display the generated immersive environment to the user in real time. The device features motion sensors and inside-out tracking cameras to detect head and body movements, dynamically updating viewing angles and perspective within the rendered scene; and a bidirectional feedback module integrated into the head-mounted device and connected to the semantic interpretation processing unit; the module is configured to interpret corrective commands, gestures, or supplementary comments from the user to refine or modify specific scene elements without interrupting the real-time visualization; The system continuously updates the virtual scene as the narrative develops, ensuring temporal synchronization between speech input and rendered output below a defined latency threshold, thus enabling a natural, dialogic construction of complex three-dimensional virtual environments.
Owner:GOUNDER MOHAN SELLAPPA DR BENGALURU +3

Visual representation global modeling method and system based on feature extraction

The invention discloses a visual representation global modeling method and system based on feature extraction, and the method comprises the steps: processing input visual data through employing a visual vocabulary bag feature coding algorithm, generating a feature vector, and obtaining space-time continuous feature representation through the dimension reduction reconstruction of a space-time continuous visual autoencoder; pixel and semantic category mapping is established by using a pixel-level semantic feature mapping algorithm, a pixel-level semantic feature map is generated, and the step comprises the sub-steps of dimension analysis, model construction and the like; inputting the feature map into a high-dimensional visual representation intelligent analysis platform, and screening through sub-steps of segmentation, standardization, correlation analysis and the like to obtain high-dimensional screening features; and on the basis of the screening features, constructing and optimizing a global feature incidence matrix through a visual representation global modeling technology, and outputting a global model. The system comprises corresponding function units, the visual data global association rule is accurately captured, and adaptability and analysis accuracy in a complex scene are enhanced.
Owner:ZHENJIANG ZHIGU HIGH END EQUIP RES INST CO LTD

Vision vocabulary guided multimodal large model hallucination optimization method and storage medium

The present application relates to the technical field of large model training, and discloses a visual vocabulary guided multimodal large model hallucination optimization method and a storage medium, training samples including images, text prompts, preferred answers and rejected answers are constructed, a visual target detection model is used to perform image detection on the images according to the preferred answers, relevant indexes consistent with visual words are extracted, and the visual weights of each visual word are calculated accordingly, different types of words are differentially weighted according to the visual weights, a dynamic adjustment coefficient is defined according to the proportion of the corrected segment, and the dynamic adjustment coefficient is combined with the probability difference and the differential weighting to construct a DPO loss function, so as to guide the multimodal large model to generate answers consistent with images and texts. The differential optimization weight is dynamically assigned, so that the multimodal large model pays attention to the words highly related to the image semantics, thereby generating information description more consistent with visual facts, and improving the consistency and authenticity of text generation.
Owner:COMMUNICATION UNIVERSITY OF CHINA

Visual vocabulary guided multi-modal large model illusion optimization method and storage medium

The invention relates to the technical field of large model training, and discloses a visual vocabulary-guided multi-modal large model illusion optimization method and a storage medium, and the method comprises the steps: constructing a training sample comprising an image, a text prompt, an optimal answer and a rejection answer, carrying out the image detection of the image through a visual target detection model according to the optimal answer, and obtaining a visual target detection result; according to the method, related indexes consistent with visual words are extracted, the visual weight of each visual word is calculated according to the indexes, different types of words are subjected to differential weighting according to the visual weights, a dynamic adjustment coefficient is defined according to the proportion of a corrected fragment, and the dynamic adjustment coefficient, probability difference and differential weighting jointly construct a DPO loss function. And guiding the multi-mode large model to generate an answer with consistent image-text. According to the method, the multi-modal large model is dynamically endowed with differentiated optimization weights, so that the multi-modal large model pays attention to vocabularies highly related to image semantics, information description more conforming to visual facts is generated, and the consistency and authenticity of text generation are improved.
Owner:COMMUNICATION UNIVERSITY OF CHINA

Method and system for generating a goods picking and placing strategy

The application provides a kind of goods grabbing and placing strategy generation method and system, method includes the sampling processing to original video data, obtains continuous video frame sequence;Video frame sequence is encoded, and the discrete visual vocabulary sequence is generated;The time sequence dependence modeling of visual vocabulary sequence is carried out to learn the continuous change rule of object action in the video content represented by visual vocabulary sequence;Based on the visual vocabulary corresponding to the current state of the goods to be operated and the visual vocabulary corresponding to the expected target state, and using the continuous change rule of object action, the latent action vector representing the action execution sequence required for the transition from the current state to the target state is generated;The real-time physical state information of the robot obtained is mapped with the latent action vector, to generate the control instruction for driving the robotic arm to perform goods grabbing or placing operation, effectively reduces the dependence on a large amount of manually labeled data, and improves the self-adaptation and generalization ability of the model to unseen scenarios.
Owner:SENAD TECH CO LTD

Image data augmentation method and device based on generative adversarial network

The image data augmentation method and device based on the generative adversarial network of the present disclosure can extract the underlying semantic features of training images to form visual vocabulary of the training images; can label target annotation words of a target scene according to the connotation of the target scene, and the connotation and extension of the visual vocabulary; can construct a generative adversarial network, and train the generative adversarial network by using the training images, the target annotation words and random noise until the image data output by the generative adversarial network meets the requirements, and obtain augmented image data. The method can break through the conversion bottleneck of sample semantic scenes, solve the small sample image classification and recognition problem of multi-source information, improve the quality of augmented image data, and improve the performance of artificial intelligence algorithms.
Owner:BEIJING LINJIN SPACE AIRCRAFT SYST ENG INST

Goods grabbing and placing strategy generation method and system

The invention provides a method and system for generating a goods grabbing and placing strategy, and the method comprises the steps: carrying out the sampling processing of original video data, and obtaining a continuous video frame sequence; encoding the video frame sequence to generate a discretized visual vocabulary sequence; performing time sequence dependence modeling on the visual vocabulary sequence to learn a continuous change rule of object actions in video contents represented by the visual vocabulary sequence; on the basis of the visual vocabulary corresponding to the current state of the goods to be operated and the visual vocabulary corresponding to the expected target state, generating a potential action vector representing an action execution sequence required for transferring from the current state to the target state by utilizing a continuous change rule of object actions; and the potential action vector and the obtained real-time physical state information of the robot are mapped, and a control instruction for driving the mechanical arm to execute goods grabbing or placing operation is generated, so that the dependence on a large amount of manual annotation data is effectively reduced, and the self-adaption and generalization ability of the model to unseen scenes is improved.
Owner:SENAD TECH CO LTD

A bow graph matching method and system based on spectral clustering

The application discloses a BOW graph matching method and system based on spectral clustering, and the method comprises the following processes: extracting node features and topological features of a citation network graph; using an optimized K-means++ algorithm obtained by combining a spectral clustering algorithm with a genetic algorithm to optimize K values, to convert node features and topological feature descriptors of the citation network graph into words, and realizing construction of a dictionary; using a local constraint coding mode to code features of the dictionary, to obtain a visual vocabulary histogram; and classifying the visual vocabulary histogram, to realize the BOW graph matching method based on spectral clustering. The application uses the spectral clustering algorithm to cluster high-dimensional data sets, and then uses the K-means algorithm to perform two-stage clustering in a low-dimensional solution space, so that the problems of poor processing effect on high-dimensional data and low classification effect are solved.
Owner:XI'AN UNIVERSITY OF ARCHITECTURE AND TECHNOLOGY

"Action to Intent" AI Workflow for Translating Multimedia Edits into Multimodal Intent Explanations

A method and system for translating visual edits into personalized textual intent statements. The system extracts visual feature vectors characterizing design elements from unedited and edited digital assets, allows user correction of detected features to build personalized visual vocabularies, and generates joint multimodal representations stored as vector embeddings in a graph database. By maintaining comprehensive edit history, performing similarity-based retrieval of user-specific examples, and employing a multimodal large language model, the system generates intent statements that map manipulated design elements to intended design principles. This bridges communication between visually-oriented creators and non-visually-oriented audiences, reducing the burden of textual explanation for visual thinkers while improving communication quality in creative, educational, and professional workflows.
Owner:KELKAR UMA +1