Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

29 results about "Visual property" patented technology

Fragment shader for creator signatures in a virtual asset marketplace

Various implementations relate to methods, systems, and computer-readable media to dynamically apply creator signatures to virtual assets within a virtual platform. According to one aspect, a computer-implemented method includes receiving a request to display a virtual asset associated with a creator, where the request is linked to an avatar in a virtual environment hosted on the platform. A creator signature, which is defined by visual properties unique to the creator, is retrieved and stored separately from the virtual asset. The method includes rendering the virtual asset by rasterizing it and applying a fragment shader to overlay the creator signature as a dynamic visual element that changes over time based on predefined parameters. The rendered virtual asset with the dynamic signature is displayed within the virtual environment, providing visual attribution to the creator while preventing unauthorized embedding of the signature in the asset data.
Owner:ROBLOX CORP

Leakage detection model construction method, leakage detection method and device

This application provides a method for constructing a leakage detection model, a leakage detection method, and an apparatus, comprising: preprocessing multiple consecutive frames of original images of a device to be detected to obtain a preprocessed image; performing feature extraction on the preprocessed image to obtain a first feature map of the preprocessed image; determining multiple different types of visual attribute values ​​in the preprocessed image, wherein the magnitude of the visual attribute values ​​is positively correlated with the probability of leakage in the device to be detected; extracting target visual features corresponding to each visual attribute value from the first feature map according to the magnitude of each visual attribute value, and generating a second feature map based on the target visual features, wherein the larger the visual attribute value, the greater the contribution of the corresponding target visual feature to the second feature map; determining the leakage probability of the device to be detected based on the second feature map; and constructing a leakage detection model based on the leakage probability. This method can reduce the false negative rate of minor leaks.
Owner:YILIAN CLOUD COMPUTING (HANGZHOU) CO LTD +1

3D model generation using multiple textures

Methods and systems for generating 3D assets for use in, for example, extended reality (XR) experiences are disclosed. A system receives a plurality of textures associated with an object, each of the plurality of textures corresponding to a different view of the object, and automatically generates an initial three-dimensional (3D) model of the object based on an initial alignment of the plurality of textures to respective portions of the initial 3D model. The system receives input adjusting the initial alignment of the plurality of textures to the respective portions of the 3D model, and combines the plurality of textures into a single texture based on the input, the single texture defining visual properties of the object from the plurality of views. The system stores the 3D model in association with the single texture.
Owner:SNAP INC

Dynamic relationship-based avatar garment generation

PCT designated stageWO2026147888A1PersonalizationComputer graphics (images)
The described system facilitates personalized avatar interactions by dynamically generating and displaying modified garments. The system determines the initiation of an interaction function by a first user with a second user within an interaction platform. It accesses avatar data for the first user, including visual attributes and a garment associated with the first avatar, as well as avatar data for the second user, including visual attributes and a garment associated with the second avatar. An image is generated featuring the first avatar wearing its garment alongside the second avatar wearing its garment. This image is applied to the first avatar's garment, creating a modified version that visually incorporates the second avatar. The system then displays the first avatar wearing the modified garment alongside the second avatar wearing its original garment, enabling enhanced user engagement through visually personalized and contextual representations of user interactions.
Owner:SNAP INC

User interface generation methods, computing systems, and storage media

This specification provides a method, computing system, and storage medium for generating a user interface. In the method, the computing system receives an interface generation request from a client. This request includes target text described by an operator in natural language. An intelligent agent group then performs an interface generation task based on the target text and a semantic knowledge base to obtain at least one user interface. The semantic knowledge base includes predefined information for multiple interface elements, each comprising multiple interface components and multiple visual attributes. The predefined information for each interface element characterizes its style and usage description.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

A low-light image enhancement quality evaluation method based on multi-modal and multi-annotation

PendingCN122454373AData setImage manipulation
The present application relates to the field of image processing and quality evaluation, and particularly relates to a low-light image enhancement quality evaluation method based on multi-modal and multi-annotation, comprising the following steps: constructing a low-light image enhancement quality evaluation data set; inputting the low-light enhancement image and its natural language description into a feature extraction module to obtain global visual features, global text features, visual features and semantic features of five quality attributes of brightness, color, noise, exposure and naturalness; constructing a visual attribute graph and a text attribute graph based on the global visual features, the global text features, the visual features and the semantic features, and extracting visual attribute representations and semantic attribute representations through a graph attention network; obtaining fusion features through a cross-modal alignment and fusion module; inputting the fusion features into a regression module to generate a quality score; the present application can accurately evaluate the quality of low-light images and provide an interpretable evaluation basis according to the influence of each attribute on the final quality.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Systems and methods for navigating an extended reality history

In an example system for navigating extended reality history, the system captures and stores a plurality of snapshots of one or more extended reality sessions. The system retrieves the snapshots and identifies a plurality of entities within the snapshots. The system determines that a degree of similarity between visual attributes of a first snapshot and visual attributes of the snapshots temporally adjacent to the first snapshot is lower than a degree of similarity between visual attributes of a second snapshot and visual attributes of the snapshots temporally adjacent to the second snapshot. Based at least in part on the determined lower degree of similarity of the first snapshot, the system assigns a higher weight to the first snapshot than to the second snapshot. Based on the identified entities and the assigning, the system identifies at least one salient snapshot to generate for presentation.
Owner:ADEIA GUIDES INC

A method, device, equipment, medium and program product for detecting a fake video

The application discloses a counterfeit video detection method, device, equipment, medium and program product. The method comprises the following steps: acquiring at least two target image frames of a to-be-detected video; wherein the target image frames are obtained by screening based on the differences between different image frames in the to-be-detected video; acquiring image description information of the target image frames; performing a preset each authenticity detection task on target detection data of the to-be-detected video to obtain a detection result corresponding to each authenticity detection task; wherein the target detection data at least comprises the image description information; and determining a counterfeit detection result of the to-be-detected video based on each detection result. The application can identify the authenticity of the content of the to-be-detected video, avoid the training set coverage limitation caused by recognizing the visual attributes based on the face and the picture, has better generalization, and can further improve the accuracy of detecting various counterfeit videos.
Owner:CHINA MOBILE INTERNET CO LTD +1

Image-text recognition translation system based on semantic recognition

The application discloses a picture-text recognition and translation system based on semantic recognition, which comprises a document analysis module, an image local text recognition module, a context perception module, a translation instruction generation module and a picture-text reorganization module; the document analysis module can separate original picture-text mixed arrangement documents into continuous text sets and image region sets; the image local text recognition module can obtain local text segments and determine basic semantic vectors by processing the image region sets; the context perception module can calculate the semantic correlation degree S of the basic semantic vectors and each paragraph vector; the translation instruction generation module can generate translation input sequences; and the picture-text reorganization module can generate translated image sets through the visual attributes and position information of original image regions, and then obtain bilingual translation documents. The application can make full use of complete semantic contexts provided by document main texts by the context perception module to realize the translation of image local texts.

A method for learning multi-view auxiliary representation and moment retrieval and related devices

This invention belongs to the field of computer vision and pattern recognition technology, and discloses a time retrieval method and related apparatus for learning multi-view auxiliary representations, aiming to solve the technical problem of unreliable time retrieval results in existing multimodal retrieval methods. The technical solution of this invention includes: constructing a multi-view auxiliary representation set based on text features; injecting target-specific visual attributes from the video into the auxiliary representation while maintaining semantic consistency to obtain an auxiliary representation that fuses visual attributes; achieving cross-modal association between the auxiliary representation and video features through multi-view semantic alignment to obtain enhanced visual features; and completing target time location through a detection head optimized by a loss function. The technical solution disclosed in this invention, through a multi-view auxiliary representation framework, actively injects visual evidence into the language representation, reduces the risk of overfitting sparse text cues, establishes semantic alignment of visual perception, and significantly improves the robustness and localization accuracy of time retrieval.
Owner:XI AN JIAOTONG UNIV

A Road Defect Retrieval and Diagnosis Method Based on Multimodal Knowledge Graph

This invention discloses a road defect retrieval and diagnosis method based on a multimodal knowledge graph, comprising the following steps: S1: constructing a multimodal knowledge graph of road defects; S2: extracting features based on visual-semantic joint embedding; S3: storing the mapped visual vectors as visual attributes of entities in the graph structure, and performing multimodal fusion alignment; S4: performing cross-modal retrieval, multi-hop reasoning, and logical error correction. The beneficial effects of this invention are: by constructing a multimodal knowledge graph of road defects with visual feature nodes, the alignment of the visual feature space and the semantic feature space is achieved. This allows for the identification of defects while simultaneously using the knowledge graph for reasoning, outputting the causes of defects and maintenance measures, and correcting false detections in visual recognition, thus achieving accurate cross-modal retrieval and intelligent diagnosis.
Owner:成都圭目机器人有限公司 +2

Meta-learning based multi-modal large model task-driven target detection method and system

PendingCN122263000ATaking into account accuracyConsider flexibilityBiological modelsLinguistic modelSemantic representation
The application provides a multi-modal large model task-driven target detection method and system based on meta learning, relates to the technical field of target detection, and comprises the following steps: obtaining multi-modal input data and performing preprocessing to obtain unified semantic representation; a general candidate generation network and a large language model are used to generate a candidate region set and a visual attribute word group set; region feature extraction is performed on each candidate target frame in the candidate region set to obtain a candidate visual embedding sequence; the visual attribute word group set is subjected to vectorization processing to obtain an attribute text embedding sequence; the candidate visual embedding sequence and the attribute text embedding sequence are jointly calibrated to obtain a calibrated candidate visual embedding sequence and a calibrated attribute text embedding sequence; a task-related score is calculated through a trainable scoring function, and all candidate frames in the candidate region set are screened to obtain a target detection result. The application realizes target detection guided by task description and has the ability of cross-task rapid adaptation.
Owner:HUBEI UNIV OF TECH

A Method and System for Audio Spectrum Fluidized Interactive Presentation Based on Shader Computation Power

This invention discloses an audio spectrum fluidized interactive presentation method and system based on shader computing power, belonging to the field of image processing technology. It includes: real-time acquisition of audio streams and user interaction data to extract features; generation of the total dynamic field of audio interaction through nonlinear coupling in a GPU parallel computing shader; inputting this field into a preset neural network model for forward inference in a fragment shader to complete the fluid neurophysical evolution; separating the audio semantic layer in the rendering pipeline and performing competitive visual attribute mapping to generate images; and asynchronous scheduling of each processing pipeline through an asynchronous computing engine. This invention employs full-pipeline GPU computing and neural operator evolution, effectively resolving the contradiction between fluid simulation and real-time rendering, eliminating frequent data transmission, and achieving low-latency interactive feedback while ensuring physical realism, significantly improving the immersiveness and expressiveness of audio visualization.
Owner:CHENGDU LIBI TECH CO LTD

A cross-modal fashion commodity retrieval method and system based on multi-granularity feature fusion and a storage medium

ActiveCN121614632BAlgorithmEngineering
The application discloses a kind of multi-granularity feature fusion-based cross-modal fashion commodity retrieval method, system and storage medium, and the structured semantic enhancement module is converted into natural language description and extracts key visual attribute word sequence by commodity label;Attribute perception double-tower feature extraction network is constructed, wherein image tower adopts hierarchical visual Transformer to capture detail texture and global semantics, and text tower introduces attribute weighting layer based on static dictionary to strengthen visual related vocabulary;Independent pair probability alignment module is used, with pair Sigmoid loss function replacing traditional Softmax loss, independently optimizing the matching probability of each pair of image-text sample, significantly reducing memory occupation and adapting multi-label data;Attribute hard filtering and vector soft matching are fused in online retrieval stage, and retrieval accuracy and efficiency are considered.This application effectively improves the identification ability of fashion commodity fine-grained attribute and the accuracy and response speed of cross-modal retrieval.
Owner:YUZHEN (SHANGHAI) INFORMATION TECHNOLOGY CO LTD

Eddy current (EC) defect view and defect classification

Presentations showing eddy current (EC) test results can be generated and displayed. For example, amplitude or phase values ​​associated with an eddy current measurement signal can be presented graphically, such as by aligning an index of the amplitude or phase value with a shape representing the object under test, where the index corresponds to the location on or within the shape from which the eddy current measurement signal was obtained. In another example, the visual attributes of the displayed index of the amplitude or phase value can be assigned based on the defect class, such as using different colors corresponding to different defect classes, when the visual attributes are defined by regions within the impedance plane. The use of attributes to identify defects can be carried out either separately or in combination with alignment with the shape of the index.
Owner:EVIDENT CANADA INC

Interactive image generation method, device, and medium

The present disclosure relates to an interactive image generation method, device and medium, the method comprising: obtaining a first character draft of a first virtual character; wherein the first character draft is generated based on extracting a character visual element corresponding to each character visual attribute in a first character building attribute group from a first character description text of the first virtual character, the character visual attribute being constructed based on the first character building attribute group comprising a plurality of character visual attributes for defining the appearance of the first virtual character; generating a second character draft of a second virtual character based on the first character draft and a second character description text of the second virtual character; calling an image generation model to generate an interactive image containing the first virtual character and the second virtual character interacting in a set interactive state based on the first character draft and the second character draft.
Owner:YIDIAN LINGXI INFORMATION TECHNOLOGY (GUANGZHOU) CO LTD

Generative AI to create time series prediction radiotherapy treatment planning systems

ActiveUS12670993B2Data setEngineering
A server monitors a sequence of screen captures from a radiotherapy treatment planning platform's user interface, operated by medical professionals. The server generates a training dataset comprising a time series of the screen captures and trains a machine learning model using this dataset. The trained model is configured to predict visual attributes of the user interface, determining the next screen's attributes based on prior interactions. When executed, the model predicts future visual attributes of the interface as a user interacts with the current screen, offering real-time guidance for navigating the treatment planning platform.
Owner:SIEMENS HEALTHINEERS INTERNATIONAL AG

Creative video generation method and system based on open-source AI agent

The application provides a creative video generation method and system based on an open-source AI agent, and relates to the technical field of artificial intelligence. First, a creative demand description text stream and a material resource storage path set are obtained, the former including a plurality of description text segment units with continuous input and semantic pointing markers, and the latter including corresponding external material storage address information. Then, semantic analysis and creative element disassembly are performed on the creative demand description text stream to generate narrative logic structure features and visual attribute constraint features. Next, an open-source agent collaborative arrangement framework is called to perform multi-agent collaborative task distribution and execution arrangement to generate a creative execution instruction sequence set. The creative execution instruction sequence set is used to drive a functional agent instance to generate a split-screen video segment data set. Finally, splicing is performed based on the narrative logic structure features to generate a creative video data stream consistent with the creative demand semantics, greatly improving the efficiency and quality of creative video generation.
Owner:SHANGHAI YUNQUE INTELLIGENT TECH CO LTD

Mapping color to data for data bound objects

ActiveUS12670636B2GraphicsData class
Embodiments are disclosed for binding colors to data visualizations on a digital canvas. In some embodiments, a method of binding colors to data visualizations includes receiving a data set including data associated with a variable. A chart, including a plurality of graphic objects, is generated based on the variable of the data set and a visual property of the plurality of graphic objects. A data type associated with the variable determined and first colors are assigned to the plurality of graphic objects based on the data type using a color binding. A selection of second colors to be assigned to the plurality of graphic objects is received and the chart is updated using the second colors.
Owner:ADOBE INC

Railway hand signal processing method and system based on multi-modal perception fusion and storage medium

The application relates to the technical field of computer vision, in particular to a railway hand signal processing method and system based on multi-modal perception fusion and a storage medium. The multi-modal hand signal instruction category determination is performed by fusing the posture angle parameters, the trajectory parameters and the types and visual attributes of signal equipment, compared with the scheme of single reliance on human posture estimation or pixel-level template comparison, the multi-modal features complement each other; by calling a standard action parameter set corresponding to the recognized hand signal instruction category, the action specification is converted into quantitative evaluation indexes of five dimensions of completion degree, speed, posture angle accuracy, equipment use correctness and visibility, a score result is output, and reliable quantitative evaluation of the hand signal is realized. The application aims to solve the problem of how to realize reliable quantitative evaluation of hand signal actions.
Owner:KUNMING UNIV OF SCI & TECH

Dynamic relationship-based avatar garment generation

The described system facilitates personalized avatar interactions by dynamically generating and displaying modified garments. The system determines the initiation of an interaction function by a first user with a second user within an interaction platform. It accesses avatar data for the first user, including visual attributes and a garment associated with the first avatar, as well as avatar data for the second user, including visual attributes and a garment associated with the second avatar. An image is generated featuring the first avatar wearing its garment alongside the second avatar wearing its garment. This image is applied to the first avatar's garment, creating a modified version that visually incorporates the second avatar. The system then displays the first avatar wearing the modified garment alongside the second avatar wearing its original garment, enabling enhanced user engagement through visually personalized and contextual representations of user interactions.
Owner:SNAP INC

Diffusion model multi-person image generation

Methods and systems for generating personalized images using one or more diffusion models are disclosed. The methods and systems access first and second artificially personalized images generated by first and second generative machine learning models, wherein the first generative machine learning model is trained to generate a first artificially personalized image containing a first person's depiction, and the second generative machine learning model is trained to generate a second artificially personalized image containing a second person's depiction. The methods and systems generate a foreground image by combining the first person's depiction in the first artificially personalized image and the second person's depiction in the second artificially personalized image. The methods and systems access and generate a new artificial image containing the foreground image on a background having visual attributes corresponding to background information.
Owner:SNAP INC

Systems and methods for improved creation of extended reality worlds, experiences, simulations and learning activities

A system and method enable creation and modification of digital objects within an extended reality (XR) environment using voice commands in combination with detected physical user inputs. An XR hardware device detects a spoken command requesting creation or modification of a digital object and detects a physical input indicative of a spatial location within the XR environment. Spatial coordinates corresponding to the physical input are determined, and a representation of the digital object is displayed at the determined spatial coordinates. Visual attributes of the digital object may be assigned or modified based on additional spoken instructions and user interactions, including gesture-based inputs and measurement selection interfaces.
Owner:CURIOXR INC

Image processing method, electronic device and cloud server

The application discloses an image processing method, an electronic device and a cloud server, and belongs to the technical field of artificial intelligence. The method is executed by the electronic device, and the method comprises the following steps: obtaining identity recognition information of a first object in a first image and image description information of the first image; outputting object description information of the first image based on the identity recognition information and the image description information; the object description information is used for describing the identity and visual attribute information of the first object; wherein the identity recognition information is obtained by performing identity recognition on the first object by the electronic device; and the image description information is obtained by performing image recognition on the first image by the cloud server.
Owner:VIVO MOBILE COMM CO LTD

Knowledge editing evaluation sample construction method and system in a multi-modal scenario

The application discloses a kind of multi-modal scene under knowledge editing evaluation sample construction method and system, belong to artificial intelligence technical field.The method is aimed at each knowledge point, first, filter triplets visual question and answer sample as reliability sample;Further, by modifying the fine-grained attribute of target entity in the sample, generate generalization sample for testing deep reasoning;Introduce the dual similarity filtering mechanism of text semantics and visual attribute, accurately filter irrelevant candidate samples from large-scale multi-modal problem library as locality sample;Finally, three kinds of samples are packaged as standardized test tuples, forming the evaluation sample set covering reliability, generalization and locality.The present application realizes the whole process automation construction, solves the limitation of single evaluation dimension and dependence on artificial annotation, and provides an efficient benchmark for accurately evaluating the knowledge editing effect of multi-modal large model.
Owner:ZHEJIANG UNIV +1

Artificial intelligence-based graphic story books conversion

PCT designated stageWO2026143202A1GraphicsAlgorithm
A system and method for converting graphic story books into multi-dimensional content using artificial intelligence is disclosed. Features are extracted from a graphic story book that includes a sequence of image panels. The features include image features representing objects within each panel and text features representing textual content. An AI model is used to enrich the content based on the extracted features by generating audio content that includes verbal descriptions of scenes and utterances from characters. Visual attributes of characters are analyzed to determine appropriate voice profiles for each character, and speeches are generated for the characters based on their respective voice profiles. Video clips are created by combining image panels that represent sequences of related movements. The enriched content includes a time dimension where different portions are presented according to a timeline.
Owner:TRILOGY 1 LLC

Artificial intelligence-based graphic story books conversion

PendingUS20260189768A1GraphicsAlgorithm
A system and method for converting graphic story books into multi-dimensional content using artificial intelligence is disclosed. Features are extracted from a graphic story book that includes a sequence of image panels. The features include image features representing objects within each panel and text features representing textual content. An AI model is used to enrich the content based on the extracted features by generating audio content that includes verbal descriptions of scenes and utterances from characters. Visual attributes of characters are analyzed to determine appropriate voice profiles for each character, and speeches are generated for the characters based on their respective voice profiles. Video clips are created by combining image panels that represent sequences of related movements. The enriched content includes a time dimension where different portions are presented according to a timeline.
Owner:TRILOGY1 LLC

Object detection based on text input that includes both target object classes and target visual attributes

Implementations improve object classification / detection by leveraging visual attributes. An image depicting instance(s) of object class(es) is obtained with a textual snippet that includes: noun(s) identifying target object class(es); and adjective(s) describing target visual attribute(s). The textual snippet may be encoded as text embedding(s) that represent target object class(es) and visual attribute(s) in a shared embedding space. The image may be processed using an image encoder to generate image encoder output tokens (IEOTs) that are used to generate object visual embedding(s) in the shared embedding space. The text embedding(s) and the object visual embeddings may be used to classify the IEOTs as depicting an instance of the target object class(es) having target visual attribute(s). The IEOTs may also be processed using a localization head to predict annotation(s) for the digital image.
Owner:DEERE & CO