Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

194 results about "Visual property" patented technology

Method of image processing for three-dimensional reconstruction in an extended reality environment and a head mounted display

Disclosed is method of image processing for three-dimensional reconstruction in an extended reality environment. The method includes defining a set of visual-attributes derived from reference images; identifying presence of at least one visual-attribute in a displayable content of target images to be used for the three-dimensional reconstruction; and modifying the displayable content of the target images, by concealing or displaying the identified at least one visual-attribute, for the three-dimensional reconstruction in the extended reality environment.
Owner:VARJO TECH OY

Image classification method and system based on attribute level optimal transmission alignment

The invention discloses an image classification method and system based on attribute-level optimal transmission alignment, and the method comprises the steps: dividing an input image into image blocks, adding a class mark vector, and forming the class mark vector and a visual feature vector through an encoder; secondly, based on the class mark vectors, visual attribute features are extracted, prompt statements are constructed, and text attribute features are obtained through a text encoder; then constructing an attribute-level cost matrix based on the visual attribute features and the text attribute features, and calculating to obtain attribute-level similarity between the visual attributes and the text attributes; according to the text attribute features, an attribute-level prediction probability is obtained through calculation, and a global feature prediction probability is obtained by applying a self-attention mechanism; and finally, carrying out weighted fusion on the attribute-level prediction probability and the global feature prediction probability to obtain a final category prediction probability. According to the method, collaborative learning of image semantics and attribute features is realized, so that image classification has higher accuracy and interpretability.
Owner:HANGZHOU DIANZI UNIV

Simulation scene generation method and device, electronic equipment and readable storage medium

The invention provides a simulation scene generation method and device, electronic equipment and a readable storage medium, and the method comprises the steps: carrying out the grid modeling of a target road element based on the real physical data of the target road element, and obtaining a grid-modeled target road element; performing Gaussian three-dimensional modeling on the grid modeled target road elements to obtain an initial Gaussian primitive set of the target simulation scene; and inputting the initial Gaussian primitive set and the target road video data into a trained simulation scene optimization model, optimizing visual attributes of each Gaussian primitive in the initial Gaussian primitive set, and obtaining a target simulation scene corresponding to the target road video data. Therefore, the geometric structure and the visual appearance of the object in the scene are decoupled, the geometric accuracy of the object is ensured through the grid model, and the visual attribute of the object is optimized through the trained simulation scene optimization model, so that the generation confidence and the generation efficiency of the simulation scene can be effectively improved.
Owner:BEIJING SAIMO TECH CO LTD

Systems, Methods, and Graphical User Interfaces for Scanning and Modeling Environments

A computer system displays graphical objects overlaying a representation of a field of view of one or more cameras, including displaying a first graphical object that represents one or more estimated spatial properties of a first physical feature that has been detected in a respective portion of the physical environment, and a second graphical object at that represents one or more estimated spatial properties of a second physical feature that has been detected in the respective portion of the physical environment. The computer system changes one or more visual properties of the first graphical object in accordance with variations in a respective predicted accuracy of the estimated spatial properties of the first physical feature, and changes the one more visual properties of the second graphical object in accordance with variations in a respective predicted accuracy of the estimated spatial properties of the second physical feature.
Owner:APPLE INC

System and Method for Managing Avatars for Use in Multiple 3D Rendering Platforms

A system includes a memory for storing a source digital-asset representation and at least one parameter table defining a non-linear mapping function. The system further includes a processor that is configured to receive context descriptors of a target rendering platform, execute an adaptive transformation engine that, in response to the context descriptors, applies the non-linear mapping function to convert geometry, materials, animation sets and physics attributes of the source digital asset into a target-platform representation, apply a stylization routine that remaps visual attributes in accordance with the context descriptors, and apply a precision-enhancement routine that increases a resolution of the target-platform representation to produce an adapted digital asset. The system further includes an output interface configured to supply the adapted digital asset to the target rendering platform at run-time.
Owner:TROY KELVIN JOHN +2

Computer vision property evaluation

Systems and methods are disclosed for computer vision property evaluation. In certain embodiments, a method may comprise executing a computer vision property evaluation operation via a computing system. The computer vision property evaluation operation may include identifying an image of a selected property feature using a first neural network (NN) of the computing system, cropping the image of the selected property feature in a selected way to produce a cropped image using a second NN of the computing system, generating a categorization of the cropped image based on identified details of the selected property feature, and generating a classification of a property corresponding to the selected property feature based on the categorization.
Owner:QUANTARIUM GRP LLC

Visual communication design evaluation system and method based on data analysis

The invention relates to the technical field of computer aided design and visual computing, and provides a visual communication design evaluation system and method based on data analysis. The method comprises the following steps: acquiring original design data, and extracting a spatial layout feature matrix, a color distribution feature vector and a semantic content feature set; calling a design specification knowledge base based on the design purpose information; performing simulated attention distribution analysis and design compliance analysis to generate corresponding features; constructing a transmission efficiency prediction model to generate a comprehensive transmission efficiency score and an initial design evaluation report; performing interference effect joint analysis to obtain a comprehensive interference effect quantitative evaluation result; and performing efficiency optimization correction on the initial design evaluation report based on the result, and outputting a final design evaluation report integrating spatial layout and visual attribute optimization suggestions. According to the method, the interference effect is accurately quantified through conjoint analysis of spatial layout and semantic attributes, and the accuracy and reliability of design conveying efficiency evaluation are improved.
Owner:CHANGCHUN ARCHITECTURE & CIVILENGEERING CO LLEGE

Methods and systems of tag location detection in an inventory environment based on visual attributes of tags using computer vision

A method comprises registering, by an application, a tag identifier received from a tag in an inventory environment with location data indicating a location of the tag based on a visual attribute of the tag, initiating, by the application, a scan of the tag to obtain tag data from the tag and to capture an image depicting the tag by transmitting a signal to the tag after registering the tag identifier with the location data of the tag, and triggering, by the application, activation of a light emitting diode (LED) on the tag to indicate whether a reader device is in a read range of the tag, in which a visual feature of the LED indicates whether the reader device is in the read range of the tag, and wherein the signal is used to activate the LED on the tag.
Owner:T MOBILE INNOVATIONS LLC

Multi-label field adaptive method based on prompt driving

The invention discloses a multi-label field adaptive method based on prompt driving, which belongs to the technical field of computer vision and comprises the following steps: generating multi-dimensional semantic text description containing visual attributes, semantic levels and category names of common scenes for each type; embedding the semantic text description into an optimizable category vector, and embedding the optimized category vector into a CLIP text prompt; extracting and projecting multi-layer style statistical characteristics of the image, and synchronously realizing image-text cross-modal alignment and source-target domain distribution alignment in a frame through a cross-domain style mapping network and various alignment losses; semantic priori knowledge of CLIP is combined with a label co-occurrence mode of a source domain, so that the model can sense and adapt to a possibly changed label dependency relationship in a target domain; performing semantic propagation on label embedding by using a graph convolutional network; the whole system is jointly optimized through a multi-task loss function, and end-to-end cross-domain multi-label classification is achieved.
Owner:NANJING UNIV OF POSTS & TELECOMM

Techniques for training identity transfer models

Systems, devices, and methods are provided for training and / or inferencing using machine-learning models. In at least one embodiment, an identity embedding and a content embedding are extracted from training media data, such as training audio data and / or training visual data. The identity embedding may encode information relating to a person's voice and / or visual properties. A decoder may be trained to create synthetic media, such as audio and visual content, based on identity and content embeddings.
Owner:AMAZON TECH INC

Text condition clothing image generation method and system based on diffusion model

The invention provides a text condition clothing image generation method and system based on a diffusion model, and the method comprises the steps: firstly obtaining a natural language input sequence containing clothing style types, fabric texture and decoration detail description elements, and carrying out the hierarchical semantic modeling processing of the natural language input sequence; generating a composite text guide vector containing the basic attribute features and the association relationship features, then generating an initial image tensor which has the same spatial size as the target clothing image and has random disturbance distribution, and taking the composite text guide vector as a conditional constraint; and performing multi-stage iterative denoising processing on the initial image tensor by using a pre-trained diffusion generation model to generate a continuous image sequence containing an intermediate generation state, and finally extracting an image in a final evolution state from the sequence as a target clothing image. Visual attributes and description elements of a natural language input sequence form a structured corresponding relation, and a garment image conforming to complex natural language description can be accurately generated.
Owner:SHANGHAI JIUZHIRUN INFORMATION TECH CO LTD

Vehicle-mounted three-dimensional scene rendering method and device, storage medium and program product

The invention provides a vehicle-mounted three-dimensional scene rendering method and device, a storage medium and a program product, and the method comprises the steps: obtaining the perception data of a target object, and constructing a nonlinear motion state space representing the motion process of the target object; performing prediction updating on the target object based on the state transition model of the nonlinear motion state space to obtain prediction state information and a prediction covariance matrix at the current rendering moment; extracting variance information related to the position from the covariance matrix to calculate an uncertainty factor, introducing the uncertainty factor into a graphic rendering pipeline, and dynamically adjusting at least one type of visual attribute of the three-dimensional model to realize uncertainty prompt; and when new perception data of the target object is obtained, performing fusion updating on the prediction state information, and performing smooth transition on a rendering result based on the correction amount to suppress rendering jump. According to the method, high-frame-rate continuous rendering display is realized under the condition of low-frequency updating of the sensor, and the interpretability and the safety of a rendering result are improved.
Owner:SHANGHAI GEOMETRICAL PERCEPTION & LEARNING CO LTD

Attitude-limitation-free Gaussian sputtering method for single-view 3D photography

The invention provides an attitude-limitation-free Gaussian sputtering method for single-view 3D photography, and the method comprises the steps: constructing an attention skipping attitude estimation network SPENet which comprises a large-kernel expansion convolution module LDC and an attention skipping feature interaction module SAFI, and achieving the unsupervised precise camera attitude estimation; a t distribution dual space Gaussian embedding model t-DSGE is constructed, Gaussian representation is expanded to a space and a visual attribute space so as to improve local detail presentation and scene coherence, and KL divergence loss is combined so as to maintain the coherence of the local detail presentation and the scene coherence. According to the method, unsupervised camera pose estimation is carried out by using a large-kernel expansion convolution module LDC and a skipping attention feature interaction module SAFI, the initialization process of a training scene is accelerated on the basis of ensuring the camera pose estimation precision, and the expression ability of 3DGS is enhanced by using a t-distribution dual space Gaussian embedding model t-DSGE, so that the accuracy of camera pose estimation is improved. And the precision of synthesizing the new view by the three-dimensional Gaussian sputtering method under the sparse pose-free view is improved.
Owner:NANJING UNIV OF SCI & TECH

Vehicle image-text retrieval method and system based on multistage semantic graph alignment and attribute enhancement

The invention belongs to the technical field of intelligent traffic, and relates to a vehicle image-text retrieval method and system based on multistage semantic graph alignment and attribute enhancement. According to the method, fusion visual embedding is obtained according to the vehicle video, and fusion text embedding is obtained according to the text data; obtaining an updated visual attribute embedding vector and an updated text attribute embedding vector based on fusion visual embedding and fusion text embedding; performing semantic matching on the updated visual attribute embedding vector and the updated text attribute embedding vector, and aligning vehicle cross-modal semantic attributes from three different levels by utilizing multi-granularity semantics to obtain final visual attribute embedding and final text attribute embedding; and according to the final visual attribute embedding and the final text attribute embedding, obtaining the similarity between the vehicle video and the artificial text description, wherein the vehicle video with the highest similarity is a vehicle video retrieval result. According to the method, the target vehicle can be quickly and accurately positioned, and higher robustness and higher retrieval success rate are shown.
Owner:CHANGAN UNIV

Ray tracing between AR and real objects

Aspects of the present disclosure involve a system for performing ray tracing between augmented reality (AR) and real-world objects. The system accesses, by the mobile device, a video depicting a first object. The system obtains, by the mobile device, a three-dimensional (3D) model of the first object. The system applies, by the mobile device, a ray tracing process to the 3D model of the first object to estimate an optical effect on a portion of the first object relative to a second object that is depicted in the video. The system modifies a visual property of the portion of the first object based on the optical effect relative to the second object.
Owner:SNAP INC

User interface for interacting with an affordance in an environment

Various implementations disclosed herein include devices, systems, and methods for indicating a distance to a selectable portion of a virtual surface. In various implementations, a device includes a display, a non-transitory memory and one or more processors coupled with the display and the non-transitory memory. In some implementations, a method includes displaying a graphical environment that includes a virtual surface, wherein at least a portion of the virtual surface is selectable. In some implementations, the method includes determining a distance between a collider object and the selectable portion of the virtual surface. In some implementations, the method includes displaying a depth indicator in association with the collider object. In some implementations, a visual property of the depth indicator is selected based on the distance between the collider object and the selectable portion of the virtual surface.
Owner:APPLE INC

Text-driven digital human audio and video generation method

The invention discloses a text-driven digital human audio and video generation method, relates to the technical field of intelligent virtual digital human generation, and provides the following scheme: obtaining a driving text and a reference picture, preprocessing the driving text to generate a phoneme sequence containing pronunciation time sequence information, and sending the phoneme sequence to the reference picture; and extracting key face feature points in the reference picture through a face feature point positioning technology, extracting tone reference feature vectors associated with tones in the reference picture through a deep learning model, and inputting the phoneme sequence into a text-to-speech model to generate an original audio. According to the method, the timbre feature vectors and the face feature points are synchronously extracted from the single reference picture, and a multi-modal feature cross fusion technology is combined, so that the problem of timbre and visual attribute splitting in a traditional method is solved, the unified digital human identity of personalized timbre and visual image is driven by only one picture, extra sound recording is not needed, and the user experience is improved. The use threshold is reduced.
Owner:HANGZHOU DIANZI UNIV

Fragment shader for creator signatures in a virtual asset marketplace

Various implementations relate to methods, systems, and computer-readable media to dynamically apply creator signatures to virtual assets within a virtual platform. According to one aspect, a computer-implemented method includes receiving a request to display a virtual asset associated with a creator, where the request is linked to an avatar in a virtual environment hosted on the platform. A creator signature, which is defined by visual properties unique to the creator, is retrieved and stored separately from the virtual asset. The method includes rendering the virtual asset by rasterizing it and applying a fragment shader to overlay the creator signature as a dynamic visual element that changes over time based on predefined parameters. The rendered virtual asset with the dynamic signature is displayed within the virtual environment, providing visual attribution to the creator while preventing unauthorized embedding of the signature in the asset data.
Owner:ROBLOX CORP

A weakly supervised small sample object detection system and method based on localization pre-training

The present application belongs to the technical field of machine learning, and particularly relates to a weakly supervised small sample target detection method and system based on positioning pre-training and progressive optimization strategy. The present application introduces a weakly supervised learning mechanism into a small sample deep target detection framework, and establishes a weakly supervised small sample target detection system with high accuracy. The method framework of the present application is simple, convenient to use, strong in scalability and strong in interpretability, and the weakly supervised small sample target detection results on two mainstream visual attribute data sets all exceed those of existing methods. The present application can provide support for basic frameworks and algorithms for target detection technology in military and industrial application fields, and can also be easily extended to other small sample learning tasks.
Owner:FUDAN UNIVERSITY

Image processing method and device, electronic equipment and storage medium

The embodiment of the invention discloses an image processing method and device, electronic equipment and a storage medium, which can be applied to various scenes such as cloud technology, artificial intelligence, intelligent traffic and auxiliary driving, and the method comprises the steps: firstly obtaining a to-be-processed target image, carrying out the depth detection of the target image, and obtaining a depth image of the target image, the depth image comprises the depth value of each pixel point in the target image, and then rendering the target image to a set display area according to the depth image, so that the target image is more stereoscopic; the method comprises the following steps: acquiring a depth image, dynamically adjusting a visual attribute parameter of each pixel point according to a depth value of each pixel point contained in the depth image to obtain an adjusted target image, and rendering the adjusted target image to a set display area to form a three-dimensional dynamic image. According to the technical scheme, the user experience and the image processing efficiency can be improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Video generation method and device, storage medium and computer program product

The embodiment of the invention discloses a video generation method. The method comprises the steps of obtaining a to-be-processed video and visual attribute information of a user; the visual attribute information represents the focusing capability of eyes of the user on light; on the basis of the visual attribute information, determining parallax information of pixels of each frame of image in the to-be-processed video; determining a plurality of target images based on the parallax information and each frame of image; the target three-dimensional video is determined based on the plurality of target images, so that the problem that the generated 3D video is not matched with the visual characteristics of the user when the 3D video is generated in the related technology is solved, and the watching experience of the user is improved. The embodiment of the invention further discloses video generation equipment, a storage medium and a computer program product.
Owner:MIGU CO LTD +1

Leakage detection model construction method, leakage detection method and device

The invention provides a leakage detection model construction method and a leakage detection method and device, and the method comprises the steps: carrying out the preprocessing of continuous multi-frame original images of to-be-detected equipment, and obtaining a preprocessed image; performing feature extraction operation on the preprocessed image to obtain a first feature map of the preprocessed image; multiple different types of visual attribute values in the preprocessed image are determined, and the sizes of the visual attribute values are in positive correlation with the leakage possibility of the to-be-detected equipment; according to the visual attribute values, target visual features corresponding to the visual attribute values are extracted from the first feature map, a second feature map is generated based on the target visual features, and the larger the visual attribute values are, the larger the contribution degree of the corresponding target visual features to the second feature map is; based on the second feature map, determining the leakage probability of the to-be-detected equipment; and constructing a leakage detection model based on the leakage probability. The method can reduce the omission ratio of tiny leakage.
Owner:YILIAN CLOUD COMPUTING (HANGZHOU) CO LTD +1

Combined zero sample recognition method based on shared learnable soft query vector and related equipment

The invention discloses a combined zero sample recognition method based on a shared learnable soft query vector and related equipment. The method comprises the following steps: acquiring an initial visual feature, an initial text attribute feature, an initial text object feature and an initial text combination feature; performing visual alignment and decoupling by sharing the soft query vector to obtain a target visual attribute feature, a target visual object feature and a target visual combination feature; performing text alignment and decoupling by sharing the soft query vector to obtain a target text attribute feature, a target text object feature and a target text combination feature; obtaining the similarity between the target visual attribute feature and the target text attribute feature, the similarity between the target visual object feature and the target text object feature, and the similarity between the target visual combination feature and the target text combination feature; and carrying out weighted fusion on the similarity to obtain a combined zero sample recognition result. The method can improve the recognition capability of unseen combinations, and can be widely applied to the technical field of computer vision.
Owner:GUANGZHOU UNIVERSITY

Product image generation based on diffusion model

Methods and systems are disclosed for generating an extended reality (XR) try-on experience based on an image produced by a diffusion model. The system receives an image depicting a real-world object and generates a prompt comprising a textual description of a fashion item. The system analyzes the image and the textual description of the fashion item using a generative machine learning model to generate an artificial image that depicts an artificial object that resembles the real-world object wearing an artificial fashion item matching the textual description of the fashion item. The system identifies an object comprising a real-world product image that matches visual attributes of the artificial fashion item and replaces the artificial fashion item in the artificial image with the object to generate an output image.
Owner:SNAP INC

Remote sensing image semi-supervised change detection method and system based on frequency domain convolution and local feature enhancement

The invention discloses a remote sensing image semi-supervised change detection method based on frequency domain convolution and local feature enhancement. The method mainly comprises a frequency domain convolution branch and a local feature enhancement module. On the basis that deep features and shallow features are extracted by using a convolutional neural network, deep frequency domain features and deep spatial domain features are extracted and interacted by using frequency domain convolution branch, so that a global dependency relationship is captured and difference features are highlighted; and through a local feature enhancement module, self-similarity in shallow features is further aggregated, rich and fine visual attributes and structural information are extracted from a local region, and the ability of the model to extract valuable geometric and visual attributes is enhanced. Meanwhile, the features processed by the frequency domain convolution branch and local feature enhancement module are fused with the original features, and the feature learning ability of the model to the unlabeled data is further enhanced. And finally, inputting the multi-scale fusion into a classifier to realize semi-supervised change detection of the high-resolution remote sensing image.
Owner:WUHAN UNIV

Leakage detection model construction method, leakage detection method and device

This application provides a method for constructing a leakage detection model, a leakage detection method, and an apparatus, comprising: preprocessing multiple consecutive frames of original images of a device to be detected to obtain a preprocessed image; performing feature extraction on the preprocessed image to obtain a first feature map of the preprocessed image; determining multiple different types of visual attribute values ​​in the preprocessed image, wherein the magnitude of the visual attribute values ​​is positively correlated with the probability of leakage in the device to be detected; extracting target visual features corresponding to each visual attribute value from the first feature map according to the magnitude of each visual attribute value, and generating a second feature map based on the target visual features, wherein the larger the visual attribute value, the greater the contribution of the corresponding target visual feature to the second feature map; determining the leakage probability of the device to be detected based on the second feature map; and constructing a leakage detection model based on the leakage probability. This method can reduce the false negative rate of minor leaks.
Owner:YILIAN CLOUD COMPUTING (HANGZHOU) CO LTD +1

A method for cluttered scene object grasping based on visual-linguistic-action joint modeling

The application discloses a method for cluttered scene target object grasping based on visual-language-action joint modeling. The application uses object-centered representation to realize a method for cluttered scene target object grasping based on visual-language-action joint modeling, processes the object-centered representation through a pre-trained visual-language model and a grasping model, obtains visual-language features and grasping features of each bounding box, and uses a transformer to implement cross-attention mechanisms among visual-language-action multimodality, generates visual-language-action cross-attention features, and then generates decisions and executes, so that higher sample utilization is realized, and additional data collection and training in the simulation-physical migration process are avoided; compared with a two-stage strategy, visual attributes and planner screening rules for language-visual matching need not be artificially designed, so that more flexible language instructions can be adapted, and better task generalization is achieved.
Owner:ZHEJIANG UNIV

Audio frequency spectrum fluidization interaction presentation method and system based on shader computing power

The invention discloses an audio frequency spectrum fluidization interaction presentation method and system based on shader computing power, and belongs to the technical field of image processing, and the method comprises the steps: obtaining an audio stream and user interaction data extraction features in real time; generating a sound interaction total dynamic field in the GPU parallel computing shader through nonlinear coupling; inputting the field into a preset neural network model in a fragment shader for forward reasoning to complete fluid neural physical evolution; separating an audio semantic layer in the rendering pipeline and executing competitive visual attribute mapping to generate an image; and performing asynchronous scheduling on each processing pipeline through an asynchronous calculation engine. According to the method, full-pipeline GPU calculation and neural operator evolution are adopted, the contradiction between fluid simulation and real-time rendering can be effectively solved, frequent data transmission is abandoned, low-delay interaction feedback is achieved while the physical reality sense is guaranteed, and the immersion sense and expressive force of audio visualization are remarkably improved.
Owner:CHENGDU LIBI TECH CO LTD

Multi-level intelligent agent construction method for regional integrated energy system

The invention belongs to the technical field of energy intelligence, and provides a regional integrated energy system multi-level agent construction method, which comprises the following steps: carrying out space-time alignment processing on real-time operation data to obtain a synchronous data stream; respectively mapping parameters in the synchronous data stream into specific visual attribute primitives to obtain a multi-level state semantic image; performing visual semantic segmentation on the multi-level state semantic image to obtain visual semantic units, and constructing a spatial correlation map between the visual semantic units; performing visual difference comparison on the spatial correlation map and a historical same-period spatial correlation map to obtain an abnormal visual semantic unit; performing multi-frame tracking verification on the abnormal visual semantic unit, and integrating an abnormal unit identifier and an abnormal type into an abnormal event report; and carrying out hierarchical decision making on the abnormal event report and the spatial correlation graph to obtain a collaborative regulation and control instruction. According to the invention, the efficiency of cooperative management of the regional integrated energy system can be improved.
Owner:STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO +1