Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

160 results about "Visual property" patented technology

Image classification method and system based on attribute level optimal transmission alignment

The invention discloses an image classification method and system based on attribute-level optimal transmission alignment, and the method comprises the steps: dividing an input image into image blocks, adding a class mark vector, and forming the class mark vector and a visual feature vector through an encoder; secondly, based on the class mark vectors, visual attribute features are extracted, prompt statements are constructed, and text attribute features are obtained through a text encoder; then constructing an attribute-level cost matrix based on the visual attribute features and the text attribute features, and calculating to obtain attribute-level similarity between the visual attributes and the text attributes; according to the text attribute features, an attribute-level prediction probability is obtained through calculation, and a global feature prediction probability is obtained by applying a self-attention mechanism; and finally, carrying out weighted fusion on the attribute-level prediction probability and the global feature prediction probability to obtain a final category prediction probability. According to the method, collaborative learning of image semantics and attribute features is realized, so that image classification has higher accuracy and interpretability.
Owner:HANGZHOU DIANZI UNIV

Simulation scene generation method and device, electronic equipment and readable storage medium

The invention provides a simulation scene generation method and device, electronic equipment and a readable storage medium, and the method comprises the steps: carrying out the grid modeling of a target road element based on the real physical data of the target road element, and obtaining a grid-modeled target road element; performing Gaussian three-dimensional modeling on the grid modeled target road elements to obtain an initial Gaussian primitive set of the target simulation scene; and inputting the initial Gaussian primitive set and the target road video data into a trained simulation scene optimization model, optimizing visual attributes of each Gaussian primitive in the initial Gaussian primitive set, and obtaining a target simulation scene corresponding to the target road video data. Therefore, the geometric structure and the visual appearance of the object in the scene are decoupled, the geometric accuracy of the object is ensured through the grid model, and the visual attribute of the object is optimized through the trained simulation scene optimization model, so that the generation confidence and the generation efficiency of the simulation scene can be effectively improved.
Owner:BEIJING SAIMO TECH CO LTD

Systems, Methods, and Graphical User Interfaces for Scanning and Modeling Environments

A computer system displays graphical objects overlaying a representation of a field of view of one or more cameras, including displaying a first graphical object that represents one or more estimated spatial properties of a first physical feature that has been detected in a respective portion of the physical environment, and a second graphical object at that represents one or more estimated spatial properties of a second physical feature that has been detected in the respective portion of the physical environment. The computer system changes one or more visual properties of the first graphical object in accordance with variations in a respective predicted accuracy of the estimated spatial properties of the first physical feature, and changes the one more visual properties of the second graphical object in accordance with variations in a respective predicted accuracy of the estimated spatial properties of the second physical feature.
Owner:APPLE INC

System and Method for Managing Avatars for Use in Multiple 3D Rendering Platforms

A system includes a memory for storing a source digital-asset representation and at least one parameter table defining a non-linear mapping function. The system further includes a processor that is configured to receive context descriptors of a target rendering platform, execute an adaptive transformation engine that, in response to the context descriptors, applies the non-linear mapping function to convert geometry, materials, animation sets and physics attributes of the source digital asset into a target-platform representation, apply a stylization routine that remaps visual attributes in accordance with the context descriptors, and apply a precision-enhancement routine that increases a resolution of the target-platform representation to produce an adapted digital asset. The system further includes an output interface configured to supply the adapted digital asset to the target rendering platform at run-time.
Owner:TROY KELVIN JOHN +2

Visual communication design evaluation system and method based on data analysis

The invention relates to the technical field of computer aided design and visual computing, and provides a visual communication design evaluation system and method based on data analysis. The method comprises the following steps: acquiring original design data, and extracting a spatial layout feature matrix, a color distribution feature vector and a semantic content feature set; calling a design specification knowledge base based on the design purpose information; performing simulated attention distribution analysis and design compliance analysis to generate corresponding features; constructing a transmission efficiency prediction model to generate a comprehensive transmission efficiency score and an initial design evaluation report; performing interference effect joint analysis to obtain a comprehensive interference effect quantitative evaluation result; and performing efficiency optimization correction on the initial design evaluation report based on the result, and outputting a final design evaluation report integrating spatial layout and visual attribute optimization suggestions. According to the method, the interference effect is accurately quantified through conjoint analysis of spatial layout and semantic attributes, and the accuracy and reliability of design conveying efficiency evaluation are improved.
Owner:CHANGCHUN ARCHITECTURE & CIVILENGEERING CO LLEGE

Methods and systems of tag location detection in an inventory environment based on visual attributes of tags using computer vision

A method comprises registering, by an application, a tag identifier received from a tag in an inventory environment with location data indicating a location of the tag based on a visual attribute of the tag, initiating, by the application, a scan of the tag to obtain tag data from the tag and to capture an image depicting the tag by transmitting a signal to the tag after registering the tag identifier with the location data of the tag, and triggering, by the application, activation of a light emitting diode (LED) on the tag to indicate whether a reader device is in a read range of the tag, in which a visual feature of the LED indicates whether the reader device is in the read range of the tag, and wherein the signal is used to activate the LED on the tag.
Owner:T MOBILE INNOVATIONS LLC

Multi-label field adaptive method based on prompt driving

The invention discloses a multi-label field adaptive method based on prompt driving, which belongs to the technical field of computer vision and comprises the following steps: generating multi-dimensional semantic text description containing visual attributes, semantic levels and category names of common scenes for each type; embedding the semantic text description into an optimizable category vector, and embedding the optimized category vector into a CLIP text prompt; extracting and projecting multi-layer style statistical characteristics of the image, and synchronously realizing image-text cross-modal alignment and source-target domain distribution alignment in a frame through a cross-domain style mapping network and various alignment losses; semantic priori knowledge of CLIP is combined with a label co-occurrence mode of a source domain, so that the model can sense and adapt to a possibly changed label dependency relationship in a target domain; performing semantic propagation on label embedding by using a graph convolutional network; the whole system is jointly optimized through a multi-task loss function, and end-to-end cross-domain multi-label classification is achieved.
Owner:NANJING UNIV OF POSTS & TELECOMM

Techniques for training identity transfer models

Systems, devices, and methods are provided for training and / or inferencing using machine-learning models. In at least one embodiment, an identity embedding and a content embedding are extracted from training media data, such as training audio data and / or training visual data. The identity embedding may encode information relating to a person's voice and / or visual properties. A decoder may be trained to create synthetic media, such as audio and visual content, based on identity and content embeddings.
Owner:AMAZON TECH INC

Text condition clothing image generation method and system based on diffusion model

The invention provides a text condition clothing image generation method and system based on a diffusion model, and the method comprises the steps: firstly obtaining a natural language input sequence containing clothing style types, fabric texture and decoration detail description elements, and carrying out the hierarchical semantic modeling processing of the natural language input sequence; generating a composite text guide vector containing the basic attribute features and the association relationship features, then generating an initial image tensor which has the same spatial size as the target clothing image and has random disturbance distribution, and taking the composite text guide vector as a conditional constraint; and performing multi-stage iterative denoising processing on the initial image tensor by using a pre-trained diffusion generation model to generate a continuous image sequence containing an intermediate generation state, and finally extracting an image in a final evolution state from the sequence as a target clothing image. Visual attributes and description elements of a natural language input sequence form a structured corresponding relation, and a garment image conforming to complex natural language description can be accurately generated.
Owner:SHANGHAI JIUZHIRUN INFORMATION TECH CO LTD

Vehicle-mounted three-dimensional scene rendering method and device, storage medium and program product

The invention provides a vehicle-mounted three-dimensional scene rendering method and device, a storage medium and a program product, and the method comprises the steps: obtaining the perception data of a target object, and constructing a nonlinear motion state space representing the motion process of the target object; performing prediction updating on the target object based on the state transition model of the nonlinear motion state space to obtain prediction state information and a prediction covariance matrix at the current rendering moment; extracting variance information related to the position from the covariance matrix to calculate an uncertainty factor, introducing the uncertainty factor into a graphic rendering pipeline, and dynamically adjusting at least one type of visual attribute of the three-dimensional model to realize uncertainty prompt; and when new perception data of the target object is obtained, performing fusion updating on the prediction state information, and performing smooth transition on a rendering result based on the correction amount to suppress rendering jump. According to the method, high-frame-rate continuous rendering display is realized under the condition of low-frequency updating of the sensor, and the interpretability and the safety of a rendering result are improved.
Owner:SHANGHAI GEOMETRICAL PERCEPTION & LEARNING CO LTD

Vehicle image-text retrieval method and system based on multistage semantic graph alignment and attribute enhancement

The invention belongs to the technical field of intelligent traffic, and relates to a vehicle image-text retrieval method and system based on multistage semantic graph alignment and attribute enhancement. According to the method, fusion visual embedding is obtained according to the vehicle video, and fusion text embedding is obtained according to the text data; obtaining an updated visual attribute embedding vector and an updated text attribute embedding vector based on fusion visual embedding and fusion text embedding; performing semantic matching on the updated visual attribute embedding vector and the updated text attribute embedding vector, and aligning vehicle cross-modal semantic attributes from three different levels by utilizing multi-granularity semantics to obtain final visual attribute embedding and final text attribute embedding; and according to the final visual attribute embedding and the final text attribute embedding, obtaining the similarity between the vehicle video and the artificial text description, wherein the vehicle video with the highest similarity is a vehicle video retrieval result. According to the method, the target vehicle can be quickly and accurately positioned, and higher robustness and higher retrieval success rate are shown.
Owner:CHANGAN UNIV

User interface for interacting with an affordance in an environment

Various implementations disclosed herein include devices, systems, and methods for indicating a distance to a selectable portion of a virtual surface. In various implementations, a device includes a display, a non-transitory memory and one or more processors coupled with the display and the non-transitory memory. In some implementations, a method includes displaying a graphical environment that includes a virtual surface, wherein at least a portion of the virtual surface is selectable. In some implementations, the method includes determining a distance between a collider object and the selectable portion of the virtual surface. In some implementations, the method includes displaying a depth indicator in association with the collider object. In some implementations, a visual property of the depth indicator is selected based on the distance between the collider object and the selectable portion of the virtual surface.
Owner:APPLE INC

Text-driven digital human audio and video generation method

The invention discloses a text-driven digital human audio and video generation method, relates to the technical field of intelligent virtual digital human generation, and provides the following scheme: obtaining a driving text and a reference picture, preprocessing the driving text to generate a phoneme sequence containing pronunciation time sequence information, and sending the phoneme sequence to the reference picture; and extracting key face feature points in the reference picture through a face feature point positioning technology, extracting tone reference feature vectors associated with tones in the reference picture through a deep learning model, and inputting the phoneme sequence into a text-to-speech model to generate an original audio. According to the method, the timbre feature vectors and the face feature points are synchronously extracted from the single reference picture, and a multi-modal feature cross fusion technology is combined, so that the problem of timbre and visual attribute splitting in a traditional method is solved, the unified digital human identity of personalized timbre and visual image is driven by only one picture, extra sound recording is not needed, and the user experience is improved. The use threshold is reduced.
Owner:HANGZHOU DIANZI UNIV

Fragment shader for creator signatures in a virtual asset marketplace

Various implementations relate to methods, systems, and computer-readable media to dynamically apply creator signatures to virtual assets within a virtual platform. According to one aspect, a computer-implemented method includes receiving a request to display a virtual asset associated with a creator, where the request is linked to an avatar in a virtual environment hosted on the platform. A creator signature, which is defined by visual properties unique to the creator, is retrieved and stored separately from the virtual asset. The method includes rendering the virtual asset by rasterizing it and applying a fragment shader to overlay the creator signature as a dynamic visual element that changes over time based on predefined parameters. The rendered virtual asset with the dynamic signature is displayed within the virtual environment, providing visual attribution to the creator while preventing unauthorized embedding of the signature in the asset data.
Owner:ROBLOX CORP

Video generation method and device, storage medium and computer program product

The embodiment of the invention discloses a video generation method. The method comprises the steps of obtaining a to-be-processed video and visual attribute information of a user; the visual attribute information represents the focusing capability of eyes of the user on light; on the basis of the visual attribute information, determining parallax information of pixels of each frame of image in the to-be-processed video; determining a plurality of target images based on the parallax information and each frame of image; the target three-dimensional video is determined based on the plurality of target images, so that the problem that the generated 3D video is not matched with the visual characteristics of the user when the 3D video is generated in the related technology is solved, and the watching experience of the user is improved. The embodiment of the invention further discloses video generation equipment, a storage medium and a computer program product.
Owner:MIGU CO LTD +1

Leakage detection model construction method, leakage detection method and device

The invention provides a leakage detection model construction method and a leakage detection method and device, and the method comprises the steps: carrying out the preprocessing of continuous multi-frame original images of to-be-detected equipment, and obtaining a preprocessed image; performing feature extraction operation on the preprocessed image to obtain a first feature map of the preprocessed image; multiple different types of visual attribute values in the preprocessed image are determined, and the sizes of the visual attribute values are in positive correlation with the leakage possibility of the to-be-detected equipment; according to the visual attribute values, target visual features corresponding to the visual attribute values are extracted from the first feature map, a second feature map is generated based on the target visual features, and the larger the visual attribute values are, the larger the contribution degree of the corresponding target visual features to the second feature map is; based on the second feature map, determining the leakage probability of the to-be-detected equipment; and constructing a leakage detection model based on the leakage probability. The method can reduce the omission ratio of tiny leakage.
Owner:YILIAN CLOUD COMPUTING (HANGZHOU) CO LTD +1

Combined zero sample recognition method based on shared learnable soft query vector and related equipment

The invention discloses a combined zero sample recognition method based on a shared learnable soft query vector and related equipment. The method comprises the following steps: acquiring an initial visual feature, an initial text attribute feature, an initial text object feature and an initial text combination feature; performing visual alignment and decoupling by sharing the soft query vector to obtain a target visual attribute feature, a target visual object feature and a target visual combination feature; performing text alignment and decoupling by sharing the soft query vector to obtain a target text attribute feature, a target text object feature and a target text combination feature; obtaining the similarity between the target visual attribute feature and the target text attribute feature, the similarity between the target visual object feature and the target text object feature, and the similarity between the target visual combination feature and the target text combination feature; and carrying out weighted fusion on the similarity to obtain a combined zero sample recognition result. The method can improve the recognition capability of unseen combinations, and can be widely applied to the technical field of computer vision.
Owner:GUANGZHOU UNIVERSITY

Product image generation based on diffusion model

Methods and systems are disclosed for generating an extended reality (XR) try-on experience based on an image produced by a diffusion model. The system receives an image depicting a real-world object and generates a prompt comprising a textual description of a fashion item. The system analyzes the image and the textual description of the fashion item using a generative machine learning model to generate an artificial image that depicts an artificial object that resembles the real-world object wearing an artificial fashion item matching the textual description of the fashion item. The system identifies an object comprising a real-world product image that matches visual attributes of the artificial fashion item and replaces the artificial fashion item in the artificial image with the object to generate an output image.
Owner:SNAP INC

Leakage detection model construction method, leakage detection method and device

This application provides a method for constructing a leakage detection model, a leakage detection method, and an apparatus, comprising: preprocessing multiple consecutive frames of original images of a device to be detected to obtain a preprocessed image; performing feature extraction on the preprocessed image to obtain a first feature map of the preprocessed image; determining multiple different types of visual attribute values ​​in the preprocessed image, wherein the magnitude of the visual attribute values ​​is positively correlated with the probability of leakage in the device to be detected; extracting target visual features corresponding to each visual attribute value from the first feature map according to the magnitude of each visual attribute value, and generating a second feature map based on the target visual features, wherein the larger the visual attribute value, the greater the contribution of the corresponding target visual feature to the second feature map; determining the leakage probability of the device to be detected based on the second feature map; and constructing a leakage detection model based on the leakage probability. This method can reduce the false negative rate of minor leaks.
Owner:YILIAN CLOUD COMPUTING (HANGZHOU) CO LTD +1

A method for cluttered scene object grasping based on visual-linguistic-action joint modeling

The application discloses a method for cluttered scene target object grasping based on visual-language-action joint modeling. The application uses object-centered representation to realize a method for cluttered scene target object grasping based on visual-language-action joint modeling, processes the object-centered representation through a pre-trained visual-language model and a grasping model, obtains visual-language features and grasping features of each bounding box, and uses a transformer to implement cross-attention mechanisms among visual-language-action multimodality, generates visual-language-action cross-attention features, and then generates decisions and executes, so that higher sample utilization is realized, and additional data collection and training in the simulation-physical migration process are avoided; compared with a two-stage strategy, visual attributes and planner screening rules for language-visual matching need not be artificially designed, so that more flexible language instructions can be adapted, and better task generalization is achieved.
Owner:ZHEJIANG UNIV

Audio frequency spectrum fluidization interaction presentation method and system based on shader computing power

The invention discloses an audio frequency spectrum fluidization interaction presentation method and system based on shader computing power, and belongs to the technical field of image processing, and the method comprises the steps: obtaining an audio stream and user interaction data extraction features in real time; generating a sound interaction total dynamic field in the GPU parallel computing shader through nonlinear coupling; inputting the field into a preset neural network model in a fragment shader for forward reasoning to complete fluid neural physical evolution; separating an audio semantic layer in the rendering pipeline and executing competitive visual attribute mapping to generate an image; and performing asynchronous scheduling on each processing pipeline through an asynchronous calculation engine. According to the method, full-pipeline GPU calculation and neural operator evolution are adopted, the contradiction between fluid simulation and real-time rendering can be effectively solved, frequent data transmission is abandoned, low-delay interaction feedback is achieved while the physical reality sense is guaranteed, and the immersion sense and expressive force of audio visualization are remarkably improved.
Owner:CHENGDU LIBI TECH CO LTD

Multi-level intelligent agent construction method for regional integrated energy system

The invention belongs to the technical field of energy intelligence, and provides a regional integrated energy system multi-level agent construction method, which comprises the following steps: carrying out space-time alignment processing on real-time operation data to obtain a synchronous data stream; respectively mapping parameters in the synchronous data stream into specific visual attribute primitives to obtain a multi-level state semantic image; performing visual semantic segmentation on the multi-level state semantic image to obtain visual semantic units, and constructing a spatial correlation map between the visual semantic units; performing visual difference comparison on the spatial correlation map and a historical same-period spatial correlation map to obtain an abnormal visual semantic unit; performing multi-frame tracking verification on the abnormal visual semantic unit, and integrating an abnormal unit identifier and an abnormal type into an abnormal event report; and carrying out hierarchical decision making on the abnormal event report and the spatial correlation graph to obtain a collaborative regulation and control instruction. According to the invention, the efficiency of cooperative management of the regional integrated energy system can be improved.
Owner:STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO +1

Evaluating bias in generative models

In implementations of systems for evaluating bias in generative models, a computing device implements a bias system to generate a modified digital image by processing an input digital image using a first machine learning model trained on training data to generate modified digital images based on input digital images. The bias system computes a first latent representation of the input digital image and a second latent representation of the modified digital image using a second machine learning model trained on training data to compute latent representations of digital images. A bias score is determined for a visual attribute based on the first latent representation and the second latent representation. The bias system generates an indication of the bias score for the visual attribute for display in a user interface.
Owner:ADOBE INC

3D model generation using multiple textures

Methods and systems for generating 3D assets for use in, for example, extended reality (XR) experiences are disclosed. A system receives a plurality of textures associated with an object, each of the plurality of textures corresponding to a different view of the object, and automatically generates an initial three-dimensional (3D) model of the object based on an initial alignment of the plurality of textures to respective portions of the initial 3D model. The system receives input adjusting the initial alignment of the plurality of textures to the respective portions of the 3D model, and combines the plurality of textures into a single texture based on the input, the single texture defining visual properties of the object from the plurality of views. The system stores the 3D model in association with the single texture.
Owner:SNAP INC

Intelligent visual adaptation for image and / or video applications to enhance user experience

A computer-implemented method including receiving an input image; adapting, based on one or more visual adaptation criteria, one or more visual properties of the input image to generate a synthesized image, wherein the adapting includes generating one or more masked image patches, each comprising a respective masked portion of the input image; processing the one or more masked image patches using one or more computer vision models to generate respective ones of one or more output image patches; and generating the synthesized image based on the input image and the one or more output image patches, the synthesized image comprising at least a first portion extending a view of the input image and a second portion modifying content within the input image; and rendering, on a display device, the synthesized image.
Owner:FUTUREWEI TECHNOLOGIES INC

Self-adaptive video watermarking method based on Web front end

The invention discloses a self-adaptive video watermarking method based on a Web front end. Establishing a standardized percentage coordinate system based on the video effective area; the method comprises the following core steps: acquiring a preview canvas size and a device pixel ratio (DPR), and executing rendering precision compensation calculation to generate a vectorization logic component; sensing transparency in real time, and triggering self-adaptive visualization enhancement assistance based on background brightness contrast when a threshold value is met; and capturing an interaction instruction, updating a proportional parameter in real time based on bidirectional mapping, and executing boundary clamping correction. Atomized attribute stripping is performed in response to the save instruction to restore the original visual attributes and generate a structured configuration information stream. According to the method, engineering pain points such as heterogeneous resolution geometric distortion, high-transparency watermark interaction difficulty and cross-end display fuzziness are solved through a visual means of DPR compensation and threshold perception, and watermark positioning precision, editing interaction efficiency and cross-platform visual consistency are remarkably improved.
Owner:HANGZHOU ARTECH

Dynamic relationship-based avatar garment generation

PCT designated stageWO2026147888A1PersonalizationComputer graphics (images)
The described system facilitates personalized avatar interactions by dynamically generating and displaying modified garments. The system determines the initiation of an interaction function by a first user with a second user within an interaction platform. It accesses avatar data for the first user, including visual attributes and a garment associated with the first avatar, as well as avatar data for the second user, including visual attributes and a garment associated with the second avatar. An image is generated featuring the first avatar wearing its garment alongside the second avatar wearing its garment. This image is applied to the first avatar's garment, creating a modified version that visually incorporates the second avatar. The system then displays the first avatar wearing the modified garment alongside the second avatar wearing its original garment, enabling enhanced user engagement through visually personalized and contextual representations of user interactions.
Owner:SNAP INC

Curved screen display method and device

The invention relates to the technical field of curved screen fault diagnosis and compensation, and discloses a curved screen display method and device. The method comprises the following steps: capturing an original display signal flow containing a pixel signal, a driving time sequence and an embedded sensor signal, and generating a three-dimensional spatio-temporal data volume after synchronous alignment and standardized packaging; a dynamic anomaly perception map is constructed based on the data body, nodes of the map are composed of visual attributes and logic state units, and edges of the map are defined by physical quantity change relations. Next, nodes in the map are associated and mapped to a curved screen physical model, abnormal association strength is calculated, and a screen physical area and a signal driving channel corresponding to an abnormal source node are backtracked and positioned according to the abnormal association strength; and generating a targeted pixel signal compensation sequence and a driving parameter compensation sequence according to a positioning result. According to the method, accurate diagnosis and directional compensation of display abnormity from a phenomenon to a physical source and a circuit channel are realized.
Owner:BEIJING LINGBAN WORKSHOP TECHNOLOGY CO LTD

Image processing method and device

The embodiment of the invention discloses an image processing method and device. According to the image processing method, an initial image needing to be processed, a processing area of the initial image and target characters needing to be presented in the processing area are obtained, visual attributes of characters in the area except the processing area are determined, and the initial image is subjected to diffusion processing in the processing area to generate an area image; the content of the area image is the target character with the visual attribute of the character of the area outside the processing area, and the target character generated through diffusion processing covers the original character of the processing area and has the same style attribute as the character which is not modified in the initial image, so that the modification of the existing character in the initial image is realized; and the overall styles of the modified characters and the characters in the initial image are unified, so that the modification trace between the modified characters and the non-modified parts is eliminated, and the viewing feeling of a reader on the modified new image is improved.
Owner:ZHUHAI KINGSOFT OFFICE SOFTWARE +2

Virtual resource processing method and device, storage medium, equipment and program product

The invention discloses a virtual resource processing method and device, a storage medium, equipment and a program product. The method comprises the following steps: acquiring physiological data of a user; the physiological data are preprocessed to obtain feature change data, and the feature change data can represent physiological changes of the user; material control parameters and / or physical control parameters of the virtual resources are / is calculated based on the feature change data; and adjusting visual attributes of the virtual resources according to the material control parameters and / or adjusting physical attributes of the virtual resources according to the physical control parameters. According to the method, the physiological data can be effectively converted into real-time parameters in the game through data analysis and quantification, reliable technical support is provided for development of physiological feedback games, and therefore better game experience can be provided for users.
Owner:NETEASE (HANGZHOU) NETWORK CO LTD