Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

5296results about "Editing/combining figures or text" patented technology

Text-driven CAD modeling method and system based on diffusion and visual language model

The invention relates to the technical field of computer aided design, in particular to a text-driven CAD modeling method and system based on a diffusion and visual language model.The method comprises the steps that natural language text description is obtained, and CAD semantic features of the natural language text description are extracted; carrying out geometric standardization on the CAD semantic features by adopting a fine-tuning diffusion model, and generating a CAD view image conforming to engineering specifications; carrying out fusion by adopting a fine-tuned visual language model to generate a parameterized CAD construction sequence; a three-mode alignment mechanism is adopted, and the semantic consistency of the CAD semantic features, the CAD view images and the CAD construction sequences is checked; performing verification and post-processing on the CAD construction sequence, and outputting an executable Python code or STEP file; the CAD modeling method disclosed by the invention performs explicit modeling based on flexible modal description, and has the characteristics of high geometric constraint and high usability.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

System and method of protecting facial privacy using text-guided makeup via adversarial latent search

Disclosed are a method and system to protect user facial privacy against unknown face recognition levels without compromising on a user's online experience. An input source to input an original face image. A training circuit configured to train a generator model to output an image that resembles the original face image. An optimizer configured to generate a protected face image based on the trained model that fools a black-box face recognition model, while imitating a makeup style. A display device to display the protected face image online.
Owner:MOHAMED BIN ZAYED UNIV OF ARTIFICIAL INTELLIGENCE

Image generation method and device based on theme information, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business scenes such as financial science and technology and medical health, and discloses an image generation method, device and equipment based on theme information and a medium. The method comprises the steps that input information is analyzed to generate theme information and copywriting information, a cue word set is generated, and a figure image set and a background image set are generated; fitting the segmented figure image with the background image to form a head image candidate set, and selecting a head image matching template frame to generate a basic image; decomposing the copywriting to generate a sub-module initial picture set, and adding a gradual change effect to form a sub-module picture set; the adjusted sub-module pictures are obtained based on size adjustment, and a pre-synthesized image is generated through splicing; and identifying the blank area to draw a title text to obtain a final image. Through theme analysis, template matching, image splicing, modular copywriting processing and blank drawing, poster generation efficiency is improved, layout flexibility is enhanced, and visual unification is realized.
Owner:CHINA PING AN PROPERTY INSURANCE CO LTD

Ai-driven creation of custom stickers from messages in chat interfaces

PendingUS20250378602A1Mathematical modelsNatural language analysisEngineeringVisual expression
This disclosure relates to techniques for generating and utilizing custom stickers in a digital communication environment. A technique involves receiving a text-based message input during a chat session and using a generative language model (e.g., a Large Language Model, or LLM) to create a text prompt. This prompt is then used by a generative image model to produce a custom sticker. The generated sticker is sent to a client device where it is displayed in a sticker tray alongside other selectable stickers. Users can select and send these stickers directly within their chat interface, enriching communication with visually expressive and contextually relevant imagery.
Owner:SNAP INC

Creation content generation system based on image recognition and large language model fusion

The invention discloses a creation content generation system based on image recognition and large language model fusion, and particularly relates to the technical field of creation content generation, and the system firstly completes the fact extraction and brand anchor point construction of an input image in a unified coordinate and scale system, and forms a structured fact package in one-to-one correspondence with an original image; then, performing protagonality scoring and ambiguity gating on the figure instance, and outputting an explainable and calibratable protagonality judgment result; on this basis, the condition controlled generation and template selection module converts the fact constraint into a controlled text packet and a format instruction packet, and keeps explicit mapping with a fact packet; the system further executes cross-modal consistency and compliance verification based on image facts, and machine-readable verification and minimum cost correction are carried out on text and graph entities, geometrical relationships and brand elements; and finally, solidifying the key intermediate quantity, the parameters and the judgment basis into an evidence chain through a chain type index, and introducing online adaptive learning in a compliance boundary to realize mild updating and rollback release.
Owner:HANGZHOU SHUANGHEDAN NETWORK TECH CO LTD

Multimodal sentiment analysis method based on diffusion model and self-paced learning

The invention provides a multi-modal sentiment analysis method based on a diffusion model and self-paced learning. The method comprises the following steps: firstly, dividing a data set into a missing image modal data set and a complete modal data set according to image modal integrity; thirdly, constructing a feature alignment diffusion model, and performing image generation; training the diffusion model by adopting a self-paced learning strategy and a missing image data set; and based on the trained diffusion model, guiding a reverse process through text features to generate feature representation of the missing image. And carrying out weighted fusion on the generated image features and text features by using an attention mechanism, and dynamically adjusting contribution weights of all modalities to generate a complete multi-modal feature representation. And finally, integrating a missing modal completion result and the complete modal features to form a unified multi-modal representation, inputting the unified multi-modal representation into a multi-modal sentiment classification module, and outputting a sentiment classification result. According to the method, the problem of multi-modal sentiment analysis under random missing of image modals is effectively solved, and the generation quality and semantic consistency are improved.
Owner:GUILIN UNIV OF ELECTRONIC TECH

Using artificial intelligence to generate images of product based on user input

The disclosed technology includes a computer-implemented technique for generating images of physically producible products in response to user input describing a conceptual product. The system receives user input—such as natural language text, speech, or images—via a user interface, configures a prompt for a generative artificial intelligence (AI) system, and generates an image representing a version of the conceptual product with distinct physical attributes. Each version is associated with a unique identifier, enabling selection, modification, and purchase of the conceptual product. The system supports multiple product categories, including jewelry, home décor, and fashion, and extracts physical attributes to determine manufacturability and pricing. User feedback is incorporated to improve AI performance. The invention enables presentation of selectable product versions and initiates manufacturing processes based on user selections, supporting unstructured user input and multiple data modalities.
Owner:ARCADE STUDIO INC

General nerve drawing method and system based on illumination function generation model

The invention discloses a general nerve drawing method and system based on an illumination function generation model, and belongs to the technical field of computer graphics, and the method comprises the steps: collecting light source information containing multi-view observation data, and collecting scene information; constructing an illumination function generation model comprising a light source coding module and a light source decoding module for converting the light source information into neural illumination representation and performing joint inference based on the neural illumination representation and scene features of the drawing points; and generating a final drawn image conforming to the real illumination distribution based on the inference result. According to the method, generalization neural drawing across light sources and scenes can be achieved, images with off-line rendering quality can be generated at the cost close to real-time calculation, the sense of reality, stability and rendering efficiency are considered, and the method has good expansibility and wide application value.
Owner:ZHEJIANG UNIV

Image inversion and editing using rectified flow neural networks

Systems and methods for performing image modification. In particular, the system can, using a rectified flow neural network, perform an image inversion and image editing process to generate a modified image that has been modified according to a conditioning input received by the system.
Owner:GOOGLE LLC

Systems and methods for image generation with machine learning models

Disclosed herein are methods, systems, and computer-readable media for regenerating a region of an image with a machine learning model based on a text input. Disclosed embodiments involve accessing a digital input image. Disclosed embodiments involve generating a masked image by removing a masked region from the input image. Disclosed embodiments involve accessing a text input corresponding to an image enhancement prompt. Disclosed embodiments include providing at least one of the input image, the masked region, or the text input to a machine learning model configured to generate an enhanced image. Disclosed embodiments involve generating, with the machine learning model, the enhanced image based on at least one of the input image, the masked region, or the text input.
Owner:OPENAI OPCO LLC

Solar irradiance prediction system and method based on dual attention multi-mode fusion

The invention discloses a solar irradiance prediction system and method based on dual attention multi-modal fusion. The method comprises the following steps: performing spatial feature extraction on a sky imaging observation map after preprocessing and time synchronization by using a MobileNetV2 network to obtain a sky imaging time sequence feature sequence; inputting the sky imaging time sequence feature sequence into a DAT model to obtain a global sky imaging time sequence feature sequence, a local sky imaging time sequence feature sequence and a local convolution position coding residual term; obtaining an aggregated sky imaging time sequence feature sequence according to the global sky imaging time sequence feature sequence, the local sky imaging time sequence feature sequence and the local convolution position coding residual item; performing time-dependent modeling on the meteorological observation data by adopting a long short-term memory network to obtain a meteorological observation data feature sequence; and performing feature fusion on the aggregated sky imaging time sequence feature sequence and the meteorological observation data feature sequence through a trans-attention mechanism to obtain a fused joint feature sequence for solar irradiance prediction.
Owner:WUHAN UNIV OF TECH

Ultrasonic operation intelligent training method and equipment based on multiple modes

The invention relates to the technical field of medical simulation training, in particular to an ultrasonic operation intelligent training method and device based on multiple modalities, and the method comprises the steps: obtaining real-time six-degree-of-freedom pose data of an ultrasonic probe held by an operator, and generating an ultrasonic image in real time through a first deep learning model in cooperation with scene parameters; wherein the first deep learning model is trained to learn and establish a continuous mapping relation from an ultrasonic probe pose space to an ultrasonic image space; performing section classification on the ultrasound image by using a second deep learning model, and determining deviation information for a non-standard section; and based on the real-time six-degree-of-freedom pose data and a division result of a preset standard section and an ultrasonic probe pose, generating visual guide information for guiding an operator to adjust the probe pose, and displaying the visual guide information. In this way, the problems that an existing virtual training image is discontinuous and guiding is inaccurate are solved, and meanwhile the training efficiency and the reality sense of ultrasonic operation are remarkably improved.
Owner:BEIJING ANZHEN HOSPITAL AFFILIATED TO CAPITAL MEDICAL UNIV +1

Feedback Predictions for Machine-Learned Generative Models

Aspects of the disclosed technology include computer-implemented systems and methods for machine-learned multimodal models for feedback predictions for synthetic content. A machine-learned multimodal model is configured to generate a feature map based at least in part on fusion of image information and text information from a synthetic image and a text prompt. The model is configured to generate a set of text tokens based at least in part on fusion of the image information and the text information. The model is configured to generate at least one misalignment or implausibility heatmap based at least in part on the at least one feature map. The model is configured to generate at least one predicted misalignment sequence based at least in part on the set of text tokens.
Owner:GOOGLE LLC

Condition-based image editing

A computer system and a computer-implement method include obtaining a source image and a modification input that indicates a target edit to the source image and generating a modification encoding representing the target edit. An image generation model generates an output image that depicts the source image with the target edit based on the source image and the modification encoding. The image generation model is trained to perform a pose modification task and a part replacement task.
Owner:ADOBE INC

Dynamic generation method of digital animation character based on generative model

The invention discloses a digital animation role dynamic generation method based on a generative model. The method comprises the following steps: inputting a role original image and scene background information, extracting skeleton key points and scene features, analyzing an action sequence, predicting a motion track, calculating an optimal position of a role in a picture, and carrying out dynamic adjustment. And according to the role position, obtaining morphological feature data, evaluating the quality grade, and optimizing the coordination of the role image. And finally, comprehensively scoring by adopting a multi-dimensional quality evaluation system, and determining the visual presentation quality of the role. According to the method, deep fusion of role actions, positions and scenes is realized, the dynamic expressive force and visual coordination of animation pictures are improved, and an efficient solution is provided for generating high-quality animation contents.
Owner:HEBEI XIONGAN PEPSI HENGXING NETWORK TECHNOLOGY CO LTD

Retrieval augmented text-to-image generation

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for generating an output image using a text-to-image model and conditioned on both the input text and image and text pairs selected from a multi-modal knowledge base. In one aspect, a method includes, at each of multiple time steps: generating a first feature map for the time step; selecting one or more neighbor image and text pairs based on their similarities to the input text; for each of the one or more neighbor images and text pairs, generating a second feature map for the neighbor image and text pair; applying an attention mechanism over the one or more second feature maps to generate an attended feature map; and generating an updated intermediate representation of the output image for the time step.
Owner:GOOGLE LLC

Face image reconstruction method based on semantic identity feature decoupling and consistency retention of diffusion model

The invention discloses a face image reconstruction method based on semantic identity feature decoupling and consistency reservation of a diffusion model, and the method comprises the steps: 1, obtaining and preprocessing a face image set of identity labeling, and generating a face feature point distribution diagram, a semantic mask diagram and a description text; 2, multi-modal features are extracted and fused through a semantic identity extraction network; 3, carrying out noise adding and de-noising processing by utilizing a diffusion model, and combining a reconstructed network and semantic identity loss optimization; and 4, face image reconstruction is completed. According to the method, in the face image reconstruction process, the driving requirements of semantic information such as texts for image editing can be accurately captured, fine-grained semantic features and identity features are decoupled, the core identity features of the face can be effectively reserved, and loss of identity consistency caused by semantic editing is avoided; therefore, technical support is provided for application scenes with high requirements on face identity accuracy in the field of computer vision, and the reliability and practicability of face image reconstruction are improved.
Owner:ANHUI UNIV

Artificial Intelligence (AI) agent inputs using User Interfaces (UIs)

Systems and methods for Artificial Intelligence (AI) agent inputs using User Interfaces (UIs) includes operating an Artificial Intelligence (AI) agent system that includes an agent core connected to memory, one or more tools, and a planner; receiving an input from a user, wherein the input includes any of a prompt from the user and a selection from a User Interface (UI); and generating, via the AI agent, an answer based on the input.
Owner:ZSCALER INC

Photovoltaic panel defect inspection method and system based on hydrogen energy unmanned aerial vehicle

The invention discloses a photovoltaic panel defect inspection method and system based on hydrogen energy unmanned aerial vehicles, and the method comprises the steps: scheduling two hydrogen energy unmanned aerial vehicles to work synchronously, and collecting visible light and infrared thermal image sequences of the front and back surfaces of a double-sided photovoltaic panel; analyzing a photovoltaic array CAD / BIM drawing, extracting geometric coordinates, topological structures and numbers of components, and establishing a local coordinate system; edge contours are extracted from front and back visible light images in a differentiated mode and matched with drawing coordinates to construct pixel-physical coordinate mapping; an abnormal area is positioned through the temperature gradient of the infrared thermogram, and an assembly number is bound; a thermoelectric coupling decoupling model is constructed, real defect data is obtained through heat conduction compensation and current balance correction, and a defect type is identified in combination with a double-current convolutional network; and superposing the front and back defect data to a view layer corresponding to the drawing, generating an independently labeled defect distribution diagram, and outputting the defect distribution diagram to a management platform. The problem of misjudgment caused by thermoelectric interference during inspection of the double-sided photovoltaic panel is solved, and the defect positioning and recognition precision is improved.
Owner:BEIJING YUANSHEN ENERGY SAVING TECH +1

System and method for splitting an image across a plurality of tiles

A system and method for editing and outputting an image, for example, for wall or other décor. Tools enable a user to parse a single image, such as a photograph, substantially automatically across a grid of multiple tiles. In addition to parsing, various kinds of image editing are provided as a function of tools in an innovative graphical user interface operating on a smart phone or one or more computing devices. Various operations are performed on an image that has been parsed across a grid of multiple tiles to provide for a custom output, such as for displaying an image uniquely on a wall or other surface.
Owner:TRACER IMAGING LLC

Generative model experience using open prompt

Described is a system for a generative model XR Experience using open prompt by receiving a first prompt of a first user via a user interface of a user device indicating a user's intent, processing the first prompt using a first machine learning model to generate a second prompt that is applied to a second machine learning model, the second prompt indicative of attributes associated with the first prompt, capturing an image of the first user via a camera feed of the user device, processing a combination of the image of the first user with the second prompt using the second machine learning model to generate a plurality of images, and applying the plurality of images to the live camera feed of the user device.
Owner:SNAP INC

Texture surface defect generation method and system based on mask perception image redrawing network

The invention belongs to the related technical field of image processing, and discloses a texture surface defect generation method and system based on a mask perception image redrawing network, and the method comprises the steps: (1) obtaining a low-frequency structure feature based on a defect sample image of a to-be-processed product category, a binary mask for marking a defect position, and a convolutional coding module; (2) a context sliding window attention modeling module carries out long-range dependence modeling on the feature sequence, and attention calculation is carried out between effective feature marks to obtain high-frequency detail features; (3) inputting the low-frequency structural features and the high-frequency detail features into a collaborative feature fusion module for fusion to obtain fusion features; (4) a style modulation decoding module performs up-sampling on the fusion features based on the comprehensive style tensor; and (5) generating a defect image based on the trained generator network model, the defect-free sample image of the to-be-processed product category and a user-defined defect mask. According to the invention, the authenticity of defect generation is improved.
Owner:HUAZHONG UNIV OF SCI & TECH

Methods and systems for preserving image features during image editing

Described embodiments generally relate to a computer-implemented method for editing an image. The method includes accessing an image; identifying at least a first area of the image and a second area of the image; configuring a model to generate an edited image based on the first area of the image and the second area of the image, wherein the edited image comprises a first area of the edited image and a second area of the edited image; wherein the model is configured to generate the edited image such that the first area of the edited image differs from the first area of the image less than the second area of the edited image differs from the second area of the image.
Owner:CANVA PTY LTD

Artificial intelligence semantic processing system and method for digital media creation

The invention provides an artificial intelligence semantic processing system and method oriented to digital media creation, and relates to the technical field of artificial intelligence semantic process.The artificial intelligence semantic processing method comprises the steps that predicate argument relation pairs of language texts are extracted, object space relation pairs of sketch images are extracted at the same time, and a basic semantic unit set is constructed; the integrity and accuracy of cross-modal semantic understanding are ensured, further, semantic units are clustered by using a dynamic routing algorithm, a semantic concept cluster with a clear importance weight is generated, deep mining and structured representation of creation intentions are realized, and the creation intentions are quickly and accurately understood. An initial semantic relation graph is constructed, a graph attention network is used for dynamic reweighting, finally, an enhanced dynamic semantic graph is generated, complex association and a hierarchical structure between semantic concepts are effectively captured, finally, hierarchical analysis is carried out on the semantic graph, and a structured semantic blueprint is output, so that the dynamic semantic graph is obtained. And a reliable semantic processing technology is provided for creation of high-quality digital media contents.
Owner:HUNAN INST OF INFORMATION TECH

Stylized visual text editing method, system and equipment and storage medium

The invention discloses a stylized visual text editing method, a stylized visual text editing system, stylized visual text editing equipment and a storage medium, which are corresponding schemes, and the related schemes aim to solve the problem of style consistency existing in image text editing of an existing diffusion model, and the stylized visual text editing efficiency is improved by combining visual features of a font image and an input text image. The method comprises the following steps of: extracting style embedded information from a text image, and inputting the style embedded information as an enhanced style condition into a diffusion model to realize fine control on a diffusion process, so that the diffusion model can generate a text image with high readability and style consistency, and can realize maintenance of an original text style or style migration based on a reference image.
Owner:UNIV OF SCI & TECH OF CHINA

Text rendering for image generation models

A method, apparatus, non-transitory computer readable medium, and system for image processing include obtaining an image generation prompt comprising a text to be generated in a synthetic image, generating a first image feature based on the image generation prompt, where the first image feature represents the text, and generating a synthetic image based on the image generation prompt and the first image feature, where the synthetic image includes the text.
Owner:ADOBE INC

Methods, apparatuses and computer program products for providing tuning-free personalized image generation

A system and method to generate a target image from a reference image are provided. The system may receive, via a LDM, a reference image and a text prompt. The system may extract, via a trained vision encoder in the LDM, a vision control signal from an object in the reference image. The vision control signal indicates an identity of the object. The system may extract, via trained text encoders in the LDM, text control signals associated with the text prompt. The system may generate, via cross attention summation of an output of a vision cross attention unit(s) associated with the vision control signal and an output of text cross attention units associated with the text control signals, spatial features indicative of the reference image and the text prompt. The system may output, via a decoder in communication with the LDM, a target image based on the generated spatial features.
Owner:META PLATFORMS INC

First-view-angle drilling method and device based on memory enhancement and storage medium

The invention relates to the technical field of robot perception and data generation technologies, in particular to a first-view-angle drilling method and device based on memory enhancement and a storage medium, and the method comprises the steps: extracting a plurality of target frames from a historical video stream stored in a space intelligent machine, obtaining a plurality of memory elements based on the plurality of target frames, performing three-dimensional reconstruction on each target frame, constructing a target world model, planning a first visual angle track of the target robot in the target world model, performing imaging simulation on the target robot, performing real-time rendering on a first visual angle video of the target robot, and performing real-time rendering on a second visual angle video of the target robot based on the same time axis and action script as the first visual angle video. A third-person video corresponding to the space intelligent machine is generated, and a drilling result is output; according to the method, the long-term memory data of the space intelligent machine is ingeniously used, high-quality drilling data can be quickly generated, the reliability of the first-view video is improved, and a reliable basis is provided for training and testing of a robot algorithm.
Owner:BEIJING QIDAISONG TECH CO LTD

Concrete precast slab vibrating and leveling robot and method

The invention relates to the technical field of building construction automation, in particular to a concrete precast slab vibrating and leveling robot and method.The robot comprises a truss-type mechanical arm, a first truss-type Y-axis mechanical arm and a second truss-type Y-axis mechanical arm, and the truss-type mechanical arm comprises an X-axis movement mechanism, a first truss-type Y-axis mechanical arm and a second truss-type Y-axis mechanical arm; the first telescopic mechanical arm and the second telescopic mechanical arm are respectively mounted at the moving ends of the first truss type Y-axis mechanical arm and the second truss type Y-axis mechanical arm; the vibrating actuator is connected with the tail end of the first telescopic mechanical arm; the leveling actuator is connected with the tail end of the second telescopic mechanical arm; the visual device is used for photographing the concrete precast slab mold table entering the operation station according to a preset period to obtain a concrete slab image corresponding to the preset period, and sending the concrete slab image to the electrical control device; and the electrical control device is used for controlling the vibration actuator and the leveling actuator to sequentially complete vibration operation and leveling operation of the concrete prefabricated slab based on the BIM drawing of the prefabricated slab and the obtained concrete slab image.
Owner:CHINA STATE CONSTR HAILONG TECH CO LTD +1