Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

9176results about "Editing/combining figures or text" patented technology

Vision generation method and device based on semantic association modeling, equipment and medium

The invention relates to the technical field of voice semantics, can be applied to business scenes of financial science and technology, medical health, poster design and the like, and discloses a visual sense generation method and device based on semantic association modeling, equipment and a medium. Generating a demand text containing theme and style parameters; semantic features in the demand text are extracted, semantic association weights are constructed, and element layout coordinates are optimized in combination with spatial distribution constraints; and encoding the layout information into a control matrix, fusing the control matrix with the initial noise, adjusting a noise reduction process through an encoding and decoding network, and generating target visual content highly matched with the semantic meaning of the user instruction. According to the method, the layout optimization function is constructed, the diffusion model is guided to focus the semantic salient region in space, language model output and the visual generation process are closely combined, structured response and space mapping of user semantic requirements are achieved, and the expression consistency and personalized adaptation capacity of visual content generation are improved.
Owner:SHENZHEN PINGAN COMM TECH CO LTD

Cooperative generation method for dynamic visual content based on cognitive logic chain

The invention discloses a dynamic visual content collaborative generation method based on a cognitive logic chain, and belongs to the technical field of visual content generation, and the method comprises the following steps: S1, user intention analysis and data input; s2, dynamically constructing a cognitive logic chain; s3, intelligent scheduling of the multi-modal generation module; s4, cross-modal content collaborative generation is carried out; s5, collaborative editing and real-time feedback are carried out; s6, iterative optimization of logic chain driving; s7, multi-dimensional quality evaluation: constructing an evaluation matrix containing semantic consistency, visual attraction and user participation degree, predicting a content propagation effect in combination with a deep learning model, and generating a quantitative improvement suggestion report; and S8, updating the self-adaptive knowledge reversely marking the cognitive logic chain according to the finally adopted content version, extracting a new association rule, and injecting the new association rule into the rule base. Through deep semantic analysis and dynamic logic chain construction, the system accurately captures a core creation target of a user and converts the core creation target into an executable visual strategy.
Owner:SHUCHUANGUANHU (HANGZHOU) INFORMATION TECHNOLOGY CO LTD

Advertisement copywriting generation method and device based on multi-modal fusion, equipment and medium

The invention discloses an advertisement copywriting generation method and device based on multi-modal fusion, equipment and a medium, and relates to the technical field of advertisement marketing. The method comprises the steps of obtaining multi-modal data of a target video, the multi-modal data comprising visual data, auditory data and related metadata of the video, and extracting associated data from a local knowledge base; and preprocessing the multi-modal data and the local knowledge base data, and converting the multi-modal data and the local knowledge base data into feature forms which can be used for analysis. According to the method, semantic calibration is carried out on multi-modal input by means of a local knowledge base, it is ensured that generated content strictly follows domain knowledge constraints, the problem of deviation caused by the fact that a traditional model depends on implicit knowledge is solved, multi-modal information ambiguity is eliminated, collaborative semantic generation of texts, images and structured data is achieved, content dimensions are enriched, and the method has the advantages of being high in practicability and easy to popularize. And an efficient solution is provided for the landing of the intelligent generation technology in the vertical field.
Owner:SHANGHAI WANGMAI INFORMATION TECH GRP CO LTD

Cold chain cargo transportation management visualization method and system

PendingCN120851756AMeasurement devicesForecastingCold chainTransportation energy
The invention relates to the field of transportation visualization, in particular to a cold chain cargo transportation management visualization method and system. The method comprises the following steps of collecting cold chain cargo state parameters, performing cargo quality trend evaluation, and generating state data streams of a plurality of cargoes; performing dynamic positioning and tracking on the cold-chain transport vehicle, performing space-time mapping processing according to the state data flow, and constructing a real-time tracking trajectory flow; performing future aging deviation prediction on the real-time tracking trajectory flow to generate an aging deviation prediction result; carrying out transportation risk comprehensive assessment according to the aging deviation prediction result, and generating a transportation risk assessment report; and carrying out path planning calculation and client notification pushing according to the transportation risk assessment report so as to execute intelligent transportation management. According to the invention, the transportation energy consumption and time cost are reduced, the resource allocation efficiency is improved, and the cargo real-time tracking visualization effect is improved.
Owner:SHENZHEN QIANHAI YUESHI INFORMATION TECH CO LTD

Rich-Media Document Auxiliary Generation Apparatus

Disclosed in the present disclosure is a rich-media document auxiliary generation apparatus. The apparatus comprises a material extraction module, a theme sorting module, a semantic retrieval module, a structured data text generation module, an illustration recommendation module and a video composition module. The present disclosure uses intelligent means to assist a user to efficiently generate a high-quality rich-media composite document, thereby quickly and accurately describing a theme event in an all-round way.
Owner:10TH RES INST OF CETC

Electrical audio signal processing systems and devices

According to an aspect of the present invention, there is provided an electrical audio signal processing system and device, comprising: a computer graphics processing and selective visual display system with a screen; an eye tracking device; a processor; one or more computer memory devices; wherein the processor is arranged for operations comprising: measuring the user's eye movements to ascertain the specific word on which the user is fixated, by the eye tracking device; modifying the display at the user's current fixation point; applying a delay between the presentation of successive graphic elements based on the user's calculated rate to accommodate the user's required time; and presenting elements to the user at a rate based upon the user's required time.
Owner:DECHARMS RICHARD CHRISTOPHER

Medical image report generation method based on multi-modal large model preference alignment technology

The invention relates to the technical field of medical images, in particular to a medical image report generation method based on a multi-modal large model preference alignment technology, which comprises the following steps of: based on medical image data, collecting known medical report data corresponding to an image, analyzing focus description and image explanation vocabularies in a report text, and generating a medical image report; and establishing image text associated data. According to the method, by combining the medical image data and the corresponding report, high automation of data processing is achieved, it is ensured that the generated medical report is more accurate in details, analysis of the relation between the image and the text is deepened through the VLM model, accurate feature vector mapping is achieved, the scientificity of text generation is improved, and the user experience is improved. Through a fine preference alignment process, diagnosis preferences and actual operation habits of experts can be accurately reflected in report generation, and diagnosis words, sentence pattern structures and information arrangement sequences of the experts are analyzed so as to better match annotations and diagnosis thinking modes of the experts.
Owner:砺进(杭州)科技有限公司

Ai-based visual content collage generation

A data processing system implements receiving, via a user interface of a client device, images for generating a collage image; generating captions for the images; constructing a first prompt by appending the captions to a first instruction string including instructions to a generative language model to extract a theme from the captions; providing the first prompt to the generative language model and receiving the theme therefrom; constructing a second prompt by appending the theme to a second instruction string including instructions to a text-to-image model to use the theme to create a background image with placeholders; providing the second prompt to the text-to-image model and receiving the background image therefrom; identifying the placeholders in the background image; creating the collage image by fitting the images into the identified placeholders; providing the collage image to the client device; and causing the user interface to display the collage image.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Multi-modal data driven general report generation method and system based on large model

The invention discloses a multi-modal data driven general report generation method and system based on a large model, and belongs to the technical field of intelligent report generation. Firstly, texts, images and sensor data related to a report theme are obtained and subjected to standardized preprocessing; analyzing the report generation instruction, and matching and querying a task modal mapping library matching modal configuration scheme according to a task demand; quantitatively evaluating the data quality of each modal, dynamically calculating the final decision weight of each modal in combination with the basic weight, and distributing the final decision weight to a corresponding processing path to form dominant, supplementary and reference data; inputting the dominant data and the supplementary data into a multi-modal model for analysis to obtain a preliminary conclusion with confidence score, performing consistency judgment, if no conflict exists, performing fusion to form a comprehensive conclusion, and if the conflict exists, combining a quality evaluation result and an arbitration rule to complete conflict judgment; and finally, inputting the comprehensive conclusion and the reference data into a large language model to generate a report text, and outputting a complete report after typesetting and proofreading.
Owner:NANJING ANCIENT NETWORK TECH CO LTD

Video Query Contextualization

Systems and methods for video query contextualization can include a router model that determines how to process and respond to the query associated with the video. The systems and methods can include obtaining an input query and video data, processing the input query and the video data with the router model to generate a video clip and routing data, and the routing data can then be utilized to determine which processing system to utilize to process the video clip and the input query. The video clip can then be processed with the determined processing system to generate a query response that may be provided to the user.
Owner:GOOGLE LLC

Medical image generation method and device based on bimodal fusion, equipment and medium

The invention discloses a medical image generation method and device based on bimodal fusion, equipment and a medium, and the method comprises the steps: respectively extracting the visual features of a medical image and the semantic features of text description through an image encoder and a text encoder; mapping the two types of features to a shared semantic space by adopting comparative learning to realize cross-modal alignment; a first path captures hierarchical semantic information through a convolutional neural network, a second path retains local significant features through maximum pooling, and two paths of outputs are fused layer by layer to construct spatial context features; projecting the cross-modal features into a spatial feature map through a multi-layer perceptron, splicing the spatial feature map with coding features, and inputting the spliced spatial feature map into a decoder for up-sampling reconstruction; and performing joint optimization on the comparison loss and the structural similarity loss to realize end-to-end training. And finally, the SSIM index of the generated medical image is improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Picture editing method and system based on multi-modal condition adaptation

The invention provides a picture editing method and system based on multi-modal condition adaptation, and the method comprises the steps: obtaining a first text vector: obtaining a picture description corresponding to a picture through processing, processing the picture description through an editor, and obtaining a first text vector; obtaining a first text vector based on the picture information; a second text vector acquisition step: processing the used editing instruction to obtain a second text vector based on the editing instruction; a fusion feature acquisition step: fusing the first text vector and the second text vector through weight to obtain a fusion feature; an image potential feature code acquisition step; in the image editing step, the injected condition information is received, meanwhile, the received potential features and potential noise of the image are denoised, and the image desired by the user is gradually generated in the iterative denoising process under the guidance of the received condition information; and an image restoration step. According to the invention, the stability, controllability and accuracy of image editing can be improved.
Owner:SHENZHEN EMDOOR DIGITAL TECH

Music segment tagging, sharing, and image generation

A method of automated generation of contextually-relevant images for a music segment includes receiving at least one of basic metadata information and lyric information for the music segment, generating a first prompt for a computer-implemented machine-learning language model based on the at least one of the basic metadata information and the lyric information, receiving context information from the computer-implemented machine-learning language model in response to the first prompt, generating a second prompt for the computer-implemented machine-learning language model based on the context information, generating a third prompt by providing the second prompt as an input to the computer-implemented machine-learning language model, and generating an image descriptive of the music segment by providing the third prompt as an input to a computer-implemented machine-learning image generation model.
Owner:HOOK MEDIA LLC

Defect picture generation method and device, electronic equipment and storage medium

The invention relates to a defect picture generation method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining a workpiece image, text prompt information and a mask image, the mask image is generated based on the workpiece image and mask information, and the mask information is used for representing the shape and position of a to-be-generated defect on the workpiece image; the workpiece image, the text prompt information and the mask image are input into a pre-trained defect image generation model, a target defect image is obtained, the defect image generation model is used for extracting a first generation feature from the workpiece image and the text prompt information, extracting a first shape feature from the mask image, and generating a second generation feature; based on the first generation feature and the first shape feature, a target defect image is generated, and the defect generated in the target defect image meets the constraint of the mask information on the defect shape and the defect position. In this way, accurate control over the defect shape and the defect position can be achieved, and the target defect image meets the actual requirement.
Owner:SHENZHEN XINRUN FULIAN DIGITAL TECH CO LTD

Text-driven CAD modeling method and system based on diffusion and visual language model

The invention relates to the technical field of computer aided design, in particular to a text-driven CAD modeling method and system based on a diffusion and visual language model.The method comprises the steps that natural language text description is obtained, and CAD semantic features of the natural language text description are extracted; carrying out geometric standardization on the CAD semantic features by adopting a fine-tuning diffusion model, and generating a CAD view image conforming to engineering specifications; carrying out fusion by adopting a fine-tuned visual language model to generate a parameterized CAD construction sequence; a three-mode alignment mechanism is adopted, and the semantic consistency of the CAD semantic features, the CAD view images and the CAD construction sequences is checked; performing verification and post-processing on the CAD construction sequence, and outputting an executable Python code or STEP file; the CAD modeling method disclosed by the invention performs explicit modeling based on flexible modal description, and has the characteristics of high geometric constraint and high usability.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Avatar creation user interface

The present disclosure generally relates to creating and editing avatars, and navigating avatar selection interfaces. In some examples, an avatar feature user interface includes a plurality of feature options that can be customized to create an avatar. In some examples, different types of avatars can be managed for use in different applications. In some examples, an interface is provided for navigating types of avatars for an application.
Owner:APPLE INC

Fine-tuning diffusion-based generative neural networks using singular value decompositions for text-to-image generation

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for fine-tuning diffusion-based generative neural networks in compact parameter spaces for text-to-image generation. In one aspect, a method performed by one or more computers for fine-tuning a diffusion-based generative neural network to obtain a fine-tuned version of the diffusion-based generative neural network is described. The method includes: for each of a number of neural network layers of the diffusion-based generative neural network: obtaining an initial weight matrix including a number of pre-trained weights parametrizing the neural network layer: performing a singular value decomposition on the initial weight matrix; and re-parametrizing the neural network layer with new weights that depend on spectral sifts; and training the spectral shifts of each of the number of neural network layers of the diffusion-based generative neural network to obtain the fine-tuned version of the diffusion-based generative neural network.
Owner:GOOGLE LLC

Multi-modal consistency verification and harmonization models

A system may access content comprising text content, visual content, and / or audio content. The system may perform, based on a harmonization model and / or consistency model, a harmonization check and / or a consistency check on the content. The system may recognize, based on the harmonization check and / or the consistency check, a conflict to be corrected. The system may identify a property of the content that should be changed based on the recognized conflict. The system may generate a corrective action based on the property of the content that should be changed.
Owner:REVE AI INC

Learning continuous control for 3d-aware image generation on text-to-image diffusion models

A method, apparatus, non-transitory computer readable medium, and system for image processing include obtaining a text prompt describing an element and an attribute value for a continuous attribute of the element, embedding the text prompt to obtain a text embedding in a text embedding space, embedding the attribute value to obtain an attribute embedding in the text embedding space, and generating a synthetic image based on the text embedding and the attribute embedding, where the synthetic image depicts the continuous attribute of the element based on the attribute value.
Owner:ADOBE INC

Map construction method and system based on laser vision dynamic weighted fusion

The invention provides a map construction method and system based on laser vision dynamic weighted fusion, and belongs to the technical field of data processing, and the method comprises the steps: obtaining an RGB image and a depth image through a vision sensor, and obtaining point cloud data through a laser radar; aligning the depth image with the RGB image, and removing invalid depth pixels; based on an ORB-SLAM2 algorithm, performing visual SLAM, extracting features of the depth image and the RGB image, and outputting sparse visual point cloud data; based on a Gmapping algorithm, performing laser radar SLAM, and outputting a 2D occupied grid map; projecting the sparse visual point cloud data into a 2D occupation grid map, and calculating the semantic occupation probability of each grid; calculating the geometric occupancy probability of each grid through an anti-sensor model; performing weighted fusion on the semantic occupancy probability and the geometric occupancy probability, and calculating a fusion occupancy probability; and generating a fusion map according to the fusion occupation probability.
Owner:SHIHEZI UNIVERSITY

Ai-based shape-adaptive consistent visual effect generation

A data processing system implements constructing a first prompt including a font mask of a reference character (RC) and a style prompt, sending the first prompt to a text2image model to iteratively generate salient content and concentrate the salient content within the font mask of RC as a first image of RC; concatenating two of the first images as a second image; generating a combined font mask of the font mask of RC and a font mask of a target character (TC); constructing a second prompt including the combined font mask and the second image, sending the second prompt to the model to iteratively generate salient content and in-paint the salient content within a half of the combined font mask as a third image of RC and TC; cropping a styled TC image from the third image using the font mask of TC; providing the styled TC image to a client device.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Direct regression encoder architecture and training

Systems and methods train and apply a specialized encoder neural network for fast and accurate projection into the latent space of a Generative Adversarial Network (GAN). The specialized encoder neural network includes an input layer, a feature extraction layer, and a bottleneck layer positioned after the feature extraction layer. The projection process includes providing an input image to the encoder and producing, by the encoder, a latent space representation of the input image. Producing the latent space representation includes extracting a feature vector from the feature extraction layer, providing the feature vector lo the bottleneck layer as input, and producing the latent space representation as output. The latent space representation produced by the encoder is provided as input to the GAN, which generates an output image based upon the latent space representation.
Owner:ADOBE INC

Generating a modified digital image utilizing a human inpainting model

The present disclosure relates to systems, methods, and non-transitory computer-readable media that modify digital images via scene-based editing using image understanding facilitated by artificial intelligence. For example, in one or more embodiments the disclosed systems utilize generative machine learning models to create modified digital images portraying human subjects. In particular, the disclosed systems generate modified digital images by performing infill modifications to complete a digital image or human inpainting for portions of a digital image that portrays a human. Moreover, in some embodiments, the disclosed systems perform reposing of subjects portrayed within a digital image to generate modified digital images. In addition, the disclosed systems in some embodiments perform facial expression transfer and facial expression animations to generate modified digital images or animations.
Owner:ADOBE INC

Apparatus and method for generating photorealistic synthetic images

A method and apparatus for generating photorealistic synthetic images by receiving multiple forms and multiple instances of user input corresponding to a user's visual idea, executing an iterative image search to identify pre-existing images semantically aligned with the user's visual idea, and using an image synthesis deep learning model to generate at least one synthetic image based on the multiple forms and instances of user inputs.
Owner:BERSERQ PTE LTD +3

Systems and methods for contextual and semantic summarization

A system may, in a first pass: divide content to be summarized into a plurality of chunks and, for each chunk: execute a language model with the chunk and an instruction to summarize the chunk, generate, based on the executed language model, a summary of the chunk. In a subsequent pass, the system may: generate a plurality of groups of summaries, each group of summaries from among the plurality of groups of summaries comprising two or more summaries, each summary corresponding to a respective chunk, for each group of summaries from among the plurality of groups: execute a language model with the group of summaries and an instruction to summarize the group of summaries, generate, based on the executed language model on the group of summaries, a group summary. The system may iteratively repeat the subsequent pass for group summaries until a summary of the content is reached.
Owner:REVE AI INC

Image relighting using machine learning

A method, apparatus, non-transitory computer readable medium, and system for image generation includes obtaining an input image and an input prompt, where the input image depicts an object and the input prompt describes a lighting condition for the object, generating relighted image features based on the input image and the input prompt, where the relighted image features represent the object with the lighting condition, and generating a synthetic image based on the relighted image features, where the synthetic image depicts the object with the lighting condition.
Owner:ADOBE INC

Method for editing facial image attributes based on text of diffusion model

The invention provides a method for editing facial image attributes based on a text of a diffusion model. The method realizes flexible facial editing with high quality and identity consistency. The method comprises the following steps: constructing a description control face diffusion model comprising a noise prediction network and a variational auto-encoder, wherein the noise prediction network comprises a plurality of text alignment face Transform modules and an optimization residual feature module; the method comprises the following steps: inputting an original image and target text description based on a pre-training model, optimizing a target embedding generated by the input target text description by a text alignment face Transform module to obtain an optimized embedding, and performing model fine tuning after optimization is completed; and performing linear interpolation on the target embedding and the optimized embedding to obtain an initial editing result, introducing ArcFace-Loss as identity loss, extracting face features of an original image and the initial editing result through a pre-training model, calculating feature similarity and minimizing the loss, and ensuring consistency of face identities after editing.
Owner:DALIAN NATIONALITIES UNIVERSITY

Adaptive rendering method and system for characters and images in AI digital human virtual and real scene fusion

The invention discloses an adaptive rendering method and system for characters and images in AI digital human virtual and real scene fusion, and the method comprises the following steps: S1, collecting Chinese text and image data, and extracting semantic and regional features; s2, constructing a Shenchang differential equation converter model, and inputting text and image features to generate fusion representation; s3, extracting an attention weight matrix, and determining the spatial arrangement position of the Chinese text in the image; s4, optimizing the structural parameter vector by adopting a luminous insect colony optimization algorithm to obtain a fusion model; s5, generating an image-text rendering instruction set based on the preliminary arrangement and optimization model; s6, in combination with the digital human action state and background image information, attitude mapping and illumination adjustment are completed to generate a rendering frame; and S7, inputting the rendering frame into a rendering engine, and outputting the video content fused by the Chinese characters and the images. According to the method, accurate fusion of Chinese semantics and images is realized, and the naturalness and interactivity of AI digital human image-text presentation are remarkably improved.
Owner:BAIGE ONLINE (XIAMEN) DIGITAL TECHNOLOGY CO LTD

Online resource text and image layout system based on natural language processing

The invention discloses an online resource text and image layout system based on natural language processing, and belongs to the technical field of natural language processing. An image semantic understanding module; a layout generation and optimization module; an interactive adjustment module; a multi-modal content fusion module; the aesthetic evaluation and compliance verification module is used for automatically evaluating the aesthetic degree and compliance of the generated layout; and the dynamic response type adaptation module is used for adjusting the layout in real time according to the sizes of terminal equipment and a screen, ensuring the cross-platform consistency, generating a multi-resolution layout scheme based on a Flexbox response type rule, and dynamically adjusting the element stacking sequence, the personalized recommendation module and the real-time cooperation and version management module through reinforcement learning. And the edge calculation module improves the efficiency of large-scale layout generation, supports an offline or low-delay scene, compresses a CNN model by using a knowledge distillation or quantification technology, and adapts to edge equipment. The invention also discloses a method.
Owner:BEIJING DIGITAL FUTURE TECHNOLOGY CO LTD