Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

8588 results about "Image generation" patented technology

Image Generation is your video production partner because we care as much as you do, and we have the experience and ability to bring that passion to the screen on every project. Kevin McKeever, owner of Image Generation, produces and shoots every project personally, giving you a single point of contact.

System and Method for Multi-Modal Hyperspectral Image Generation with Cross-Modal Attention and Adaptive Quality Assurance

A system and method are disclosed for generating hyperspectral images from multi-modal sensor data including RGB, LiDAR, thermal, and near-infrared inputs. Training data includes hyperspectral images and corresponding multi-modal measurements. Spectral band grouping is performed based on correlation coefficients. A multi-modal decomposition network with cross-modal attention mechanisms generate reconstructed hyperspectral images by fusing complementary sensor information. A fine-tuning network creates reconstructed RGB images. A comprehensive quality assurance system analyzes spectral consistency, cross-modal coherence, and fusion artifacts to generate quality metrics. Missing data compensation strategies handle corrupted sensor inputs using information from other modalities. The system includes temporal integration for video sequences and multi-resolution processing for different sensor resolutions. Quality metrics guide network weight adjustments to improve reconstruction accuracy while maintaining robustness to sensor failures and environmental variations.
Owner:ATOMBEAM TECH INC

Multi-modal semantic and physical law driven remote sensing image generation method

The invention discloses a multi-modal semantic and physical law driven remote sensing image generation method, belongs to the technical field of computer vision and remote sensing image generation, and aims to solve the problems of insufficient cross-modal semantic alignment, low reliability of a generation result and insufficient physical mechanism fusion. The four-stage method comprises the following steps: firstly, rejecting low-quality samples from original data and unifying a spatial scale; then, extracting a multi-modal semantic vector by adopting a BLIP model and a CLIP model, and introducing a remote sensing physical rule to carry out vector optimization; then position coding and physical constraint conditions are embedded in the submerged space, and multi-source information joint modeling is achieved through a cross-modal encoder; and finally, by taking text description, physical priori knowledge and diffusion time steps as joint conditions, performing de-noising reasoning based on a Transform architecture, and completing back diffusion reconstruction by means of a trans-attention mechanism. According to the method, physical rationality and semantic consistency are improved, and a more reliable technical normal form is provided for remote sensing image generation in the fields of disaster monitoring, military simulation and the like.
Owner:CHINA UNIV OF MINING & TECH +2

Construction progress dynamic optimization method and system based on BIM and computer vision

The invention discloses a construction progress dynamic optimization method and system based on BIM and computer vision, and particularly relates to the technical field of building construction management, and the method comprises the steps: carrying out the automatic registration of a BIM model and a construction site image; processing the construction site image by adopting a visual identification algorithm to generate a visual identification result; constructing a four-dimensional dynamic BIM model, and mapping a visual identification result to a corresponding component in real time through multi-feature similarity calculation; the progress deviation is monitored by using key path dynamic identification and a deviation propagation matrix, and the risk is predicted by combining a Bayesian network and Monte Carlo simulation. The BIM and computer vision technologies are fused, a construction progress optimization system integrating automatic registration, dynamic monitoring, risk prediction and intelligent decision making is constructed, and the problems that traditional manual inspection data collection is low in efficiency, progress monitoring is lagged, risk prejudgment is fuzzy and resource allocation is extensive are solved; accurate monitoring, risk early warning and resource optimization configuration of the construction progress are realized.
Owner:ZHEJIANG LIDE ENGINEERING CONSULTING CO LTD

Transformer iron core detection method based on computer vision

The invention relates to the technical field of industrial component detection, in particular to a transformer iron core detection method based on computer vision, which comprises the following steps of: acquiring an iron core image, extracting key pixel characteristics, screening a directional scattering abnormal region to generate an interference map, extracting a consistent gradient region correction image to generate a reconstruction map, and positioning a symmetric disturbance generation structure map by integral gray difference. And analyzing an overlapping relation by a superposition structure graph to generate an abnormal component graph, and evaluating a risk level by matching a reference index to generate an early warning graph layer. Interference reflection and structural features can be distinguished through linkage analysis of the pixel direction vector and the brightness change frequency, correction of a distorted area in an image is realized based on a gray statistical stable value, and the distortion of the image is corrected by constructing a symmetric point map and analyzing the change trend of a gradient difference value sequence. And the structural overlapping relation is quantitatively judged by combining a component mapping profile diagram, so that the relevance between an abnormal region and a key component is clearly expressed, and the grading evaluation capability of various fault risks in the iron core is improved.
Owner:JIANGSU WEILAN DIGITAL INTELLIGENCE TECH CO LTD

Infrared and visible light image fusion method based on cross-domain Transform

The invention relates to an infrared and visible light image fusion method based on a cross-domain Transform, and belongs to the field of computer image processing. The method comprises the following steps: respectively carrying out preprocessing operation on an infrared image and a visible light image to obtain a training data set; an end-to-end image generator network is designed, an encoder module is used for extracting deep semantic features of an infrared image and a visible light image, a fusion module introduces an axial attention mechanism to enhance the global modeling capability of the features, and feature fusion is carried out in combination with information of a spatial domain and a frequency domain; the fused features are gradually recovered to an image space through a decoder module, and a fused image is generated; constructing a fusion loss function module, and guiding the network to focus a significant feature difference between the source image and the fusion image based on a comparative learning idea; and finally, inputting the infrared and visible light image Y channel into the network model, generating a fusion image, completing a training process, and realizing unified optimization of fusion performance and visual quality.
Owner:FUZHOU UNIV

Cross-modal joint source-channel coding and decoding method adaptable to changeable scenarios

PCT designated stageWO2026031415A1Internal combustion piston enginesBiological modelsChannel decoderImage signal
The present invention relates to the technical field of cross-modal image signal reconstruction. Disclosed is a cross-modal joint source-channel coding and decoding method adaptable to changeable scenarios. The method comprises: first designing a Transformer encoder-based cross-modal channel coding and decoding optimization solution, so as to achieve the performance improvement and robustness of a channel encoder and a channel decoder; then designing a cross-modal source coding and decoding optimization solution for haptic-to-image generation based on a latent diffusion model, so that under an image signal loss scenario, haptic information is used to guide image generation; and finally, incorporating transfer learning technology, so as to reduce additional training costs caused by a system facing changeable cross-modal communication scenarios such as a changeable channel signal-to-noise ratio and different transmission tasks. Under cross-modal changeable communication scenarios, the joint source-channel coding and decoding method provided in the present invention can solve the problems of the inability of a receiving end to well complete image reconstruction, and additional model training costs caused by changeable channel environments and scenarios.
Owner:NANJING UNIV OF POSTS & TELECOMM

Unmanned aerial vehicle autonomous inspection orthoimage generation method

The invention discloses a method for generating an autonomous inspection orthoimage of an unmanned aerial vehicle, and relates to the technical field of unmanned aerial vehicle surveying and mapping and autonomous navigation. Global optimization of an air route is realized through a path optimization strategy fused by a genetic algorithm in combination with grid characteristics of an inspection area and image parameter constraints, and candidate waypoints are used as nodes; redundant waypoints are screened through leg smoothness factors, route complexity is reduced, a genetic algorithm takes flight height and waypoint spacing as constraints, a fitness function containing flight distance, turning times and overlapping rate is constructed, a global optimal path is found through population iteration, multi-target requirements are optimized and balanced, and it is ensured that the route meets the image acquisition precision requirement and also meets the requirement of image acquisition. The flight distance can be shortened, the turning frequency is reduced, the cruising ability of the unmanned aerial vehicle is adapted, efficient propelling of the inspection task is guaranteed, meanwhile, waypoint coordinates output through simulation directly adapt to a flight control system, and it is guaranteed that actual flight parameters are consistent with planning parameters.
Owner:TUOHANG TECH CO LTD

Multi-modal image fusion method based on modal self-adaption and modal interaction compensation

The invention provides a multi-modal image fusion method based on modal self-adaption and modal interaction compensation, and the method comprises the following steps: S1, obtaining a multi-modal image fusion data set, and obtaining a training data set through preprocessing; S2, analyzing the modal difference characteristics of infrared and visible light images, and evaluating the correlation characteristics of image pairs in different scenes; s3, capturing a cross-modal feature dependency relationship through a self-attention mechanism; s4, a differential feature extraction strategy is adopted, model parameters are optimized through iterative training, and multi-modal image fusion is completed; s5, a modal interaction compensation module is additionally arranged, unit dynamic balance common features and modal exclusive features are fused, feature complementation is achieved in channel and space dimensions, parameters of the modal interaction compensation module are optimized, the model is made to learn the optimal fusion weight of the multi-modal features in a self-adaptive mode, and multi-modal fusion image generation optimization is achieved through the model; according to the invention, multi-modal image fusion can be accurately and effectively carried out.
Owner:FUZHOU UNIV

Small sample industrial defect detection system based on multi-stage diffusion model

The invention discloses a small sample industrial defect detection system based on a multi-stage diffusion model. The small sample industrial defect detection system comprises a multi-stage diffusion generation module and a defect detection module based on multi-scale attention and physical constraint. The multi-stage diffusion generation module divides the diffusion generation process into three stages of global structure reconstruction, local detail refinement and texture feature synthesis through a stage control mechanism, and gradually guides feature evolution and improves the generation effect aiming at the quality and diversity problems of defect image generation under the small sample condition; the defect detection module based on multi-scale attention and physical constraint adopts a coding structure fusing local window attention and global attention, through multi-scale feature extraction and fusion, significant features of a defect area are effectively captured, gradient smoothing loss and edge energy consistency loss are introduced at a decoder end, and the defect detection accuracy is improved. Therefore, high-quality reconstruction of a normal area and effective suppression of an abnormal area are realized.
Owner:ZHONGBEI UNIV +1

Flexible display module surface defect image recognition method

The invention relates to the technical field of industrial product surface quality detection, in particular to a flexible display module surface defect image recognition method, which comprises the following steps: acquiring a plurality of surface images of a flexible display module under different light sources and carrying out distortion removal processing on the surface images; reconstructing and generating three-dimensional reference point cloud data representing the current curved surface form of the module; re-projecting the distorted image to the ideal rigid plane according to the relationship, generating a plurality of corrected images, and generating a plurality of corrected images to eliminate geometric and luminosity distortion introduced by flexible deformation; obtaining a defect candidate area binary image; and extracting a multi-dimensional feature vector of the defect candidate region from the binary image of the defect candidate region, and classifying the feature vector by using a pre-trained defect classification model to obtain a defect identification result. Through the three-dimensional reference point cloud reconstruction and image re-projection technology, the problem of geometric distortion caused by surface deformation of the flexible display module is effectively solved, and misjudgment and missed judgment are avoided.
Owner:HUNAN HUICHENGXIN TECHNOLOGY CO LTD

Query evaluation for image retrieval and conditional image generation

Techniques are generally described for query evaluation for image retrieval and image generation. In various examples, a first encoded representation of first natural language input data may be generated. An image retrieval process may be selected from among the image retrieval process and an image generation process based at least in part on the first encoded representation of the first natural language input data. A second natural language encoder may generate a second encoded representation of the first natural language input data. The second encoded representation may be used to determine first image data stored in a first data repository. The first image data may be sent for output on a display of a first computing device.
Owner:AMAZON TECH INC

Tunnel disease identification model training method and system based on point cloud and image

The invention discloses a tunnel disease recognition model training method and system based on a point cloud and an image, and the method comprises the steps: synchronously collecting tunnel point cloud and image data, carrying out the calibration and registration, achieving the spatial alignment, preprocessing the point cloud, generating a gray-scale image and a depth image, collecting continuous images through a line-scan digital camera, generating a spliced image, and carrying out the recognition of tunnel diseases. Partitioning a large-size image after multi-image space-time synchronization; a multi-branch network is constructed, point cloud geometry and image texture features are extracted by using Point Net / 3DCNN and CNN / Transform respectively, and semantic collaborative fusion is realized through an intermediate layer fusion module; parameters are adjusted according to errors through self-adaptive training, manual labeling dependence is reduced in combination with transfer learning and the like, and convergence is accelerated through online iteration, joint loss and self-adaptive weight to improve generalization; the performance of the model is evaluated through field testing and indexes in various tunnel environments, and the structure is optimized according to data to ensure that engineering is feasible and efficient. The model can accurately identify various diseases on different tasks, and can adapt to complex and changeable working conditions in tunnel detection.
Owner:WUHAN HANNING TECH

Generative adversarial network-based MRI-PET mode conversion method and system

The invention discloses an MRI-PET mode conversion method and system based on a generative adversarial network, and belongs to the technical field of artificial intelligence medical image generation. And the multi-scale structure representation injection module injects multi-scale anatomical prior information at different stages of the encoder, and overcomes the limitations of insufficient utilization of prior information and single injection scale. And the adaptive semantic residual fusion module adopts semantic attention guidance and double-branch attention weighting, adaptively fuses fine-grained local features and global context information, harmonizes the difference between the fine-grained local features and the global context information in an abstract level and a semantic category, and solves the problems of feature conflict and semantic fuzziness in a bottleneck region. The direction sensing space-frequency discriminator realizes multi-dimensional and fine-grained adversarial supervision through a space, frequency and local image block multi-branch collaborative discrimination mechanism, and improves the structural fidelity and spectrum authenticity of a synthetic image. And the generated image is superior to the existing method in indexes such as structural similarity and peak signal-to-noise ratio, and has higher clinical practical value.
Owner:NORTHEASTERN UNIV AT QINHUANGDAO

Image multi-modal feature extraction, ground feature classification and recognition and GIS image generation method

The invention discloses an image multi-modal feature extraction method, a ground feature classification and recognition method and a GIS image generation method. Comprising the following steps: firstly, extracting texture features from a target image by adopting a multi-directional statistical method fusing rotation invariant coding of a local binary pattern and a gray-level co-occurrence matrix, and extracting color features from the target image by adopting an LAB-HSV dual-color space collaborative analysis method to obtain features of different modes of the target image; and then according to the information values of the extracted color features and texture features, adjusting the weight ratio corresponding to the color features and the texture features so as to optimize the recognition precision of the classification model on complex ground features. And finally, according to a weight ratio corresponding to the color feature and the texture feature, performing weighted fusion on the color feature and the texture feature to obtain a corresponding multi-modal feature vector.
Owner:CHONGQING GEOMATICS & REMOTE SENSING CENT

Real-time interactive image generation system based on multi-point touch canvas

The invention relates to the technical field of computer graphic interactive processing, in particular to a real-time interactive image generation system based on a multi-point touch canvas. The input acquisition unit is used for acquiring original touch data from an operating system and preprocessing the original touch data; the gesture recognition and analysis unit is used for performing high-level semantic behavior analysis on the contact data sequence processed by the preprocessing module; the interaction control and parameter mapping unit is used for receiving the semantic event output by the gesture recognition and analysis unit, analyzing the semantic event into an image control command and generating an executable command sequence; and the image generation unit generates interactive image content dynamically responded in real time based on an internal graph state management mechanism. By introducing the adaptive Kalman filtering and trajectory prediction auxiliary mechanism, the filtering intensity of the contact data can be dynamically adjusted, the efficient suppression of finger jitter and the consistent reconstruction of the contact ID are realized, and the input stability and data continuity under the multi-point touch operation are remarkably improved.
Owner:HUNAN VOCATIONAL COLLEGE OF SCI & TECH

Image relighting using machine learning

A method, apparatus, non-transitory computer readable medium, and system for image generation includes obtaining an input image and an input prompt, where the input image depicts an object and the input prompt describes a lighting condition for the object, generating relighted image features based on the input image and the input prompt, where the relighted image features represent the object with the lighting condition, and generating a synthetic image based on the relighted image features, where the synthetic image depicts the object with the lighting condition.
Owner:ADOBE INC

Text and image bidirectional alignment method and system based on multi-hop parallel reasoning

The invention discloses a text and image bidirectional alignment method and system based on multi-hop parallel reasoning, and the method comprises the steps: analyzing and marking a text, and obtaining a multi-granularity text feature group; image generation is processed, grids and target output are combined into a multi-scale visual feature group, an alignment index table is established, a sub-target chain is constructed based on visual features, the position and the dependency relationship are recorded, and explicit constraints are injected into five elements; recalling the candidate area, generating an anchor point through verification and fusion, writing the anchor point into a cross-modal index, locally completing three-layer decoupling scoring based on anchor point geometry and the cross-modal index, and outputting a three-state evidence in combination with a threshold value; the method comprises the steps of generating an evidence set through geometric and semantic consistency calibration, generating cross-hop evidence according to a rule, weighting and pooling the evidence set, generating a shared vector and traceable meta-information, inputting a task head output result and generating an evidence list and an auditing path, and according to the method, local alignment, constraint verification and evidence accumulation are completed hop by hop. And high-precision, traceable and interpretable cross-modal correspondence is realized.
Owner:SHAANXI NORMAL UNIV

Coal mining safety prediction visualization method and system

The invention relates to a coal mining safety prediction visualization method and system, and the method comprises the steps: determining a sensor deployment scheme based on mine geological structure information and operation environment characteristics; collecting multi-modal monitoring data according to the sensor deployment scheme, and generating a fusion data set of time-space alignment; extracting corresponding spatio-temporal features, and generating a corresponding risk assessment matrix based on the spatio-temporal features; constructing a four-dimensional tensor field containing a time dimension, and generating a dynamic risk visual image based on the four-dimensional tensor field; in the dynamic risk visual image generation process, receiving a feedback instruction input by an expert, and updating the coal mine safety domain knowledge graph by adopting a fuzzy cognitive mapping mode based on the feedback instruction; and optimizing the space-time convolutional network, updating the dynamic risk visualization image, obtaining an updated risk visualization image result, and executing hierarchical early warning processing. According to the method, the richness and accuracy of coal mining safety prediction visualization can be improved.
Owner:成都恒海峰科技有限公司

Vision-language model image-text pair accurate evaluation data construction method based on optimization algorithm

The invention relates to a visual-language model image-text pair accurate evaluation data construction method based on an optimization algorithm, and the method comprises the steps: firstly constructing an original image set through a manner of public data set screening, real-time equipment collection or deep generation, and reversely generating an initial cue word based on a pre-trained visual-language model; in combination with the constructed cue word template, optimizing the initial cue word by using a large language model, and generating a cue word highly matched with the picture; and then, performing optimization processing on the image-text pair data through a multi-dimensional evaluation function, performing manual verification on an optimized data set, removing low-quality or repeated image-text pairs, and finally constructing a high-quality vision-language model evaluation data set. According to the method, the matching degree and diversity of the image-text to the data are iteratively improved by adopting the optimization algorithm, the accuracy and coverage range of the evaluation data are remarkably improved, and the method can be widely applied to model performance evaluation of tasks such as image generation, visual question and answer and cross-modal retrieval.
Owner:THE THIRD RES INST OF MIN OF PUBLIC SECURITY

Image generation method and system based on style feature injection

The invention belongs to the technical field of artificial intelligence image generation, and particularly relates to an image generation method and system based on style feature injection, and the method comprises the steps: carrying out the multi-dimensional feature analysis of a reference image specified by a user side through a pre-trained multi-level style extraction model, extracting a style feature vector, and carrying out the feature extraction of the style feature vector; the style feature vector comprises a color feature vector, a texture feature vector, a composition feature vector and an illumination feature vector; obtaining a text description input by a user side, and outputting triple information including a scene entity, an entity attribute and a spatial relationship; performing weighted fusion on the triple information and the style feature vector to obtain a style enhanced semantic embedding vector; and inputting the style-enhanced semantic embedding vector into a generative adversarial network for splicing, and outputting an image. The method has the effect of remarkably improving the similarity between the generated image and the reference style.
Owner:GUANGDONG OPEN UNIV (GUANGDONG POLYTECHNIC VOCATIONAL COLLEGE)

Automobile part quality detection method and system based on artificial intelligence visual inspection

The invention discloses an automobile part quality detection method and system based on artificial intelligence visual inspection, and belongs to the field of artificial intelligence machine visual inspection, and the method comprises the steps: firstly, carrying out the registration of a collected RGB image and a depth image, and extracting a part region through a saliency detection network; two-dimensional key points are extracted based on an RGB region image and are matched with key points of a three-dimensional model, an initial three-dimensional attitude is obtained by adopting a PnP algorithm, iterative registration is performed with the three-dimensional model in combination with a point cloud generated by a depth image, and a fine three-dimensional attitude is obtained. And calculating a geometric transformation matrix from the part to a standard front view attitude according to the attitude, and performing attitude correction on the RGB and depth region image. And then matching the corrected image with a standard template image by using a feature detection and matching network so as to correct the position of the detection window. And finally, the three-dimensional size of the part is calculated in the corrected detection window in combination with the depth value, and tolerance judgment is carried out. And the precision, the robustness and the automation level of online detection of the automobile parts can be obviously improved.
Owner:XIANYANG VOCATIONAL TECHN COLLEGE

Visual token generation method and device based on shared index, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a visual token generation method, device, equipment and medium based on a shared index, and the method comprises the steps: obtaining an input image, and extracting semantic features and pixel features through a semantic encoder and a pixel encoder; calculating the distance between each feature and a codebook thereof, and carrying out weighted summation to determine a shared index; retrieving quantitative features from the codebook using a shared index; respectively generating a reconstructed image and a reconstructed semantic feature by using a pixel decoder and a semantic decoder; jointly optimizing an encoder, a codebook and a decoder based on a reconstruction result; and generating a unified visual token sequence for the target task image by using the optimized component. Through double-flow feature extraction, shared mapping quantization and joint loss optimization, global semantic information and local pixel details can be reserved in the visual token at the same time, so that the model has accurate understanding ability, high-fidelity images can be generated, and the performance of understanding and generating tasks is improved.
Owner:PING AN TECH (BEIJING) CO LTD

Personalized image generation using combined image features

Examples described herein relate to personalized image generation using combined image features. A plurality of input images is provided by a user of an interaction application. Each of the plurality of input images depicts at least part of a subject. Each input image is encoded to obtain an identity representation. The identity representations obtained from the plurality of input images are combined to obtain a combined identity representation associated with the subject. A personalized output image is generated via a generative machine learning model. The generative machine learning model processes the combined identity representation and at least one additional image generation control to generate the personalized output image. At a user device, the personalized output image is presented in a user interface of the interaction application.
Owner:SNAP INC

Control method and device based on task understanding representation, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business scenes such as personal intelligence, financial science and technology and medical health, and discloses a control method, device and equipment based on task understanding characterization and a medium, and the method comprises the steps: obtaining an environment image of a to-be-processed scene and a task instruction for specifying an operation task, a visual encoder is utilized to process an environment image to generate a visual feature vector, a language encoder is utilized to process a task instruction to generate a semantic representation, the visual feature vector and the semantic representation are fused to obtain a fused feature, and the fused feature is input into a pre-training model to generate a task understanding representation. And generating an action sequence by using an action decoder based on a diffusion model and a flow matching technology, and controlling an execution device to execute an operation according to the action sequence. According to the method, multi-modal information is fused, a task understanding mechanism is introduced, a high-reliability action sequence is generated in combination with a diffusion model and a flow matching technology, and the response ability to variable task instructions and the operation accuracy can be improved in a complex environment.
Owner:PING AN TECH (SHENZHEN) CO LTD

Training and deployment of image generation models

In some embodiments, a method receives a text prompt. A text encoder is executed on the text prompt to generate a representation. The method generates a set of images based on the representation and a set of parameters of an image generation model. The set of images is ranked using reward values that are generated by a reward model. The reward model is trained using human input that provided feedback on a quality of generated images using the image generation model. The method outputs one or more images based on the ranking in response to the text prompt.
Owner:CASTLE GLOBAL INC

Industrial defect image generation system and method based on deep learning

The invention provides an industrial defect image generation system and method based on deep learning, and relates to the technical field of artificial intelligence. A defect form adaptive module, a physical attribute modulation module and a multi-mechanism fusion module are integrated in a defect image generation module; dynamically selecting a feature extraction unit according to the defect type label based on a pre-constructed generative network model, and generating a defect feature map; generating an affine transformation parameter based on the physical attribute vector, and performing channel-by-channel linear modulation on a defect feature map of a middle layer of the generative network model, so that a multi-scale defect feature map finally generated by the generative network model contains specified defect type features and physical attribute features; and fusing the multi-scale defect feature image and the defect-free background image to obtain a fused defect image. The problems that in the prior art, industrial defect image generation is insufficient in sense of reality, poor in controllability and poor in fusion effect are solved, the method can be used for data enhancement of industrial visual inspection, and downstream model performance is improved.
Owner:CHANGZHOU XINGYU AUTOMOTIVE LIGHTING SYST CO LTD

System and Method for Enhancing Generative Artificial Intelligence (AI) Model-Based Document Search with Image Retrieval

A method, computer program product, and computing system for generating a plurality of chunks for a plurality of text portions of a document, wherein the document includes the plurality of text portions and a plurality of images. Each chunk is indexed using a word embedding. Each of the plurality of images is indexed based upon, at least in part, a position of a respective image relative to a corresponding chunk. An image placeholder is generated for each of the plurality of images. A plurality of image-enhanced embeddings is generated by inserting the image placeholder for each of the plurality of images into a respective word embedding for the corresponding chunk. The plurality of image-enhanced embeddings are provided for processing a query using a generative artificial intelligence (AI) model.
Owner:DELL PROD LP

Image generation method and device based on theme information, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business scenes such as financial science and technology and medical health, and discloses an image generation method, device and equipment based on theme information and a medium. The method comprises the steps that input information is analyzed to generate theme information and copywriting information, a cue word set is generated, and a figure image set and a background image set are generated; fitting the segmented figure image with the background image to form a head image candidate set, and selecting a head image matching template frame to generate a basic image; decomposing the copywriting to generate a sub-module initial picture set, and adding a gradual change effect to form a sub-module picture set; the adjusted sub-module pictures are obtained based on size adjustment, and a pre-synthesized image is generated through splicing; and identifying the blank area to draw a title text to obtain a final image. Through theme analysis, template matching, image splicing, modular copywriting processing and blank drawing, poster generation efficiency is improved, layout flexibility is enhanced, and visual unification is realized.
Owner:CHINA PING AN PROPERTY INSURANCE CO LTD