Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

352 results about "Image description" patented technology

An image description is a textual, audio or graphical content portraying the image in a representation intelligible by the addressees. The description should be comprehensive as well as perceptible by the target audience.

Target detection method and device, model training method and device, electronic equipment and medium

The invention relates to the technical field of data processing, and provides a target detection method and device, a model training method and device, electronic equipment and a medium. The target detection method comprises the steps that a to-be-recognized image and a query text are acquired, and the query text is used for querying a target object corresponding to the query text in the to-be-recognized image; performing image recognition on the to-be-recognized image to obtain image description features and region detection visual features; performing regional multi-modal fusion processing on the image description features and the regional detection visual features to obtain regional multi-modal fusion features; performing feature fusion processing on text features obtained based on the query text and the regional multi-modal fusion features to obtain text regional fusion features corresponding to the query text; and a target detection result is obtained based on the text features and the text region fusion features, so that the fusion degree of text semantics and image region features is improved, and the accuracy and robustness of target detection in a complex scene are improved.
Owner:BEIJING JIZHI DIGITAL TECH CO LTD

Intellectual visual question and answer method and device based on big and small model collaboration and medium

The invention discloses a knowledge-based visual question-answering method and device based on big and small model collaboration and a medium, belongs to the technical field of visual question-answering, and solves the problem of how to improve the accuracy of complex questions faced by a visual question-answering technology. In the image description extraction step, an image description and key object label set which is highly associated with a natural language problem is generated, and in the entity enhancement processing step, key entities are extracted and a clarity problem and an entity example which are used for resolving semantic ambiguity are generated. In the candidate answer generation step, a plurality of semantically complementary answers are generated under different reasoning dimensions on the basis of a multi-round generation strategy, optimized candidate answers are screened on the basis of a model built-in scoring mechanism, in the context example retrieval step, examples most related to current input are retrieved, and finally, the candidate answers are obtained. And the large language model receives a unified input prompt formed by splicing output results in the steps, and inference is performed in an autoregression mode to generate answers, so that the accuracy of the visual question-answering technology facing complex questions is effectively improved.
Owner:STATE GRID ANHUI ELECTRIC POWER CO LTD ELECTRIC POWER SCI RES INST

Multi-language object illusion relieving method based on cross-language attention mode

The invention discloses a multi-language object illusion relieving method based on a cross-language attention mode, solves the problem of how to relieve multi-language object illusion during detection of a large visual language model under non-English questions, and belongs to the technical field of multi-mode questions and answers. The method comprises the following steps: identifying a cross-modal attention head set of which a visual language model shows obviously different behaviors for English and a target language when same semantic questions in different languages are processed; constructing image description queries of English and target languages for the same image, respectively inputting the image description queries into respective visual language models for reasoning, obtaining attention output under the English and target languages, and taking an average difference between the attention output as a language migration vector of the target languages; and in the reasoning process of the target language question, intervening the attention heads in the attention head set by using the language migration vector, so that the visual understanding ability of the visual language model under the non-English question is closer to the English question.
Owner:HARBIN INST OF TECH

Image processing model

A method comprising generating image descriptions of images in an original training set of images; determining, using at least one LLM, at least one domain and / or class which is under-represented in the original training set; generating, using a second LLM and based on the determination of the at least one domain and / or class, at least one instruction for a third LLM to generate at least one text prompt; generating, using the third LLM and based on the at least one instruction, the at least one text prompt for a text-to-image model; generating, using the text-to-image model and based on the at least one text prompt, at least one synthetic image; and generating an enhanced training set of images for use in training an image processing machine learning, ML, model, the enhanced training set of images comprising the original training set of images and the at least one synthetic image.
Owner:FUJITSU LTD +1

Image text description generation method, electronic equipment and readable storage medium

The invention provides an image text description generation method, electronic equipment and a readable storage medium. According to the method, the object perception prototype learning module and the global context feature extraction module are introduced, so that fine-grained information and global semantic understanding in the image are effectively balanced. The visual backbone network module can extract multi-scale and multi-level image features and perform fusion, thereby enhancing the expression ability of the image features. The object perception prototype learning module further extracts an object prototype from the fusion features to ensure that the model can accurately capture key objects and attributes thereof in the image, and the global context feature extraction module ensures that the overall context of the image is fully understood. On the basis, the encoding and decoding module combines the global context and the object prototype to generate the text description, so that the semantic splitting phenomenon in the traditional method is avoided, and the detail information in the image is effectively reserved, thereby improving the accuracy and integrity of the image description.
Owner:WUHAN UNIV

Unified streaming processing method, system and device for multi-mode AI interactive content, medium and program product

The invention discloses a unified streaming processing method, system and device for multi-modal AI interactive content, a medium and a program product, and the method comprises the steps: receiving a request which is sent by a user and comprises an image and a text, and constructing a request context; extracting image features; carrying out image analysis and outputting an image description text; performing intention recognition and outputting an intention; multi-stage reasoning is carried out, and thinking content fragments are generated in a streaming mode; buffering and releasing thinking content fragments; generating text content fragments based on multi-stage reasoning; buffering and releasing the text content fragments; performing recommendation triggering based on the image description text and the intention; recommending content fragments; inserting a recommended position, buffering and issuing; and process event management: issuing event notifications when the beginning, any step fails and the end. The method is a universal method capable of transmitting different types of outputs in a single ordered stream, the analysis cost can be reduced, and the interaction experience and expansibility are improved.
Owner:BEIJING DIANFU TECHNOLOGY CO LTD

Automatic software testing method and system based on multi-mode AI cooperation

The invention relates to the technical field of automatic testing, in particular to an automatic software testing method and system based on multi-mode AI cooperation. According to the scheme, optical character recognition, voice recognition and natural language processing analysis tools are packaged into independent containers, and analysis results are input into a multi-mode semantic association model in real time; calculating the semantic similarity between the image description text and the corresponding voice transcription text; constructing a test demand complexity evaluation model; defining use case quality indexes, and establishing an error sample library to store manually labeled problem use cases and correction labels thereof; the intention demand is finally determined according to the screening result of the three-stage filter; and setting a test case template, optimizing a path according to the fusion quality, improving the case generation efficiency, and synchronously updating the optimized path to a knowledge base associated with the error sample library. According to the scheme, automatic software testing is achieved through multi-modal AI cooperation in combination with containerized deployment, model dynamic optimization and the like, and efficiency and recognition accuracy are improved.
Owner:RUIJIAN TECHNOLOGY (BEIJING) CO LTD

Document structure extraction and model training method and device, equipment and medium

The invention discloses a document structure extraction and model training method and device, equipment and a medium, and relates to the technical field of artificial intelligence and computer vision. The method comprises the following steps: constructing a special training data set containing data of at least two document understanding tasks (including optical character recognition, layout analysis, text positioning, regional text extraction, image description and chart title generation), and a fine tuning data set for converting a document image into a machine-readable structured text format; constructing a multi-modal large model comprising a shape adaptive cutting module, a visual encoder, a visual token compression module, a modal connector and a language decoder; pre-training the model by using the special training data set to jointly learn various document understanding tasks; and performing fine tuning on the pre-trained model by using the fine tuning data set, and adapting to a document structure extraction task to obtain a document structure extraction model. By means of the technical scheme, efficient and accurate document structure extraction can be achieved.
Owner:CETC CYBERSPACE SECURITY TECH CO LTD

Text-based image retrieval

A method, apparatus, non-transitory computer readable medium, and system for media processing include obtaining a text prompt describing content, generating, using a multi-modal encoder, a text embedding based on the text prompt, and obtaining an image depicting the content based on the text embedding. The multi-modal encoder is trained to encode image descriptions based on a similarity between a caption of a training image and a paraphrase of the caption.
Owner:ADOBE INC

Multi-modal task fine tuning method based on singular value decomposition enhanced routing function

The invention discloses a singular value decomposition-based multi-modal task fine tuning method for enhancing a routing function, which comprises the following steps of: mapping input language and visual features from a high-dimensional space to a low-rank space by using a PEFT method, performing singular value decomposition on language features in the low-rank space, performing routing function alignment through a tensor after efficient reconstruction, and performing multi-modal task fine tuning on a multi-modal task based on a singular value decomposition-enhanced routing function. And finally, after the low-rank space is recovered to the original dimension again, performing residual connection with the original language features, and outputting the features. According to the method, singular value decomposition is applied to language features before a routing function, a low-rank dominant mode of the language features is extracted, the alignment precision of vision and language features is enhanced, interference of high-dimensional noise is eliminated, and meanwhile calculation efficiency and model stability are kept. Routing calculation is carried out through the reconstructed tensor, key information in the features can be better extracted and aligned, and therefore the precision and effect of feature alignment are improved. The method is suitable for VL tasks such as visual questioning and answering and image description generation, and model performance can be obviously improved.
Owner:GUANGDONG POLYTECHNIC NORMAL UNIV

Knowledge enhancement and emotion inconsistency-based multi-mode siphonage detection method

The invention discloses a knowledge enhancement and sentiment inconsistency-based multi-modal chaffy detection method, which comprises the following steps of: obtaining a to-be-detected sample containing a text and an image, and extracting image and text features and text embedding features by using a feature encoder; and obtaining image description and texts in the image based on the image, splicing the image description and the texts to form knowledge texts, and sequentially inputting the knowledge texts into the emotion dictionary and the feature encoder to obtain emotion features and knowledge embedding features. And carrying out feature interaction among the image, the text and the emotion features by adopting cross attention, and carrying out adaptive weighting on the image and text features through a gating mechanism to obtain image-text comprehensive features and emotion features processed by the cross attention. And inputting the text embedding features and the knowledge embedding features into an emotion inconsistency module, and calculating an emotion inconsistency expression. And based on the image-text comprehensive features, emotion features subjected to cross attention processing and emotion inconsistent representation, performing chiffon prediction, and outputting a detection result. According to the invention, the method achieves better prediction performance on a multi-mode anti-tech detection public data set, and can more accurately recognize an image-text sample with irony emotion.
Owner:SHANTOU UNIV

Image context based text generation

Methods, systems, and storage media for generating contextually relevant text from image descriptions and user intent are disclosed. Exemplary implementations may: receive an image and a user-defined intent for text output; analyze the received image to generate a contextual description of the image; generate a query based on the contextual description of the image and the user-defined intent; and generate the text output based on the query.
Owner:SHUTTERSTOCK

Method for generating image description text based on large model

The invention discloses a method for generating an image description text based on a large model, which relates to the technical field of image processing, and comprises the following steps: an image preprocessing step: dividing a target level through semantic segmentation and extracting key visual information by adopting hybrid denoising and self-adaptive normalization; a feature extraction step: fusing the multi-scale visual features and the semantic features, and generating a high-dimensional fusion feature vector through cross-modal alignment; a large model initialization and adaptation step: loading the pre-training model and performing incremental fine tuning, and dynamically adjusting the Prompt template; a text generation step: generating candidate texts through logic constraint and beam search; and a text optimization adjustment step of outputting a final description text based on multi-dimensional evaluation and user preference iterative correction. According to the method, the semantic matching degree, logic coherence and common sense accuracy of image description are improved, multi-element scenes are adapted through dynamic adaptation and iterative optimization, and high-quality text description support is provided for high-precision and diversified scenes.
Owner:BEIJING LINGMANG TECH CULTURE CO LTD

Wafer defect retrieval method and device and storage medium

The invention provides a wafer defect retrieval method and device and a storage medium, and the method comprises the steps: obtaining a to-be-retrieved wafer image and related text information, the to-be-retrieved wafer image comprises defect features, and the related text information of the to-be-retrieved wafer image comprises an image description text; inputting the to-be-retrieved wafer image and the related text information into a pre-trained multi-modal feature extraction model to obtain multi-modal features of the to-be-retrieved wafer image; and retrieving a preset wafer defect database by using the multi-modal features of the to-be-retrieved wafer image to obtain similar cases of the to-be-retrieved wafer image. By means of the scheme, the accuracy and distinction degree of the detection result can be improved, and the wafer defect retrieval efficiency is improved.
Owner:SKYVERSE TECH CO LTD

Fine-grained image classification method based on large model enhancement

The invention discloses a fine-grained image classification method based on large model enhancement, which guides a model to learn a specific judgment mode of a task by directly introducing prior knowledge. A series of image descriptions are generated by performing question and answer interaction with a multi-modal large language model (MLLM). In order to filter illusion information and redundant content existing in description, a dual-guide text feature optimization module is introduced, and the quality of text features is improved through task-guided feature selection and similarity-guided feature pruning. And finally, adopting a multilayer network structure based on an attention mechanism to realize vision-language fusion for final classification prediction.
Owner:BEIJING UNIV OF TECH

Voxel-level semantic mapping method and device for human visual cortex and electronic equipment

PendingCN121170797AImage analysisSemantic analysisVisual cortexVoxel
The embodiment of the invention provides a voxel-level semantic mapping method for a human visual cortex, and the method comprises the steps: building a corresponding relation between a visual stimulation feature and a human brain voxel response through a visual language model and a coding model; adopting a BLIP image description generation and word segmentation technology to extract a plurality of noun candidate tags from a human brain image sample set highly responsive to target human brain voxels; generating a plurality of image samples consistent with the plurality of noun-type candidate tags in terms of semantics in batches, and inputting the image samples into the coding model for response evaluation; and by sorting the voxel responses of the plurality of image samples, distributing semantic correlation scores for the plurality of noun type candidate tags. According to the voxel-level semantic mapping method for the human visual cortex provided by the embodiment of the invention, the label bias can be reduced by utilizing the visual language model by constructing a voxel-level semantic mapping mechanism running in an open vocabulary space. The embodiment of the invention further provides a voxel-level semantic mapping device and electronic equipment.
Owner:BEIJING INST OF TECH

SAR (Synthetic Aperture Radar) image description method, device and equipment based on multi-modal large model driving

The invention relates to an SAR image description method, device and equipment based on multi-modal large model driving, and the method comprises the steps: constructing a target-relation-scene three-layer collaborative layered labeling strategy according to a multi-layer relation of a target in a scene; constructing a structural cue word comprising a description cue instruction, a category cue instruction and a restriction word cue instruction, inputting the SAR image and the structural cue word into a multi-modal large model, and generating hierarchical semantic description corresponding to each hierarchy in a hierarchical labeling strategy by the model according to the structural cue word, and reconstructing and fusing the semantic description of each level to obtain a preliminary semantic description, performing quality analysis on the preliminary semantic description, optimizing the structural cue word according to a quality analysis result, and correcting the preliminary semantic description according to a semantic description standard to obtain the semantic description. By adopting the method, the semantic description annotation of the SAR image can be realized aiming at the problems that the image structure is complicated, the annotation mode is limited and the data scale is insufficient.
Owner:NAT UNIV OF DEFENSE TECH

Visual language model-oriented medical image text generation method and system

The invention provides a medical image text generation method and system oriented to a visual language model, and belongs to the technical field of medical image processing. The method comprises the following steps: acquiring a labeled image of a medical image; performing connected domain analysis on the annotated image, and detecting and counting the number of tumor areas in the annotated image; for each tumor area, extracting morphological features including size, shape, position, cavity features and edge shape; according to a preset text template, the extracted morphological features are converted into structured natural language description, and final medical image text description is output. It is ensured that the generated medical image description has high consistency and specialty, and subjectivity and difference of manual writing are avoided. It can be ensured that key morphological characteristics (such as tumor size, shape and position) are accurately and completely described, and reliable supervision signals are provided for subsequent model training.
Owner:SHANDONG UNIV

Method and electronic device for distributing image description text

The present disclosure provides a method and an electronic apparatus for distributing image description text. The method including: obtaining video data including multiple images; performing image recognition processing on the multiple images to identify images with important information, the important information indicates objects and / or events in the images; in response to the objects and / or events satisfying predetermined triggering conditions, obtaining predefined keywords corresponding to the objects and / or events in the images; generating image description text indicating the objects and / or events based on the obtained predefined keywords and the images with important information; and distributing the generated image description text to terminal devices.
Owner:TP-LINK SYSTEMS INC

Ancient mural monitoring and repairing method and system and electronic equipment

The invention discloses an ancient mural monitoring and repairing method and system and electronic equipment, and the method comprises the steps: a sensing layer deploys an intelligent camera carrying a lightweight edge intelligent segmentation model, collects a mural image in real time, and carries out the damage segmentation; the network layer adopts a low-power-consumption communication technology to transmit data; the cloud layer runs a generative artificial intelligence repair model to process a complex task; an edge intelligent segmentation model is optimized, the high-frequency feature capture capability is enhanced through an external injection dynamic feature modulation component, the segmentation precision is improved by combining internal fine tuning with residual connection, and the small damage detection sensitivity is improved by adopting a cross entropy and a Dice loss function; performing context sensing repair on small damage by directly utilizing a generative repair model fine-tuned by a parameter efficient fine-tuning framework; for large damage, image description is generated through a visual language understanding model, and a multi-modal large language model is combined to generate semantic prompts to guide a generative repair model to generate repair contents with consistent styles.
Owner:CHINA UNIV OF MINING & TECH +1

Video editing method and device based on multi-modal large model, equipment and medium

The invention relates to the technical field of video editing, solves the problem that how to obtain a video clip is not involved in the prior art, and provides a video editing method, device and equipment based on a multi-modal large model and a medium, and the method comprises the steps: obtaining image description information input by a user and a to-be-edited video; performing semantic segmentation on the to-be-edited video according to the image description information to obtain a plurality of initial segmentation fragments in the to-be-edited video; for each initial boundary frame in each initial segmentation segment, if an adjacent video frame which is adjacent to the initial boundary frame and is not included in the initial segmentation segment exists in the video to be edited, obtaining a target boundary frame according to the adjacent video frame and the initial boundary frame; and determining a target video clip in the video to be edited according to the two target boundary frames in the initial segmentation clip. Dynamic adjustment and optimization of the boundary can be realized, so that the target video clip in the video can be accurately determined.
Owner:NINGBO SIMSHINE INTELLIGENT TECH CO LTD

Auto-regression image generation method, computer equipment and computer program product

The invention relates to an autoregression image generation method, computer equipment and a computer program product. The method comprises the following steps: generating text sequence representation according to an image description text; inputting the text sequence representation into a frequency autoregression generation model, and gradually predicting a plurality of frequency domain sequence representations corresponding to the to-be-generated image according to a sequence from low frequency to high frequency; decoding the frequency domain sequence representation to generate a target image; and the quality of the generated image is improved.
Owner:HANGZHOU HIGH-TECH ZONE (BINJIANG) INSTITUTE OF BLOCKCHAIN & DATA SECURITY

Visual question and answer method based on field adaptive retrieval decision

The invention provides a visual question and answer method based on a domain self-adaptive retrieval decision, which comprises the following steps of: forming an input triple (x, q, d) comprising an image, a question text and an image description text; feature modal extraction and domain identification; generating an explicit reasoning track and a preliminary answer by using a chain reasoning technology CoT; according to a preset decision rule, judging that the preliminary answer is output as a final answer or enters the next step; based on the input triad (x, q, d) and the reasoning track, image retrieval and text retrieval are executed, and an enhanced knowledge set is generated; using the enhanced knowledge set to generate a final answer through a chain reasoning technology CoT, performing credibility verification, and outputting the final answer or a preset unknown identifier according to a verification result; according to the method, through reasoning-driven adaptive retrieval and multi-modal knowledge reordering, efficient utilization and real-time supplement of external knowledge are realized, and the accuracy and robustness of visual questions and answers are effectively improved.
Owner:NANJING UNIV OF POSTS & TELECOMM

A method and system for image editing based on multimodal condition adaptation

This invention provides an image editing method and system based on multimodal condition adaptation. The image editing method includes: a first text vector acquisition step: processing to obtain an image description corresponding to the image, processing the image description with an editor to obtain a first text vector based on the image information; a second text vector acquisition step: processing an editing instruction to obtain a second text vector based on the editing instruction; a fusion feature acquisition step: fusing the first and second text vectors through weights to obtain fusion features; an image latent feature encoding acquisition step; an image editing step: receiving injected conditional information, simultaneously denoising the received image latent features and latent noise, with the received conditional information guiding the iterative denoising process to gradually generate the image desired by the user; and an image restoration step. This invention can improve the stability, controllability, and accuracy of image editing.
Owner:SHENZHEN EMDOOR DIGITAL TECH

Literature image-text pair quality control method and system, storage medium and equipment

The invention relates to the technical field of data processing, and discloses a literature image-text pair quality control method, which comprises the following steps: inputting academic literatures, identifying and extracting related descriptive characters of an image in a text, and generating a preliminary image-text pair; extracting a picture title text, calculating a semantic matching degree between the picture title text and the related descriptive text, filtering the preliminary image-text pair based on a first matching degree threshold, and generating a first candidate image-text pair; generating a semantic vector of the academic literature theme, mapping the title text to generate a picture title vector, calculating a correlation degree between the picture title vector and the semantic vector, filtering the first candidate image-text pair based on a second correlation degree threshold, and generating a second candidate image-text pair; processing the pictures in the second candidate image-text pair, taking the full text as background information, generating a corresponding picture description text, performing semantic similarity comparison on the corresponding picture description text and the related descriptive text, and outputting image-text pair data higher than a third similarity threshold value through the preset third similarity threshold value. The method can improve the control quality of literature image-text pairs.
Owner:TONGFANG KNOWLEDGE DIGITAL PUBLISHING TECH CO LTD

Neural radiation field semantic scene reconstruction method based on multi-modal information guidance

The invention belongs to the technical field of image processing, and particularly provides a neural radiation field semantic scene reconstruction method based on multi-modal information guidance, which comprises the following steps: preprocessing data not containing image description to obtain text description of an image; dividing a scene space through a K-means algorithm, and initializing a NeRF sub-network for each sub-space; and constructing a three-flow feature processing channel, and fusing the three-flow feature processing channel into a multi-modal feature. And training the initialized NeRF sub-network of each sub-space. Self-adaptive joint loss integrating luminosity loss and semantic consistency loss is designed, and different losses are emphasized according to different training periods. And through multi-modal feature fusion and a semantic guidance mechanism, leap-type improvement of the three-dimensional reconstruction quality is realized. A dynamic time-varying regulation and control mechanism is combined with a distributed architecture, so that the training efficiency is greatly improved. A vision-language-geometry three-flow cooperation method is adopted, and a multi-modal information barrier is broken through to realize deep feature fusion.
Owner:LIAONING GENERAL AVIATION ACAD +1

Text generation, model training method and apparatus

The disclosure provides a text generation method and device and a model training method and device, and relates to the technical field of computer vision. The text generation method comprises the following steps: extracting visual features of a to-be-processed image; obtaining related text of the to-be-processed image; encoding the related text of the to-be-processed image to obtain related semantic features of the to-be-processed image; and generating a description text of the to-be-processed image according to the visual features of the to-be-processed image and the related semantic features of the to-be-processed image. Through the above steps, the accuracy of the generated image description text can be improved.
Owner:JINGDONG TECH HLDG CO LTD

Image description generation method and system based on two-stage progressive fusion coding

The invention discloses an image description generation method and system based on two-stage progressive fusion coding, and the method comprises the steps: gradually interpolating features extracted by an image encoder CLIP ViT into features extracted by a corresponding image encoder Swin Transform in a first stage, so as to refine semantic representation; in the second stage, a global perception work space module is provided, and the work space integrates features from an image encoder Swin Transform and an image encoder CLIP ViT through weighted fusion; variable-length input is efficiently processed by adopting a length-independent extension module; the problems of feature representation fragmentation and non-ideal visual language alignment caused by dependence on a single visual encoder in the existing method are solved, and the method has outstanding performance in the aspects of image description generation accuracy and semantic expression richness.
Owner:TIANJIN POLYTECHNIC UNIV +1

Liver lesion image description method based on semantic segmentation network

The application discloses a liver lesion image description method based on a semantic segmentation network, comprising the following steps: building a Unet semantic segmentation network based on a lightweight network GhostNet, and constructing a liver lesion segmentation model; acquiring a historical liver lesion ultrasound image dataset, inputting the liver lesion segmentation model for training, and obtaining an optimal segmentation model; and describing the features of the lesion image based on the segmentation result. The application segments the liver lesion by constructing a lesion segmentation model, and describes based on the segmentation result, thereby improving the accuracy and reliability of liver disease diagnosis, reducing misjudgment and subjective bias; by combining the advantages of deep learning and traditional image processing, the characteristics of both are fully utilized, the performance and stability of the algorithm are improved; automatic processing and analysis of the ultrasound image are realized, the lesion description result provides accurate disease positioning and type judgment, and a better treatment scheme is provided for doctors, thereby improving the treatment effect and treatment experience of patients.
Owner:VINNO TECH (SUZHOU) CO LTD