Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

165 results about "Image description" patented technology

An image description is a textual, audio or graphical content portraying the image in a representation intelligible by the addressees. The description should be comprehensive as well as perceptible by the target audience.

Unified streaming processing method, system and device for multi-mode AI interactive content, medium and program product

The invention discloses a unified streaming processing method, system and device for multi-modal AI interactive content, a medium and a program product, and the method comprises the steps: receiving a request which is sent by a user and comprises an image and a text, and constructing a request context; extracting image features; carrying out image analysis and outputting an image description text; performing intention recognition and outputting an intention; multi-stage reasoning is carried out, and thinking content fragments are generated in a streaming mode; buffering and releasing thinking content fragments; generating text content fragments based on multi-stage reasoning; buffering and releasing the text content fragments; performing recommendation triggering based on the image description text and the intention; recommending content fragments; inserting a recommended position, buffering and issuing; and process event management: issuing event notifications when the beginning, any step fails and the end. The method is a universal method capable of transmitting different types of outputs in a single ordered stream, the analysis cost can be reduced, and the interaction experience and expansibility are improved.
Owner:BEIJING DIANFU TECHNOLOGY CO LTD

Document structure extraction and model training method and device, equipment and medium

The invention discloses a document structure extraction and model training method and device, equipment and a medium, and relates to the technical field of artificial intelligence and computer vision. The method comprises the following steps: constructing a special training data set containing data of at least two document understanding tasks (including optical character recognition, layout analysis, text positioning, regional text extraction, image description and chart title generation), and a fine tuning data set for converting a document image into a machine-readable structured text format; constructing a multi-modal large model comprising a shape adaptive cutting module, a visual encoder, a visual token compression module, a modal connector and a language decoder; pre-training the model by using the special training data set to jointly learn various document understanding tasks; and performing fine tuning on the pre-trained model by using the fine tuning data set, and adapting to a document structure extraction task to obtain a document structure extraction model. By means of the technical scheme, efficient and accurate document structure extraction can be achieved.
Owner:CETC CYBERSPACE SECURITY TECH CO LTD

Text-based image retrieval

A method, apparatus, non-transitory computer readable medium, and system for media processing include obtaining a text prompt describing content, generating, using a multi-modal encoder, a text embedding based on the text prompt, and obtaining an image depicting the content based on the text embedding. The multi-modal encoder is trained to encode image descriptions based on a similarity between a caption of a training image and a paraphrase of the caption.
Owner:ADOBE INC

Knowledge enhancement and emotion inconsistency-based multi-mode siphonage detection method

The invention discloses a knowledge enhancement and sentiment inconsistency-based multi-modal chaffy detection method, which comprises the following steps of: obtaining a to-be-detected sample containing a text and an image, and extracting image and text features and text embedding features by using a feature encoder; and obtaining image description and texts in the image based on the image, splicing the image description and the texts to form knowledge texts, and sequentially inputting the knowledge texts into the emotion dictionary and the feature encoder to obtain emotion features and knowledge embedding features. And carrying out feature interaction among the image, the text and the emotion features by adopting cross attention, and carrying out adaptive weighting on the image and text features through a gating mechanism to obtain image-text comprehensive features and emotion features processed by the cross attention. And inputting the text embedding features and the knowledge embedding features into an emotion inconsistency module, and calculating an emotion inconsistency expression. And based on the image-text comprehensive features, emotion features subjected to cross attention processing and emotion inconsistent representation, performing chiffon prediction, and outputting a detection result. According to the invention, the method achieves better prediction performance on a multi-mode anti-tech detection public data set, and can more accurately recognize an image-text sample with irony emotion.
Owner:SHANTOU UNIV

Fine-grained image classification method based on large model enhancement

The invention discloses a fine-grained image classification method based on large model enhancement, which guides a model to learn a specific judgment mode of a task by directly introducing prior knowledge. A series of image descriptions are generated by performing question and answer interaction with a multi-modal large language model (MLLM). In order to filter illusion information and redundant content existing in description, a dual-guide text feature optimization module is introduced, and the quality of text features is improved through task-guided feature selection and similarity-guided feature pruning. And finally, adopting a multilayer network structure based on an attention mechanism to realize vision-language fusion for final classification prediction.
Owner:BEIJING UNIV OF TECH

Visual question and answer method based on field adaptive retrieval decision

The invention provides a visual question and answer method based on a domain self-adaptive retrieval decision, which comprises the following steps of: forming an input triple (x, q, d) comprising an image, a question text and an image description text; feature modal extraction and domain identification; generating an explicit reasoning track and a preliminary answer by using a chain reasoning technology CoT; according to a preset decision rule, judging that the preliminary answer is output as a final answer or enters the next step; based on the input triad (x, q, d) and the reasoning track, image retrieval and text retrieval are executed, and an enhanced knowledge set is generated; using the enhanced knowledge set to generate a final answer through a chain reasoning technology CoT, performing credibility verification, and outputting the final answer or a preset unknown identifier according to a verification result; according to the method, through reasoning-driven adaptive retrieval and multi-modal knowledge reordering, efficient utilization and real-time supplement of external knowledge are realized, and the accuracy and robustness of visual questions and answers are effectively improved.
Owner:NANJING UNIV OF POSTS & TELECOMM

Tunnel face image descriptive model construction method, system, device and medium

The application discloses a tunnel face image descriptive model construction method, system, device and medium, relates to the fields of tunnel engineering and computer vision technology, and comprises the following steps: acquiring tunnel face geological sketches of each mileage and tunnel face images corresponding to the mileage; extracting geological descriptions from the tunnel face geological sketches and performing standardization processing; extracting image high-dimensional features from the preprocessed tunnel face images; constructing training samples by using the image high-dimensional features and the geological descriptions corresponding to the mileage and subjected to the standardization processing; training a preset initial model by using the training samples; and determining the initial model reaching a preset training end condition as a Bert tunnel face image description model.
Owner:中铁科学研究院集团有限公司

Training method for image generation model, and related apparatus and medium

Provided in the present disclosure are a training method for an image generation model, and a related apparatus and a medium. The method comprises: acquiring a plurality of image-text sample pairs, wherein each image-text sample pair comprises a background template image, a noise reference image, image description information of the noise reference image, and a sample object image comprising a sample object; on the basis of the noise reference image, the background template image and a template mask image, determining sample splicing image features; on the basis of contour features of a reference object in the background template image, the image description information, and the sample object image, determining denoising network control information; on the basis of the sample splicing image features and the denoising network control information, performing noise prediction by means of an image generation model, so as to obtain a noise prediction result corresponding to the noise reference image; and on the basis of comparison results between the noise reference images in the plurality of image-text sample pairs and the noise prediction results respectively corresponding thereto, training the image generation model. The present disclosure can improve the accuracy of generating a target image.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Image area content description method and device, equipment and medium

The invention relates to the technical field of visual languages, in particular to an image area content description method and device, equipment and a medium. The method comprises the following steps: acquiring a to-be-described image and an image region specified by description of the to-be-described image, and generating a region mask matrix by using an attention module of a visual encoder in a trained visual language model and combining the image region; when feature coding is carried out on an image to be described, attention weighting is carried out on features of an image region based on the region mask matrix to obtain an image feature vector; and a text encoder in a trained visual language model is used to describe the image feature vector to obtain a content description text, and the trained visual language model is obtained by loss training constructed based on comparison of the difference between the total image description and the regional cut image description. The feature in the region is endowed with a higher weight through the region attention mask to obtain a focused image feature vector, so that the focusing capability of the model is optimized, the global interference is inhibited, and the description accuracy is improved.
Owner:SHENZHEN INTELLIFUSION TECHNOLOGIES CO LTD

Dynamic mapping driven prototyping image description system

This invention discloses a dynamically mapped prototype acquisition image description system, relating to the fields of artificial intelligence and digital image processing technology. The system includes: a visual encoding module: extracting spatial and global features of the input image, encoding the spatial and global features to obtain target features; a fine-grained alignment module: clustering the target features to obtain several semantic clusters and pseudo-labels for the target features; enhancing the target features within each semantic cluster based on the pseudo-labels to obtain enhanced features; constructing a semantic prototype library containing several semantic prototypes; and obtaining the centroid of the semantic clusters as the initial value of the semantic prototypes, aligning the enhanced features to the semantic prototypes based on the initial values; and a decoding and generation module: decoding the aligned enhanced features based on the decoder of the caption generation model to obtain image words, and obtaining image captions based on the image words, thus solving the problem of the lack of explicit semantic association between visual representation and language generation in existing image caption generation systems.
Owner:CHENGDU UNIV OF INFORMATION TECH

Time-series image description method for dam defects based on local self-attention

A time-series image description method for dam defects based on local self-attention mechanism is provided, including: performing frame sampling on an input time-series image of dam defect, extracting a feature sequence using a convolutional neural network and using the sequence as an input to a self-attention encoder, where the encoder includes a Transformer network based on a variable self-attention mechanism that dynamically establishes contextual feature relations for each frame; generating description text using a long short term memory (LSTM) network based on a local attention mechanism to enable each word predicted to be feature related to an image frame, improving text generation accuracy by establishing a contextual dependency between image and text. A dynamic mechanism is added to the present application for calculating the global self-attention of image frames, and LSTM networks with added local attention directly establish the correspondence between image and text modal data.
Owner:HUANENG LANCANG RIVER HYDROPOWER CO LTD +1

Image description generation method and system based on two-stage progressive fusion coding

The application discloses an image description generation method and system based on two-stage progressive fusion coding, and the method comprises the following steps: in the first stage, the features extracted by the image encoder CLIP ViT are gradually interpolated into the features extracted by the corresponding image encoder Swin Transformer to refine the semantic representation; in the second stage, a global perception workspace module is proposed, the workspace integrates the features from the image encoder Swin Transformer and the image encoder CLIP ViT through weighted fusion; and a length-independent expansion module is used to efficiently process the variable-length input; the problems of feature representation fragmentation and unsatisfactory visual language alignment caused by the dependence of the existing method on a single visual encoder are solved, and the method has outstanding performance in the accuracy of image description generation and the richness of semantic expression.
Owner:TIANJIN POLYTECHNIC UNIV +1

Image description generation method based on fine-grained attribute learning and gated attention network

The invention relates to the technical field of minority image description generation, and discloses an image description generation method based on fine-grained attribute learning and a gated attention network, and the method comprises the steps: carrying out the salient region detection through a MaskR-CNN, extracting a region visual feature through a BLIP model, enabling the visual feature to be aligned with an attribute vocabulary through feature matching, and carrying out the recognition of the feature. An attribute-visual feature representation is generated. According to the method, the problem of scarcity of data in the field is solved by constructing a special ethnic costume data set; a fine-grained attribute learning module is provided, modeling is explicitly carried out through feature matching, local visual features and semantic attributes are associated, and the accuracy of detail recognition and attribute classification is remarkably improved; the designed Transform-LSTM-gated attention network can effectively fuse visual features based on attributes and context text prior, and dynamic utilization of cultural semantic information is realized, so that accurate description rich in cultural details is generated.
Owner:YUNNAN UNIV

An incremental training method for an object detection model, a server, and an object detection system.

This application discloses an incremental training method, server, and object detection system for an object detection model. The method includes: training an original OVD visual backbone network, an OVD predefined detection subnetwork, and a language model using a large image description dataset to obtain frozen first network parameters of the OVD visual backbone network; training the original OVD visual backbone network and the incremental subnetwork based on a pre-trained dataset to obtain an updated OVD visual backbone network and a trained incremental subnetwork, using the parameters of the trained incremental subnetwork as second network parameters; loading the frozen first network parameters of the original OVD visual backbone network and the second network parameters of the incremental subnetwork, and fine-tuning the incremental subnetwork using gradient backpropagation based on user incremental data to obtain third network parameters of the incremental subnetwork; and generating a fusion network of the OVD predefined detection subnetwork and the incremental subnetwork corresponding to the third network parameters as an object detection model.
Owner:HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD

A remote sensing image description evaluation method based on cycle consistency

This invention discloses a remote sensing image description evaluation method based on cycle consistency, comprising: generating fine-grained text descriptions of the original remote sensing image using a remote sensing multimodal large model; inputting the generated long text descriptions into a text-to-image model for image reconstruction, and introducing a weighted basis image strategy to compensate for information loss during the text-to-image conversion process; using a multimodal quality assessment model to compare and score the cycle consistency of the original image and the reconstructed image from dimensions such as information sufficiency, scene consistency, and spatial layout, and performing quantitative evaluation according to a preset six-level alignment standard, followed by normalization processing to obtain the final description quality evaluation result. This method solves the problem that existing remote sensing image description evaluation indicators fail in long text and fine-grained scenarios, achieving objective and reliable automatic evaluation without relying on reference text. It is suitable for applications such as remote sensing multimodal model performance evaluation, large-scale dataset quality screening, and intelligent geographic information perception.
Owner:HOHAI UNIV

Multimedia resource storage method and device, equipment, storage medium and program product

The application discloses a multimedia resource storage method and device, equipment, storage medium and program product, and belongs to the technical field of data storage. The method comprises the following steps: displaying a multimedia resource, wherein the multimedia resource comprises at least one image; for each first image in the multimedia resource, dividing the image content of the first image to obtain a first element and a second element, the second element comprises at least part of the image content of the first image, and the first element comprises the image content of the first image except the second element; generating an image description text according to the image content of the second element; and for each first image in the multimedia resource, storing the pixel information corresponding to the image description text and the image content of the first element.
Owner:VIVO MOBILE COMM CO LTD

Medical image report generation method and device and storage medium

The invention discloses a medical image report generation method and device and a storage medium. The method comprises the following steps: acquiring a target image examination part corresponding to a target examination object and a first image description text; and if it is detected that the first image description text has a writing error, correcting the first image description text based on a reference correction rule matched with the image description text in a correction rule library, and generating a second image description text. And identifying a plurality of medical entity words contained in the second image description text, and classifying the plurality of medical entity words to obtain classification tags corresponding to the plurality of medical entity words. According to the method, multiple pieces of structured data used for representing lesion features are generated based on the classification labels corresponding to the multiple medical entity words, and the multiple pieces of structured data are filled into the medical image report template corresponding to the target image examination part, so that the medical terms in the image description text can be more accurately filled into the template; and a high-quality image report is obtained.
Owner:BEIJING FRIENDSHIP HOSPITAL CAPITAL MEDICAL UNIV +1

Multi-granularity evidence enhanced multi-modal semantic understanding and aspect-level sentiment analysis method

The invention discloses a multi-modal semantic comprehension and aspect-level sentiment analysis method based on multi-granularity evidence enhancement. The method comprises the following steps: acquiring multi-modal input data including texts, images and target aspects; according to the multi-modal input data, utilizing a pre-trained visual language model to generate a picture description text related to the target aspect; based on the multi-modal input data, the picture description text and an emotion labeling label of the data, utilizing the visual language model to obtain global emotion evidence; based on the multi-modal input data, utilizing the visual language model to identify a key attribute set related to the target aspect; on the basis of the key attribute set, the emotion labeling label of the data and the picture description text, utilizing the visual language model to obtain fine-grained emotion evidence; and according to the text, the picture description text, the global emotion evidence and the fine-grained emotion evidence, performing emotion polarity prediction on the target aspect to obtain an emotion value.
Owner:GUANGDONG UNIV OF TECH

Alloy performance prediction method fusing vision-language-process multi-modal data

An alloy performance prediction method fusing vision-language-process multi-modal data comprises the steps that a multi-modal data set of an alloy material is constructed, and the multi-modal data set comprises three kinds of modal input including process parameters such as temperature and time of solid solution and aging treatment, SEM image visual information and SEM image description text information and corresponding alloy mechanical property true values; training a visual encoder ResNet50 model through comparative learning and training a language encoder BERT model through mask language modeling to obtain an SEM image and vector codes of language description of the SEM image; splicing and fusing process parameters such as temperature and time of solid solution and aging treatment and the vector codes obtained in the second step, and training a random forest regression device according to corresponding alloy mechanical properties; and the random forest regression device obtained through training is used for alloy performance prediction. According to the method, the structured process data, the unstructured SEM image data and the derived text description data are subjected to collaborative fusion and joint modeling for the first time, complementarity among different modal data is fully utilized, and more comprehensive and more three-dimensional digital representation of the alloy state is constructed; the multi-modal fusion framework and the feature learning mechanism are suitable for wide material systems. Meanwhile, the adopted'depth representation + random forest 'hybrid model has the advantage of high training efficiency while ensuring high prediction precision.
Owner:ZHEJIANG UNIV

Image processing method and device, readable medium and electronic equipment

The present disclosure relates to an image processing method, device, readable medium and electronic equipment, the image processing method comprises: obtaining target detection data, the target detection data comprising a to-be-detected image and image description information of the to-be-detected image, the image description information being specified position information representing a generated region in the to-be-detected image, or being a description text representing an object in the to-be-detected image; inputting the target detection data into a target image processing model to obtain an image processing result output by the target image processing model, wherein the target image processing model is configured to output a representation text representing the generated region corresponding to an image in the case that the image description information is the specified position information, and output a target position of the representation object in the to-be-detected image in the case that the image description information is the description text.
Owner:BEIJING YOUZHUJU NETWORK TECH CO LTD

Interactive graphical user interface for training digital human figures in electronic devices

1. Name of the product in this design: Interactive graphical user interface for training digital human figures in electronic devices. 2. Intended use of this design: electronic equipment. 3. The key design feature of this product is its graphical user interface. 4. The image or photo that best illustrates the design points: Design 2 Interface Change State Diagram 1. 5. Design 2 is designated as the basic design. 6. Purpose of Graphical User Interface: Graphical user interfaces are used for training digital human figures. 7. Human-computer interaction methods of the graphical user interface: Design 1's main view is the homepage of the digital human character training interaction; Design 1's interface state change diagram 1 shows the interface displayed after clicking "Character Training" in the main view of Design 1; Design 1's interface state change diagram 2 shows the interface displayed after uploading a video by clicking any position in the "Video Upload" area in Design 1's interface state change diagram 1; Design 1's interface state change diagram 3 shows the interface displayed after clicking "Confirm" in Design 1's interface state change diagram 2; Design 2's main view is the homepage of the digital human character training interaction; Design 2's interface state change diagram 1 shows the interface displayed after clicking "Character Training" in the main view of Design 2; Design 2's interface state change diagram 2 shows the interface displayed after clicking "Image Upload" in Design 2's interface state change diagram 1; Design 2's interface state change diagram 3 shows the interface displayed after uploading an image by clicking any position in the "Image Upload" area in Design 2's interface state change diagram 2; Design 2's interface state change diagram 4 shows the interface displayed after "Character Character Description" in Design 2's interface state change diagram 3. The interface displayed after entering the image description text in the input box below; Design 2 interface change state diagram 5 is the interface displayed after clicking "Confirm" in Design 2 interface change state diagram 4; Design 3 main view is the homepage of digital human image training interaction; Design 3 interface change state diagram 1 is the interface displayed after clicking "Image Training" in the Design 3 main view; Design 3 interface change state diagram 2 is the interface displayed after clicking "Create at Will" in Design 3 interface change state diagram 1; Design 3 interface change state diagram 3 is the interface displayed after clicking any position in the upload area of ​​"Appearance Reference", "Clothing Reference", "Scene Reference" and "More" in Design 3 interface change state diagram 2 and uploading the corresponding content; Design 3 interface change state diagram 4 is the interface displayed after entering the description text in the input box below "Character Image Description" in Design 3 interface change state diagram 3 and selecting "Cool and Mature Woman" below "Reference Tips"; Design 3 interface change state diagram 5 is the interface displayed after clicking "Confirm" in Design 3 interface change state diagram 4.
Owner:LIANGSHENG DIGITAL TECHNOLOGY (SHENZHEN) CO LTD

Image generation model training method and device and image generation method and device

The embodiment of the invention provides an image generation model training method and device and an image generation method and device.The image generation model training method comprises the steps that multiple page images of a target application program and image description texts of all the page images are obtained, and a style description text corresponding to the target application program is determined; constructing a training data pair according to each page image, the image description text corresponding to each page image and the style description text; and inputting the training data pair into the initial image generation model to obtain a prediction image, and training the initial image generation model according to the prediction image and the page image in the training data pair to obtain a target image generation model. The image can be generated directly according to the image description text and the style description text on the basis of the target image generation model subsequently, the time cost of manual design is greatly reduced, a large number of potential design combinations can be explored on the basis of the model, and therefore efficient, intelligent and innovative image generation is achieved.
Owner:BEIJING KANYUN SOFTWARE CO LTD

An end-to-end image description generation method based on enhanced attention mechanism

The application provides an end-to-end image description generation method based on an enhanced attention mechanism, and belongs to the technical field of artificial intelligence. An image description generation model is generated, which comprises an image feature extraction layer, a multi-granularity feature fusion encoder, a self-adaptive bidirectional graph decoding device, a linear transformation layer and a scoring and sorting layer; the image description generation model is trained using a cross-entropy loss, then a self-criticism training is used to optimize a CIDEr score, and the trained image description generation model is used to describe an image. The evaluation index is superior to the prior art, the image description method has improved image semantic understanding ability, is closer to human description habits, and has good interpretability.
Owner:CHINA UNIV OF MINING & TECH

Multi-mode large model ground penetrating radar disease data intelligent interpretation method and device

The invention discloses a multi-modal large model ground penetrating radar disease data intelligent interpretation method and device, and belongs to the computing intelligence and information processing technology. The method is based on a computing device of machine learning, artificial intelligence and a specific mathematical model, and an inference model of specific ground penetrating radar knowledge is fused for information processing. Comprising the following steps: acquiring underground disease data by using multi-frequency ground penetrating radar equipment, and converting the underground disease data into a serialization format adaptive to a multi-modal large model; and then sequentially executing detection, segmentation and image description tasks. Wherein a segmented label mask is generated by adopting a Bayesian optimization method based on a target detection result, and an image description data set is constructed by utilizing a segmentation result; and finally, performing intelligent interpretation on the data by using the multi-modal large model to generate an interpretation report. According to the method, through multi-task collaborative learning, full-process automatic interpretation from disease detection to semantic description is realized, the interpretation efficiency and reliability are improved, and an intelligent basis is provided for road maintenance decision making.
Owner:CHANGAN UNIV

Open question answering and training method and device of multi-modal large model and related equipment

ActiveCN117235232BAbility to detect the spatial position of objectsDigital data information retrievalInternal combustion piston enginesSemantic alignmentTraining phase
The application discloses a kind of open question and the training method, device and related equipment of multimodal large model, to promote multimodal large model to focus on spatial information, matched image description text with spatial information is generated for training image in pre-training phase, spatial information is used to indicate the spatial position of object contained in training image in training image, and the pre-training of multimodal large model is carried out using training image and the above-mentioned image description text with explicit object spatial information, can make that multimodal large model is further focused on the spatial position of object in image on the basis of learning the semantic alignment relationship of image and content description text, that is, make multimodal large model have the ability of detecting object spatial position.On this basis, when multimodal large model is applied to open question answering task, it can accurately give correct answer based on the ability mastered when answering questions related to spatial arrangement.
Owner:IFLYTEK CO LTD

Small language graphic data set construction method, device and medium for internet data

This invention relates to a method, device, and medium for constructing a minority language image-text dataset for internet data, comprising: acquiring an HTML webpage file; extracting a multimodal document containing minority language image-text data from the file; filtering image-text related pairs based on replaceable text corresponding to images in the multimodal document; inputting the plain images (excluding replaceable text) and the text information after removing the images, along with custom prompts, from the multimodal document into a first multimodal large model to generate corresponding image description information and image-text question-and-answer information; combining the plain images and image description information into image-text description pairs; and combining the plain images and image-text question-and-answer information into visual question-and-answer pairs; inputting the multimodal document, image-text related pairs, image-text description pairs, and visual question-and-answer pairs into a second multimodal large model to output image scene labels; and combining the above information to obtain a minority language image-text dataset. Compared with existing technologies, this invention can achieve the construction of high-quality, diverse, and standardized minority language image-text datasets.
Owner:SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT