Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

306 results about "Visual cueing" patented technology

Visual cueing. A visual cue is a signal which your brain extracts from what you see. It indicates the state of some property around you that you are interested in perceiving. Now, only 1% of what you see actually enters through your eyes (the rest is -surprisingly correct – made up by your brain).

Zero sample anomaly detection method and device and electronic equipment

The invention provides a zero sample anomaly detection method and device and electronic equipment, and relates to the technical field of image anomaly detection.The method comprises the steps that a to-be-detected image is acquired, and visual features are extracted through a CLIP model; performing visual enhancement on the visual features to obtain enhanced visual features; injecting the enhanced visual features into a learnable text prompt template to generate an adaptive text prompt; injecting the adaptive text prompt into a text encoder for encoding to obtain a text embedding representation; mapping the adaptive text prompt to a visual space to obtain a visual prompt, and inputting the visual prompt into the local visual features to obtain scale visual features; and based on the text embedding representation and the scale visual features, separately calculating an anomaly score and an anomaly positioning map of a preset scale, and fusing to obtain a detection result. According to the zero sample anomaly detection method and device and the electronic equipment provided by the invention, the generalization ability of the whole detection process is effectively improved, and the actual application requirements are further met.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Steel bridge disease detection and identification method based on large language model

The invention relates to a steel bridge disease detection and identification method based on a large language model, and belongs to the technical field of artificial intelligence and civil engineering crossing. According to the method, a cross-modal feature alignment mechanism is constructed through a pre-trained multi-modal large language model by fusing a steel bridge image and a field customized text prompt, and a cascade detection process of'component identification-disease classification-region segmentation 'is realized. Comprising the following steps: designing a structured text prompt word bank to enhance semantic consistency, and dynamically fusing general knowledge and instance features in combination with a mixed prompt mechanism; a multi-level cross-modal alignment strategy is adopted to generate an anomaly graph, and a disease area is accurately positioned; a visual prompt enhancement module is introduced to improve the multi-scale feature discrimination ability, and the robustness in a complex environment is adjusted and optimized through data self-adaption. Under the condition of few samples or even zero samples, high-sensitivity detection and pixel-level segmentation of steel bridge cracks, corrosion and other diseases are achieved, and the problems that a traditional method is low in efficiency, poor in generalization, high in labor cost and the like are effectively solved.
Owner:HEBEI UNIV OF TECH

Method and system for efficiently marking and segmenting lesion area in medical image

The invention provides a method and system for efficiently marking and segmenting a focus area in a medical image, and the method comprises the steps: S1, collecting a medical image data set, and dividing the medical image data set into a training set and a test set; s2, selecting a reference image from the training set, carrying out initial labeling, generating a visual prompt, and obtaining the reference image with the visual prompt; s3, constructing a deep convolutional neural network architecture, and obtaining a trained model through model training; s4, performing prediction segmentation on the unlabeled image in the test set by using the trained model to generate a prediction segmentation result; and S5, checking and correcting the prediction segmentation result, and improving the model performance through interactive optimization. According to the method, a doctor only needs to carry out initial labeling on a small number of focus areas, the problem that an existing fully-supervised segmentation method seriously depends on a large amount of high-quality labeling data is solved, the requirement for the high-quality labeling data is remarkably reduced, and the workload of medical staff is relieved.
Owner:SUZHOU VOXEL INFORMATION TECH CO LTD

Incorporating non-text cues for machine learning referential dialogue

A computer system, method, and program product facilitate human-computer interaction. A processor set receives a non-text visual cue and natural language instruction regarding a scene. The processor set converts the non-text visual cue into a textual location information indicating a portion of an image representing the scene. A language machine learning model is triggered by using the textual location information, the natural language instruction, and the image representing the scene as input. The language machine learning model outputs a response to the input.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Experiment task execution method and device

The invention provides an experiment task execution method and device.The method comprises the steps that task information of a target experiment task is divided based on a visual language model, a subtask sequence is obtained, and the subtask sequence is formed by arranging multiple subtasks from front to back according to the execution sequence; processing the text information corresponding to the current sub-task and the experiment image before the current sub-task is executed through the visual language model from the first sub-task in the sub-task sequence to obtain a visual prompt image, and executing the visual prompt image through a visual language action model. And processing based on the text information, the experiment image and the visual prompt image, guiding a robot to execute the current sub-task until the last sub-task in the sub-task sequence is completed, and determining that the target experiment task is completed. According to the method, the success rate, the operation safety and the regulation compliance of the experiment task are effectively improved, and the method has high universality and safety.
Owner:BEIJING ACAD OF ARTIFICIAL INTELLLIGENCE

Remote sensing image target statistical method and system fusing large language model and visual cue driving

The invention provides a remote sensing image target statistical method and system fusing a large language model and visual prompt driving. The method comprises the following steps: acquiring a remote sensing instance segmentation image to be processed and a visual prompt thereof; inputting a to-be-processed remote sensing instance segmentation image and a visual prompt thereof into the trained remote sensing image target statistical model, and outputting a remote sensing image target statistical result; the training comprises the following steps: introducing a large language model and visual cue into an encoder architecture of a GrondingDINO model to obtain a remote sensing image target statistical model; inputting a remote sensing instance segmented image and the visual cue thereof into an encoder, and outputting an image feature, a visual cue feature and a text feature; the feature intensifier carries out fusion processing on the output of the encoder; a language-guided query selection module calculates cross-modal query according to the fusion processing result, and a cross-modal decoder obtains a target statistical result of the image based on the fusion processing result and the cross-modal query; and training by using the training data and outputting the trained model.
Owner:WUHAN UNIV

Video understanding method and device based on multi-modal large model, and medium

The invention discloses a video understanding method and device based on a multi-modal large model and a medium, and the method comprises the steps: carrying out the frame-by-frame fusion of visual prompt information corresponding to a predetermined first frame and a video frame sequence through a dynamic Alpha mixing technology; extracting multi-frame features from the fused video frame sequence, and integrating the multi-frame features to generate corresponding space-time visual representation; for each current frame in the video frame sequence, fusing the visual prompt information of the previous frame corresponding to the current frame with the visual features of the current frame, and generating visual prompt information of all frames after the first frame in the video frame sequence; and according to the dynamic Alpha mixed coefficient, performing smooth transition on the visual prompt information of every two adjacent frames in the video frame sequence so as to generate a corresponding video understanding result through an autoregression language model in combination with the text instruction and the space-time visual representation.
Owner:山东浪潮智慧建筑科技有限公司

Multi-mode prompt memory unsupervised continuous anomaly detection method and system

The invention discloses an unsupervised continuous anomaly detection method and system for multi-mode prompt memory, and relates to the technical field of computer vision and industrial image detection. The method comprises the following steps: acquiring industrial images to form a data set; an unsupervised anomaly detection model is constructed, the unsupervised anomaly detection model comprises a visual branch and a text branch which are respectively used for extracting visual features and text features, learnable visual prompts and learnable text prompts are introduced into the unsupervised anomaly detection model, and then the visual features and the text features are fused by using a self-adaptive fusion mechanism to obtain an anomaly detection result; training learnable visual prompts and learnable text prompts in the unsupervised anomaly detection model by using the data set; and performing anomaly detection on the industrial data by using the trained unsupervised anomaly detection model. According to the method, multi-modal information can be fused in the learning process, the continuous learning ability is achieved, and efficient unsupervised anomaly detection of industrial products is achieved.
Owner:SHANDONG COMP SCI CENTNAT SUPERCOMP CENT IN JINAN +2

Passenger car part fault detection method and system based on visual prompt and visual large model

The invention discloses a passenger car component fault detection method and system based on visual prompt and a visual large model, and the method comprises the steps: carrying out the position marking of a target in each sample image for each type of component sample image set, splicing a panoramic image, and generating a mask image; describing the exception type of each type of part and the details of each exception type through the exception description text; extracting multi-scale image features of the panoramic image and the mask image, and fusing the multi-scale image features to obtain visual prompt features; extracting a multi-scale image feature of the to-be-detected part image, fusing the multi-scale image feature with the visual prompt feature to obtain a multi-scale image feature of the to-be-detected part, inputting the multi-scale image feature into the part positioning model, and positioning a target part image; and fusing the visual features of the target part image with the text features of the corresponding abnormal description text by using a visual large model, and predicting and outputting a fault detection result. According to the invention, precise part positioning and fault detection can be realized.
Owner:CENT SOUTH UNIV +1

Light-based visual cues that assist in medication delivery to a patient

One or more light sources may be integrated with a medication delivery device, and states of the light sources may be changed. For example, a first lighting state change may be performed, such as based on a detection that an applied force meets or exceeds a threshold break force or an interaction of a movable component with a proximal sensor integrated with the medication delivery device. The first lighting state change may provide visual feedback indicating that an injection has begun. Additionally, a second lighting state change may be performed, such as based on a detection that a resistive force meets or exceeds a threshold back force or an interaction of the movable component with a distal sensor integrated with the medication delivery device. The second lighting state change may indicate that a complete dosage of the medication has been delivered from the medication delivery device.
Owner:JANSSEN RESEARCH & DEVELOPMENT LLC

Method and device for selectively positioning sound source based on visual prompt, medium and product

The invention provides a method and a device for selectively positioning a sound source based on visual prompt, a medium and a product. The method comprises the following steps: acquiring a mixed audio signal and a prompt image; the mixed audio signal comprises audio signals corresponding to different sound events triggered by at least two different sound sources; the prompt image is associated with a target sound source, and the target sound source is a required sound source selected and positioned from different sound sources; the sound event triggered by the object prompted by the prompt image and the sound event triggered by the target sound source belong to the same sound event; processing the mixed audio signal and the prompt image through a preset cross-instance audio-visual positioning model, including estimating a target mask; and outputting a target arrival direction corresponding to the target sound source based on the target mask. According to the invention, the specific sound source corresponding to the specific sound event is selected and positioned from a plurality of different sound events triggered by a plurality of sound sources, and the application range is wide.
Owner:TRUE SPACE (ZHUHAI) TECH CO LTD

Ultrasonic image segmentation method based on edge guidance

The invention relates to the technical field of medical image processing, in particular to an ultrasonic image segmentation method based on edge guidance, and the method comprises the steps: inputting an ultrasonic image into an edge extraction branch and an image encoder; outputting an edge mask image corresponding to the image through the edge extraction branch; performing noise suppression and topological correction on the edge mask image to generate a closed edge constrained by an anatomical structure; generating prompt box information based on the closed edge, and inputting the prompt box information to a prompt encoder; fusing the image prompt features output by the prompt encoder, the image features extracted by the image encoder and the edge mask information; the fusion features are input into a decoding module, a final segmentation result is obtained, the synergistic effect of edge information and visual prompt is fully utilized, the perception ability of the model for the anatomical structure in the ultrasonic image is improved, and the accuracy and robustness of segmentation are remarkably improved. Therefore, the problems of fuzzy boundary, inaccurate prompt, weak structure identification capability and the like in related technologies are solved.
Owner:WUHAN UNIV

Method, device, and medium for training large scale object foundation model

Embodiments of the present disclosure provide a method, device, and medium for training a large scale object foundation model. The method comprises obtaining a training dataset comprising a plurality of subsets for a plurality of object perception tasks, wherein a sample in the training dataset comprises an image with an object, a prompt indicating the object, and labeled object perception information of the image. The method further comprises generating, by the image encoder, an image feature based on the image. The method further comprises generating, by the text encoder or the visual prompt encoder, a prompt embedding based on the prompt. The method further comprises generating, by the object decoder, object perception information of the object based on the image feature and the prompt embedding. In addition, the method further comprises training the object processing model based on the generated object perception information and the labeled object perception information.
Owner:LEMON INC(GB)

Keycap with an Interchangeable Keycap Top

The present disclosure teaches a keycap with an interchangeable keycap top. The keycap comprises a keycap base and a keycap top. Wherein, a first magnetic piece connects to the keycap base, and a second magnetic piece connects to the keycap top. Both magnetic pieces can be a magnet or of a magnetic material, and at least one of the magnetic pieces is a magnet. The keycap top can thus be secured onto the keycap base by attaching the first magnetic piece to the second magnetic piece. The keycap top may be textured, have a distinctive shape, or have a distinctive coloring, or have other features, to provide physical and visual cues for a user. The present disclosure also teaches a customizable keyboard and a method of customizing such keyboard using said interchangeable keycap tops.
Owner:LEBLANC KYLE +2

Multi-modal face living body detection method and device based on text enhancement

The invention discloses a multi-modal face living body detection method and device based on text enhancement, and the method comprises the steps: inputting a multi-modal face image into an image block embedding module and a visual prompt generation module, and carrying out the encoding through an image encoder, thereby obtaining an image feature and a visual prompt; inputting the dichotomy text description into a text embedding module, randomly initializing a text prompt vector, and encoding through a text encoder to obtain text features and text prompts; respectively enhancing image features and visual prompts by using a hybrid expert module and a bypass prompt enhancement module; performing information exchange on the visual prompt and the text prompt by using a text enhancement module to obtain an enhanced text prompt, and performing information exchange on the image feature and the text feature by using an image mask module to obtain an enhanced image feature; and finally calculating the similarity between the image features and the text features, and taking the category corresponding to the highest similarity as a detection result. According to the invention, the method can effectively enhance the discrimination capability of human face features, and improves the detection accuracy and generalization.
Owner:ZHEJIANG UNIV

Visual field detection equipment, system and method based on expanded reality

The invention relates to the technical field of intelligent visual detection, in particular to visual field detection equipment, system and method based on extended reality (XR), and the method comprises the steps that an image feature extraction module collects image data in real time, extracts features and eliminates optical distortion; the fixation deviation correction module analyzes eye movement data in real time, quantifies fixation deviation and triggers visual prompt and correction; the adaptive stimulation generation module dynamically encrypts to generate a test point image and adaptively adjusts the test point image according to visual field analysis and test requirements; the dynamic adjustment image module predicts and optimizes a test point presentation sequence according to image features and fixation points, automatically encrypts a high-probability defect area and quickly jumps to a key screening area; and the visual field defect identification module extracts image features, calculates the defect probability of each region in real time by using a Bayesian network in combination with subject response data, and locates a high-risk region. The system provides an efficient and accurate visual field detection scheme through the synergistic effect of multiple modules.
Owner:TONGJI HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI TECH +1

Adaptive training method and system based on Tai Chi action recognition

The invention belongs to the technical field of intelligent sports and computer vision crossing. The adaptive training method based on Tai Chi action recognition is provided, and space action features are extracted according to three-dimensional human skeleton key point coordinates of a human body; extracting multi-scale motion features of the human body in a time domain according to the spatial motion features; global attention aggregation is carried out according to the multi-scale motion features of the human body in the time domain, and a human body motion recognition result is obtained; according to a human body action recognition result, matching a corresponding Tai Chi style drawing theory; superposing multi-dimensional visual prompts into a Tai Chi motion video according to the Tai Chi motion style theory so as to intuitively prompt a user about action key points; calculating an action problem based on priori knowledge or a preset template, prompting an error position and an error reason by adopting graphs and / or characters, and giving an adjustment suggestion; according to the invention, more personalized training experience can be provided, cognitive load can be reduced, and training efficiency, training effect and cognitive learning fluency can be improved.
Owner:SHANDONG UNIV

Remote sensing target segmentation method, electronic equipment, storage medium and program product

The embodiment of the invention provides a remote sensing target segmentation method, electronic equipment, a storage medium and a program product. The method comprises the following steps: acquiring a to-be-segmented remote sensing image and a sample library; calculating the feature similarity between the to-be-segmented target semantic tag and each sample, and screening out a target sample of which the feature similarity meets a preset requirement from the sample library; generating a target reference mask based on the target sample; and converting the target reference mask into a visual prompt, and inputting the visual prompt and the to-be-segmented target semantic tag into a segmentation model to obtain a segmentation result. Through combination of sample library assistance and visual prompt guidance, accurate segmentation of a specific target in a remote sensing image is realized, and dependence on dense manual annotation is eliminated. Meanwhile, the generalization ability of the model to the same kind of targets in different scenes is enhanced through double constraints of semantic tags and visual features.
Owner:ZHONGKEHONGYUN TECH (BEIJING) CO LTD

Video action recognition model training method, video action recognition method and device

The invention relates to a video action recognition model training method, a video action recognition method and a video action recognition device. The method comprises the following steps: acquiring a sample video frame image and an action category description text corresponding to the sample video frame image, inputting the sample video frame image and the action category description text into a to-be-trained recognition model, and recognizing the action category of the to-be-trained recognition model by an image encoder in the recognition model according to a preset visual prompt vector and the sample video frame image, generating a video embedding corresponding to a sample video frame image, generating a text embedding corresponding to an action category description text by a text encoder in the recognition model based on a preset text prompt vector and the action category description text, and constructing bidirectional comparison loss by taking the video embedding and the text embedding as positive sample pairs, and updating the visual prompt vector and the text prompt vector based on the bidirectional contrast loss to obtain a trained recognition model. By adopting the method, the video action recognition accuracy can be improved.
Owner:CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1

High-performance unmanned aerial vehicle detection method based on visual cue words

The invention discloses a high-performance unmanned aerial vehicle detection method based on a visual cue word, and relates to the technical field of unmanned aerial vehicle detection, and the method comprises the steps: collecting diversified images containing a target to form a training data set, collecting a few-sample cross-domain image as cross-domain information input, inputting the training image into a pre-trained and parameter-locked backbone network, and extracting basic features; a few-sample cross-domain image is input into a cross-domain network, after features are extracted, a learnable visual prompt vector is generated, noise is added, then the prompt vector and basic features are input into a neck module through a cross attention mechanism to be integrated to obtain fusion features, and finally a detection head is input to output a classification and position regression result of a target. By reducing the training cost and improving the cross-domain generalization performance of an unmanned aerial vehicle algorithm, an efficient solution is provided for low-resource cross-domain detection tasks such as industrial quality inspection and geographical remote sensing, and meanwhile, a solution is provided for training a high-performance unmanned aerial vehicle through extremely few training parameters.
Owner:ZHEJIANG NORMAL UNIV +3

Visual language model training method, image tag prediction method, and electronic device

PCT designated stageWO2026026225A1Biological modelsFeature extractionLinguistic model
Embodiments of the present disclosure provide a visual language model training method, an image tag prediction method, and an electronic device. According to the embodiments of the present disclosure, encoding parameters respectively corresponding to an image encoder and a text encoder in a first visual language model are fixed, on the basis of initial values of visual prompts and text samples respectively corresponding to a plurality of image classification tags, a second visual language model is generated by training the first visual language model, and the visual prompts respectively corresponding to the plurality of image classification tags are mined and learned, thereby enhancing the visual representation capabilities of visual language models. Next, on the basis of the learned visual prompts, text prompts, and the text samples, a target visual language model is further trained by training the second visual language model, and adapter parameters of model adapters respectively configured for the image encoder and the text encoder are adjusted, to collaboratively optimize image feature extraction and text feature extraction. This process allows for the transfer of visual knowledge from visual prompts to text prompts.
Owner:HANGZHOU ALIBABA INT INTERNET IND CO LTD

Image anomaly detection method and device, equipment and storage medium

The embodiment of the invention provides an image anomaly detection method and device, equipment and a storage medium, and relates to the technical field of computer vision. The method comprises the following steps: inputting an abnormal detection image and a depth visual prompt feature into an image encoder to generate a classification vector and a plurality of block vectors, inputting a plurality of prompt texts into a text encoder to obtain a text vector, and calculating the similarity of the classification vector and the plurality of text vectors to obtain a global similarity score; and calculating the similarity of the block vector and the plurality of text vectors to obtain a block similarity score, further obtaining a similarity mapping result, and obtaining an anomaly detection result according to the global similarity score and the similarity mapping result. Depth information of an image, a spatial position and a structure of an object and other contents are provided by means of depth visual prompt features, so that the capability of identifying an anomaly of an unknown category is improved. The block vector is used to describe the image carefully from the angle of local pixels, so that the accuracy of anomaly recognition is remarkably improved.
Owner:PENG CHENG LAB

Method and apparatus for learning a visual prompt on a multimodal large language model

A computer implemented method for learning a visual prompt on a Multimodal Large Language Model (MLLM) for a downstream task, comprising: applying a visual prompt to a plain image, wherein the visual prompt is a set of parameters in a pixel space; obtaining a first feature embedding of the prompted image and a second feature embedding of the plain image; outputting a textual prediction corresponding to the prompted image for the downstream task by the MLLM based on a projection of the first feature embedding to a text embedding space; and optimizing the visual prompt, with parameters of the MLLM fixed, by minimizing a loss function constructed based at least on a relative entropy loss between the first feature embedding and the second feature embedding. Several other aspects are disclosed.
Owner:ROBERT BOSCH GMBH +1

System for identifying a battery based on a color of a part of the battery

Battery life cycle management is facilited with distinct visual cues of colors and color-coded combinations. The color-codes enable battery types and characteristics over the course of the life of the battery to be determined. The color-codes are based on a color system that indicates a performance rating of the batteries. A system for managing the color-coded batteries selects color-coded batteries for use based upon color, maintaining color based functionality, and color-coded battery indication for destruction of the battery at the end of its life cycle.
Owner:CPS TECHNOLOGY HOLDINGS LLC

Visual prompt multi-modal large model for multi-source remote sensing image interpretation

The invention discloses a visual prompt multi-modal large model for multi-source remote sensing image interpretation, which belongs to the technical field of crossing of remote sensing and computer vision and comprises a multi-modal content coding and integration module, a cross-domain first-stage fusion training module, a pixel level visual positioning module and a large language model. The model supports the interpretation of a remote sensing image after the remote sensing image is arbitrarily amplified and reduced, and has flexible multi-granularity vision and language interaction capability. According to the model, a large language model is used as an interface, and multi-modal content integration including multi-sensor images, visual prompts and text instructions is achieved. In addition, two types of space tasks of anaphora understanding and visual positioning are unified into a visual prompt learning framework, and comprehensive and flexible multi-granularity understanding of remote sensing data is promoted.
Owner:BEIJING INST OF TECH

Mapping characteristics of music into a visual display

A method and system for visualizing music using a perceptually conformal mapping system are provided. A music source file is input into a processor configured to carry out a series of steps on audio cues identified within the music and ultimately generate a simultaneous visual representation on a display device. The series of steps include application of one or more perceptually conformal mapping systems that essentially induce a synesthetic experience in which a person can experience music both acoustically and visually at the same time. The device extracts cues from the music that are designed to specifically capture fundamentals of human appreciation, maps them into visual cues, then presents those visual cues synchronized with the source music.
Owner:NEW RESONANCE LLC

Long video generation method, device and equipment, readable storage medium and program product

The invention discloses a long video generation method and device, equipment, a readable storage medium and a program product, relates to the technical field of communication, and aims to improve the quality of a generated long video. The method comprises: obtaining a video description text to be processed and video processing parameters, the video processing parameters comprising a video segment number N and a video frame number M included in each video segment, N being an integer greater than or equal to 2, and M being an integer greater than or equal to 1; dividing the video description text into text information corresponding to N time periods; taking the text information corresponding to each time period and the historical video visual prompt word corresponding to each time period as input of a video visual encoder, and running the video visual encoder to obtain a video clip corresponding to each time period; and splicing the video clips to obtain a long video. According to the embodiment of the invention, the quality of the generated long video can be improved.
Owner:CHINA MOBILE COMM LTD RES INST +1

Interpreting summarization model decisions based on attention

The disclosure herein describes interpreting attention-based decisions of summarization outputs generated by a deep learning model. A decision interpretation model obtains attention values defining connections between input tokens associated with a source text and output tokens for a selected portion of a summary associated with the source text. The input tokens having the highest attention values indicating the strongest connections between the input tokens of the source text and an output token of the summary are selected as primary tokens. A semantic similarity between the primary tokens for each attention head and an output token is calculated. The model selects the primary tokens having the closest semantic similarity with the summary portion. A visual cue is generated on or within a portion of the source text corresponding to the primary tokens. The visual cue identifies dominant words in the source text used to explain the summary portion.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Visual language large model perception enhancement method and system based on visual prompt

The invention discloses a visual language large model perception enhancement method and system based on visual prompt, and relates to the field of visual language large models. The problem that how to deploy a small-scale large-language model under the condition that resources are limited in the prior art becomes an urgent problem to be solved is solved. The method comprises the following steps of: segmenting an original image by adopting a segmentation component to generate a mask and an object segmentation list; respectively processing an original image and an image mask generated by a semantic segmentation device by using a visual encoder, so as to extract multi-level visual features which highlight the position and boundary of an object; layer normalization and MLP layer processing are carried out to form visual features; taking the generated mask and the segmentation result list of the object as a text instruction, and inputting the extracted multi-level visual features which highlight the position and the boundary of the object and the visual features into a visual language large model to perform autoregressive semantic generation; the method is also suitable for the technical field of improving the object perception and question-answering ability of the visual language large model without adding extra training parameters.
Owner:HARBIN INST OF TECH