Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

146 results about "Visual cueing" patented technology

Visual cueing. A visual cue is a signal which your brain extracts from what you see. It indicates the state of some property around you that you are interested in perceiving. Now, only 1% of what you see actually enters through your eyes (the rest is -surprisingly correct – made up by your brain).

Remote sensing image target statistical method and system fusing large language model and visual cue driving

The invention provides a remote sensing image target statistical method and system fusing a large language model and visual prompt driving. The method comprises the following steps: acquiring a remote sensing instance segmentation image to be processed and a visual prompt thereof; inputting a to-be-processed remote sensing instance segmentation image and a visual prompt thereof into the trained remote sensing image target statistical model, and outputting a remote sensing image target statistical result; the training comprises the following steps: introducing a large language model and visual cue into an encoder architecture of a GrondingDINO model to obtain a remote sensing image target statistical model; inputting a remote sensing instance segmented image and the visual cue thereof into an encoder, and outputting an image feature, a visual cue feature and a text feature; the feature intensifier carries out fusion processing on the output of the encoder; a language-guided query selection module calculates cross-modal query according to the fusion processing result, and a cross-modal decoder obtains a target statistical result of the image based on the fusion processing result and the cross-modal query; and training by using the training data and outputting the trained model.
Owner:WUHAN UNIV

Method, device, and medium for training large scale object foundation model

Embodiments of the present disclosure provide a method, device, and medium for training a large scale object foundation model. The method comprises obtaining a training dataset comprising a plurality of subsets for a plurality of object perception tasks, wherein a sample in the training dataset comprises an image with an object, a prompt indicating the object, and labeled object perception information of the image. The method further comprises generating, by the image encoder, an image feature based on the image. The method further comprises generating, by the text encoder or the visual prompt encoder, a prompt embedding based on the prompt. The method further comprises generating, by the object decoder, object perception information of the object based on the image feature and the prompt embedding. In addition, the method further comprises training the object processing model based on the generated object perception information and the labeled object perception information.
Owner:LEMON INC(GB)

Video action recognition model training method, video action recognition method and device

The invention relates to a video action recognition model training method, a video action recognition method and a video action recognition device. The method comprises the following steps: acquiring a sample video frame image and an action category description text corresponding to the sample video frame image, inputting the sample video frame image and the action category description text into a to-be-trained recognition model, and recognizing the action category of the to-be-trained recognition model by an image encoder in the recognition model according to a preset visual prompt vector and the sample video frame image, generating a video embedding corresponding to a sample video frame image, generating a text embedding corresponding to an action category description text by a text encoder in the recognition model based on a preset text prompt vector and the action category description text, and constructing bidirectional comparison loss by taking the video embedding and the text embedding as positive sample pairs, and updating the visual prompt vector and the text prompt vector based on the bidirectional contrast loss to obtain a trained recognition model. By adopting the method, the video action recognition accuracy can be improved.
Owner:CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1

Visual language model training method, image tag prediction method, and electronic device

PCT designated stageWO2026026225A1Biological modelsFeature extractionLinguistic model
Embodiments of the present disclosure provide a visual language model training method, an image tag prediction method, and an electronic device. According to the embodiments of the present disclosure, encoding parameters respectively corresponding to an image encoder and a text encoder in a first visual language model are fixed, on the basis of initial values of visual prompts and text samples respectively corresponding to a plurality of image classification tags, a second visual language model is generated by training the first visual language model, and the visual prompts respectively corresponding to the plurality of image classification tags are mined and learned, thereby enhancing the visual representation capabilities of visual language models. Next, on the basis of the learned visual prompts, text prompts, and the text samples, a target visual language model is further trained by training the second visual language model, and adapter parameters of model adapters respectively configured for the image encoder and the text encoder are adjusted, to collaboratively optimize image feature extraction and text feature extraction. This process allows for the transfer of visual knowledge from visual prompts to text prompts.
Owner:HANGZHOU ALIBABA INT INTERNET IND CO LTD

Visual prompt multi-modal large model for multi-source remote sensing image interpretation

The invention discloses a visual prompt multi-modal large model for multi-source remote sensing image interpretation, which belongs to the technical field of crossing of remote sensing and computer vision and comprises a multi-modal content coding and integration module, a cross-domain first-stage fusion training module, a pixel level visual positioning module and a large language model. The model supports the interpretation of a remote sensing image after the remote sensing image is arbitrarily amplified and reduced, and has flexible multi-granularity vision and language interaction capability. According to the model, a large language model is used as an interface, and multi-modal content integration including multi-sensor images, visual prompts and text instructions is achieved. In addition, two types of space tasks of anaphora understanding and visual positioning are unified into a visual prompt learning framework, and comprehensive and flexible multi-granularity understanding of remote sensing data is promoted.
Owner:BEIJING INST OF TECH

Target navigation method and device based on active 3DGS and visual language model reasoning

The invention discloses a target navigation method and device based on active three-dimensional Gaussian spatter and visual language model reasoning, and the method comprises the steps: obtaining an RGB-D image and a pose in an unknown environment, and constructing an incremental three-dimensional Gaussian spatter map as persistent memory through active perception; generating an exploration map based on the constructed 3DGS map, extracting a leading edge point and carrying out space structure adaptive clustering; generating a guidance track, carrying out free viewpoint optimization based on the track, and rendering a leading edge point first-person view angle image containing rich information; constructing a structured visual prompt, combining with a thinking chain prompt, inputting a visual language model to carry out reasoning planning, and selecting an optimal navigation target; in the navigation process, a real-time target detector is used for screening potential targets, and a new view angle is rendered in a 3DGS space through an action decision VLM for target re-verification. The navigation success rate and efficiency are improved.
Owner:ZHEJIANG UNIV OF TECH

Explore until confident: efficient exploration for embodied question answering

A method for embodied agent exploration is described. The method includes building a semantic map of a surrounding scene based on depth information and via visual prompting of a vision language model (VLM). The method also includes utilizing conformal prediction to calibrate a question answering confidence of the VLM. The method further includes performing, by an embodied agent, scene exploration utilizing knowledge of relevant regions of the scene. The method also includes determining, by the embodied agent, when to terminate the scene exploration utilizing a calibrated question answering confidence of the VLM.
Owner:TOYOTA RESEARCH INSTITUTE INC +3

Apparatus and Method for Sensory Adjustments in Electric Vehicles

A method and apparatus enables modifying the electronic controls of EVs to mimic the sensory experience of driving a performance ICE car. The method and apparatus creates a sensory “virtual cockpit” with both electronic and mechanical enhancements for a sensory experience. By downloading and implementing the method and apparatus, one may mimic, for example, the vehicle dynamics, performance horsepower, torque curves, suspension settings, oversteer and understeer behavior, steering-wheel inputs, cabin sound, subtle cabin vibrations, and audio / visual cues via a graphical user interface. These simulations replicate, in an electric vehicle, the various aspects of an ICE vehicle to mimic the whole experience of driving various ICE performance vehicles.
Owner:LOCCISANO VINCENT

Multi-modal map enhanced retrieval method and dialogue system based on feature fusion optimization

The invention discloses a feature fusion optimization-based multi-modal map enhancement retrieval method and a dialogue system. The method comprises the following steps of: respectively carrying out pre-training and fine tuning on a visual model and a language model by utilizing a domain image and text data; constructing a knowledge graph based on the text data in the knowledge base and constructing a vector database containing associated image data; performing semantic analysis and optimization on the original query of the user by using the language model and forming a structured retrieval intention; searching related sub-graphs, text semantic vector information and associated image data based on the search intention; encoding the sub-images into knowledge contexts, inputting the knowledge contexts into a dynamic prompt generator to generate visual prompts, and extracting enhanced visual features from the associated image data through a visual model; and inputting the subgraph, the text semantic vector information and the enhanced visual features into a language model for collaborative reasoning, and generating and outputting a final answer. According to the method, deep fusion and accurate retrieval of multi-modal knowledge can be realized, and the accuracy and efficiency are remarkably improved.
Owner:ZHEJIANG UNIV

Weak supervision video anomaly detection method based on vision and text double decision

The invention belongs to the field of computer vision and image processing, and provides a weak supervision video anomaly detection method based on vision and text double decisions. According to the method, local and global time sequence modeling modules are constructed, an anomaly focusing visual prompt mechanism and a learnable text prompt vector are introduced, and fine-grained anomaly recognition is realized in combination with a dual-mode memory. In the training stage, the model extracts visual and text features on the premise of no frame-level supervision based on video-level labels, constructs a category alignment graph and optimizes category embedding. In the reasoning stage, the model dynamically updates high-confidence-degree features through a positive and negative memory mechanism, and suppression of prediction deviation between semantic proximity categories is achieved. Compared with the prior art, the method has the advantages that the abnormal behavior recognition capability and the multi-class distinguishing precision of the model are remarkably improved under the weak supervision condition, the structure is simple, deployment is easy, adaptability is high and the like, and the method is suitable for efficient video anomaly detection tasks in intelligent monitoring, behavior recognition and other scenes.
Owner:DALIAN UNIV OF TECH

A multi-modal visual understanding method based on consistent learning and mixed feature extraction

This invention relates to a multimodal visual understanding method based on consistency learning and hybrid feature extraction. It includes constructing an end-to-end fine-grained consistency learning framework, introducing a hybrid region extractor, fusing local details and global semantics to generate high-quality hybrid visual cue embeddings, combining self-reconstruction loss and latent spatial consistency loss to force the model to establish explicit alignment between the input visual cue and the output segmentation label, utilizing the geometric boundary constraints of the localization task for description generation, and simultaneously optimizing localization accuracy using the semantic depth of the description task. Furthermore, it constructs a detailed localization index expression and segmentation task to enhance the model's reasoning ability for complex long text instructions. The aim is to address the problems of feature fragmentation and insufficient accuracy in existing large models for fine-grained visual localization and description tasks. Compared with existing technologies, this invention has advantages such as high accuracy and strong generalization ability in pixel-level localization and fine-grained description.
Owner:TONGJI UNIV

Facial expression recognition method and device based on multimode attitude measurement learning, and storage medium

PendingCN121861709ARealize precise identificationSolve the problem of weak generalization abilitySemantic analysisBiological modelsBiologyMachine learning
The invention discloses a facial expression recognition method and device based on multimode attitude measurement learning, and a storage medium, relates to the technical field of facial expression recognition, and solves the technical problems of weak generalization ability, insufficient cross-modal feature fusion and easy forgetting of pre-training knowledge due to data scarcity in the existing facial expression recognition technology. The method comprises the following steps: acquiring image data containing facial expressions, preprocessing the image data, constructing a text prompt word set, and constructing triple training samples of anchor points, positive samples and negative samples; a visual cue word vector is designed to be injected into a Transform layer of a pre-trained CLIP-ViT-B / 32 model, and model trunk parameters are frozen to only train visual prompt and LayerNorm layer parameters; triple loss and classification loss combined optimization is adopted; in the reasoning process, the expression category is determined by calculating the similarity between the to-be-recognized image features and various expression text prompt features, pre-training semantic knowledge and cross-modal adaptation ability of CLIP are fully utilized, and the expression recognition precision, robustness and zero sample generalization ability are remarkably improved.
Owner:DIANSHI TECH (ZHEJIANG) CO LTD +1

Directive image segmentation method based on multi-mode prompt

The invention relates to a directional image segmentation method based on multi-modal prompt, which fuses natural language prompt and visual prompt information to realize accurate segmentation of a target area in an image. Specifically, the directional image segmentation model adopts an encoding-decoding structure and is composed of an image encoder, a text encoder, a pixel encoder, a visual prompt encoder and a shared mask decoder. By introducing a multi-layer embedded module and a deformable attention mechanism in a visual cue encoder, the model can fully extract space and semantic information of a cue area in a reference image. A text encoder extracts semantic features by using a pre-training language model, and performs modal fusion with visual prompts to effectively realize information complementation. In addition, the directivity image segmentation model supports segmentation tasks guided only by texts, only by vision or by the combination of the texts and the vision, and has good adaptability and task universality.
Owner:XIAMEN UNIV

Rehabilitation behavior analysis method, device and equipment based on dynamic visual cues generation

The application discloses a rehabilitation behavior analysis method, device and equipment based on dynamic visual prompt generation, comprising: collecting action video data, extracting two-dimensional skeleton point sequence data from the action video data through a pose estimator; extracting visual features and prompt features from video modal data and skeleton point modal data through a visual classification model and a skeleton classification model respectively; pre-training the visual classification model using labeled video modal data, and then performing contrast learning on the visual classification model and the skeleton classification model using unlabeled video modal data and corresponding unlabeled skeleton modal data respectively; collecting rehabilitation training video data of a patient and extracting corresponding skeleton modal data; and inputting the video data and the corresponding skeleton modal data into the trained visual classification model and the skeleton classification model respectively for classification. The application can realize more complete action representation, reduce data dependence, and improve the accuracy of rehabilitation behavior analysis.
Owner:ZHEJIANG UNIV

Audio-visual collaborative-based early parkinson disease gait intervention method, device and equipment

The application provides a gait intervention method, device and equipment for early Parkinson disease based on audio-visual cooperation. The intervention method comprises the following steps: a gait parameter calculation step; an audio-visual cooperation prompt paradigm construction step, gait rhythm characteristics are determined based on left and right foot gait cycles to generate auditory prompts synchronized with the gait rhythm characteristics, and visual prompts in the form of virtual people in a mixed reality environment are generated based on left and right foot step heights and foot position and ground detection information; a modal fusion step, the auditory prompts and the visual prompts are synchronized in the time dimension; and an audio-visual cooperation intervention step, gait intervention is performed through the synchronized auditory prompts and the visual prompts. The application can improve the gait symmetry of early Parkinson disease patients through the accurate synchronization of visual and auditory prompts in the time dimension.
Owner:INST OF SOFTWARE - CHINESE ACAD OF SCI

Model training method and device and target detection method and device

The invention discloses a model training method and device and a target detection method and device, and belongs to the field of target detection. The method comprises the following steps: acquiring a sample image, wherein the sample image comprises at least one object and a labeling category of the at least one object; performing target detection on the sample image by adopting a pre-training model to determine a prediction category of the at least one object; and training the pre-training model according to the annotation category of the at least one object and the prediction category of the at least one object to obtain a target detection model. The pre-training model comprises a visual prompt module and a target detection module, and the visual prompt module is used for carrying out weighting processing on a plurality of fusion features of the sample image according to the annotation category of the at least one object to obtain a plurality of weighted features; the target detection module is used for performing target detection according to the plurality of weighted features to determine a prediction category of the at least one object. According to the invention, the target detection accuracy of the target detection model on the image comprising the new category of objects can be improved.
Owner:BOE TECHNOLOGY GROUP CO LTD +1

Methods, devices, computer equipment and storage media for generating pet visual data

This application provides a method, apparatus, computer device, and storage medium for generating pet visual data. The method includes receiving a data generation request, which includes visual cue text and reference visual data of a target pet; invoking a pre-trained generative model based on the data generation request, wherein the generative model includes a data fusion sub-model and a visual processing sub-model; invoking the data fusion sub-model to perform data fusion based on the visual cue text and the reference visual data to obtain a visual latent vector; and invoking the visual processing sub-model to generate visual data based on the visual latent vector to obtain target visual data corresponding to the visual cue text, wherein the target visual data includes an image and / or a video of the target pet. This method can improve the consistency of pet visual data generation.
Owner:SHENZHEN LIBRO TECH CO LTD

Multi-modal instance-level understanding method and system based on visual prompt

The invention relates to the field of computer vision, and particularly discloses a multi-modal instance level understanding method and system based on visual cues, and the method comprises the steps: carrying out the space-time segmentation of a specific instance through employing an interactive instance segmentation model and a video instance tracking model, and generating a visual prompt; then, a visual encoder is used for encoding a video with a visual prompt, and a cross-modal connection module is used for mapping visual representation to a multi-modal representation space shared with a language, so that visual features are obtained; processing the input text by using a text word segmentation device to obtain corresponding text features; and finally, utilizing a large language model to uniformly model vision and language input to obtain fine-grained description or question answers about a specific instance. Compared with the prior art, the method has the advantages that accurate positioning and tracking of the specific instance in the space-time dimension are realized, and the instance-level fine-grained understanding capability of the video multi-modal model is improved.
Owner:FUDAN UNIVERSITY

Raman image fusion multi-modal fusion classification method

The invention relates to a Raman image fusion multi-modal fusion classification method, which relates to the technical field of Raman image analysis, and comprises the following steps: injecting an exclusive prompt vector P for each domain based on a visual prompt adjustment mechanism in a visual Transform, and eliminating the influence of different acquisition conditions on visual features; based on a multi-scale residual convolutional network, in cooperation with batch normalization processing, the influence of different acquisition conditions on spectral characteristics is eliminated; fusion data of visual features and spectral features are obtained based on a multi-modal fusion algorithm, and a multi-modal fusion classification model is constructed according to the fusion data, so that the advantages of images and spectrums are integrated, and the average recognition rate of data under different collection conditions is improved.
Owner:QINGDAO SINGLE CELL BIOTECH CO LTD

Visual cue-based footwork and tactical response training system for fencing

A fencing training system delivers programmable visual cues to simulate referee commands and tactical bout scenarios. It includes a transparent panel (11) with addressable LED lights (10), mounted on a tripod (12) and controlled by a microcontroller (14). Timed color and pattern sequences prompt fencers to initiate offensive, defensive, or interpretive actions. Users can customize delay intervals, cue types, and repetition frequency. The system improves timing accuracy, footwork consistency, and decision-making within the four-meter engagement zone. Configurable preparation windows and distance markers (20, 22, 24) aid spatial judgment. The system is adaptable for saber and foil fencing and may include auditory cues, motion detection, or mobile app control. Designed for portability and real-time responsiveness, it supports solo and group training. The invention may also be adapted for other sports requiring timed decision-making in confined spaces and incorporate AI-based customization using bout performance data to personalize footwork timing and cue sequences.
Owner:CHEN COOPER YE-ZOU

Matting method and electronic device

The application provides a matting method and an electronic device. The method comprises the following steps: inputting a to-be-processed image into an image encoder of an image processing model to obtain a first feature vector, inputting a visual prompt box into a prompt encoder of the image processing model to obtain a second feature vector, the visual prompt box being obtained according to a manually selected region in the to-be-processed image; inputting the first feature vector and the second feature vector into a mask decoder of the image processing model to obtain a global matting mask and a local matting mask; combining the global matting mask and the local matting mask to determine an overall mask; and cutting out an image region where a target is located from the to-be-processed image according to the overall mask. Thus, the image processing model is used to separately process the edge region to generate a local matting mask of the edge region with high accuracy, the overall mask is generated by combining the global matting mask corresponding to the foreground region, and finally, the matting is completed according to the overall mask, so that the matting effect can be improved.
Owner:HONOR DEVICE CO LTD

Systems, apparatuses, methods, and non-transitory computer-readable storage media for facilitating user interaction with a virtual keyboard

Systems, apparatuses, methods, and computer-readable storage media are disclosed for facilitating user interaction with a virtual keyboard. A computerized method comprises: receiving contextual information associated with a virtual keyboard displayed on a user interface; determining, based on the contextual information, that a visual hint display condition is met; displaying a visual hint associated with one or more keys of the virtual keyboard based on the visual hint display condition; after displaying the visual hint, receiving additional contextual information associated with the virtual keyboard; determining, based on the additional contextual information, that a visual hint dismissal condition is met; and dismissing the visual hint from the virtual keyboard based on the visual hint dismissal condition.
Owner:HUAWEI TECH CO LTD

System to facilitate parking a vehicle straight in a desired location inside a garage

An apparatus to be viewed by a driver to aid parking a car parallel to the walls of a garage, as well as the correct distance into the garage. The driver is able to park in exactly the same spot each time, at the desired distance from other cars, walls, or other stationary objects via the aid apparatus. This helps minimize banging car doors on objects upon opening, or bumping into anything in front of the vehicle, while maximizing use of space within the garage. The apparatus is installed in a fixed position in front of the driver at eye level as they sit comfortably in a driver's seat of the vehicle. As the driver drives their car into their garage to park, they are guided by visual cues from the apparatus to park the car in the same footprint on the floor each time.
Owner:POWERS ROBERT WILLIAM

Dynamic adaptation of a given assistant output based on a given persona assigned to the automated assistant.

It provides dynamic adaptation of a given assistant output based on a given persona assigned to the automated assistant. [Solution] The method generates and then adapts a given assistant output based on a given persona assigned to the automated assistant. The given assistant output is specific to the given persona and is generated without the need to later adapt the given assistant output to the given persona. The given assistant output includes a stream of text content synthesized for audible presentation to the user and a stream of visual cues used to control the client device's display and / or the automated assistant's visual representation. The given assistant output is made to reflect a given persona using a Large-Scale Language Model (LLM) or output previously generated using an LLM.
Owner:GOOGLE LLC

Target detection method, equipment and target detection model

The invention provides a target detection method, target detection equipment and a target detection model. The method provided by the invention is realized based on a pre-trained target detection model, and the target detection model comprises a detection network, a modal mapping module and a large language model. The method comprises the following steps: inputting visual prompt information of a to-be-detected target and a to-be-detected image into a target detection model, and performing target detection on the to-be-detected image by a detection network according to the visual prompt information to obtain a plurality of feature vectors of a plurality of detected target objects; mapping the plurality of feature vectors to a feature space of a large language model by using a modal mapping module to obtain a coding feature vector; and generating a target category matched with the coding feature vector by using a large language model, and determining the target category as the category to which the plurality of target objects belong. According to the target detection method provided by the invention, the target object belonging to the same category as the to-be-detected target and the specific category to which the target object belongs can be accurately identified.
Owner:SHENZHEN INTELLIFUSION TECHNOLOGIES CO LTD

Multi-agent inquiry method and system fusing fuzzy semantics and focus image features

The invention discloses a multi-agent inquiry method and system fusing fuzzy semantics and focus image features. The method comprises the following steps: preprocessing a patient subjective description text and an original medical image to obtain a cleaned text and an enhanced image; analyzing the quantized text fuzzy semantics to generate a fuzzy semantic vector, guiding a visual model through the vector to extract focus features of an enhanced image, and generating a fuzzy visual prompt and attention thermodynamic diagram; then forming a consensus diagnosis result by multi-agent collaborative inquiry and debate reasoning based on the information; and finally, combining a diagnosis result, a thermodynamic diagram and a debate process to generate a structured report containing a diagnosis conclusion, visual evidence and a reasoning path. According to the method, accurate fusion of subjective fuzzy semantics and objective image features is realized, the medical illusion rate is reduced, the interactivity, interpretability and accuracy of diagnosis are improved, and the method is suitable for various intelligent medical auxiliary diagnosis scenes.
Owner:CENT SOUTH UNIV

Continuous sign language recognition method based on visual text prompt guidance

According to the continuous sign language recognition method based on visual text prompt guidance, frame-level semantic conditional fusion is carried out before fusion, and on the premise that an external sensor and a skeleton pipeline are not introduced, the frame-level semantic conditional fusion is carried out; frame-level visual feature sequences extracted by a video encoder are constructed in parallel, and are averagely pooled to obtain a video-level visual prompt vector and a frame-level text prompt vector obtained by a text prompt extraction module. Unified linear projection and splicing are completed in the prompt guide fusion module, visual-text prompt vectors are generated through a multi-layer perceptron, broadcast copying is carried out along the time dimension, layer normalization is carried out to complete frame-by-frame fusion after the visual-text prompt vectors are added with feature residuals output by a main model of each frame, and then the visual-text prompt vectors are input into an encoder and a CTC to be subjected to end-to-end training; the discriminability and the time sequence consistency of feature expression are improved, so that the synchronous improvement of identification optimization and feature optimization is realized, and the robustness and the identification performance of the model are enhanced.
Owner:TIANJIN UNIVERSITY OF TECHNOLOGY

Counterfeit detection method

The application discloses a forgery detection method, and relates to the technical field of forgery detection. The method is based on an image encoder to perform feature extraction on a to-be-detected image, obtain a first feature vector, a second feature vector and a visual prompt vector set, splice a randomly generated context vector, a category vector and the visual prompt vector to generate an input vector, perform feature extraction on the input vector based on a text encoder, obtain a third feature vector, determine a fusion feature vector and a first similarity based on the second feature vector and the third feature vector, determine a second similarity based on the second feature vector and the fusion feature vector, generate a heat map according to the first similarity and the second similarity, and determine a forgery detection result according to the heat map. Thus, the method can generate a heat map based on a continuous input vector to provide a detection result and mark a forgery position, thereby improving the explainability of forgery detection.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Multi-modal target detection model training method and system

The invention relates to the field of model training, in particular to a multi-modal target detection model training method and system, and the method comprises the steps: obtaining pseudo-annotation data associated with a target detection demand, and carrying out the pre-training of a target detection base model; based on a preset open source training data set, performing supervision fine tuning training on the pre-trained target detection base model; based on private data associated with a target detection demand, performing increment fine tuning training on the target detection base model subjected to supervision fine tuning training; and performing visual prompt training on the target detection base model subjected to increment fine tuning training to obtain a multi-modal target detection model. According to the technical scheme provided by the invention, training can be carried out based on multiple stages, multi-source training data is coordinated, the basic detection and recognition capability of the target detection model can be expanded by utilizing pseudo-label data with a large amount of noise, recognition enhancement and improvement of specific target categories by utilizing private data can be supported, and the method has a promotable value.
Owner:CHONGQING ZHONGKE YUNCONG TECH CO LTD