Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

392 results about "Associated image" patented technology

Music segment tagging, sharing, and image generation

A method of automated generation of contextually-relevant images for a music segment includes receiving at least one of basic metadata information and lyric information for the music segment, generating a first prompt for a computer-implemented machine-learning language model based on the at least one of the basic metadata information and the lyric information, receiving context information from the computer-implemented machine-learning language model in response to the first prompt, generating a second prompt for the computer-implemented machine-learning language model based on the context information, generating a third prompt by providing the second prompt as an input to the computer-implemented machine-learning language model, and generating an image descriptive of the music segment by providing the third prompt as an input to a computer-implemented machine-learning image generation model.
Owner:HOOK MEDIA LLC

Multi-image processing method based on multi-modal entity alignment

The invention belongs to the technical field of image processing, and discloses a multi-image processing method based on multi-modal entity alignment, which comprises the following steps: acquiring a plurality of images; retrieving rich semantic information of the entity from an external knowledge base; selecting the most representative image for each entity in the multi-image scene by using semantic information; encoding the original input of each mode; performing first-layer fusion between a visual mode and a text mode by using cross-divergent attention, then performing second-layer interaction with a structured mode, and finally aligning an entity representation of an image by using contrast loss; and outputting multiple images fusing the text modality and the visual modality. According to the invention, hierarchical interaction fusion is utilized to enhance multi-modal interaction; enhancing entity text representation by integrating the external attribute value and the context information; and the most representative image is selected by using the semantic text, so that the influence of uncorrelated images is minimized.
Owner:NAT UNIV OF DEFENSE TECH

Fire situation analysis method based on multi-dimensional data fusion

The invention discloses a fire situation analysis method based on multi-dimensional data fusion, and particularly relates to the field of fire image analysis, and the method comprises the steps: analyzing the change of the direction retention rate between adjacent frames through extracting the main direction texture vectors of a building and a vegetation region, and recognizing object state change candidate segments; in the candidate area, combining main direction disturbance and image definition reduction to construct a spatial scoring graph, performing nonlinear amplification on the spatial scoring graph, and extracting a gradient increasing path to generate a structure damage main path set; then calculating a directional included angle between paths and a space coincidence rate, constructing a trend consistency aggregation channel graph, and extracting a spreading principal axis; and finally, superposing a temperature rise area in the thermal infrared image with a spreading principal axis, extracting a dual response area to generate a fire behavior boundary prediction layer, and realizing fire behavior spreading path prediction based on fire scene related image structure damage information analysis.
Owner:TIANJIN SHENGDA SECURITY TECH CO LTD +1

Object Detection Method and System Based on User-Defined Category

The provided is a method and system for object detection based on user-defined categories. The method includes: a user inputting a natural language description and a related image, obtaining a detection target auxiliary input using an auxiliary characterization generation technique for a detection target based on a phrase boundary point modeling technique; calling a detection target characterization generation model based on a multimodal reconstruction and alignment network to obtain a plurality of text characterizations of the detection target; generating target reverse characterizations based on an image-adaptive target characterization matching estimation technique to meet custom requirements of the detection target; and optimizing a vision-language multimodal model based on feedback data of the detection target of the user under detection, and optimizing the vision-language multimodal model based on the feedback data during usage of custom object detection.
Owner:HANGZHOU MEARI TECH CO LTD

Ai-driven creation of custom stickers from messages in chat interfaces

PendingUS20250378602A1Mathematical modelsNatural language analysisEngineeringVisual expression
This disclosure relates to techniques for generating and utilizing custom stickers in a digital communication environment. A technique involves receiving a text-based message input during a chat session and using a generative language model (e.g., a Large Language Model, or LLM) to create a text prompt. This prompt is then used by a generative image model to produce a custom sticker. The generated sticker is sent to a client device where it is displayed in a sticker tray alongside other selectable stickers. Users can select and send these stickers directly within their chat interface, enriching communication with visually expressive and contextually relevant imagery.
Owner:SNAP INC

Structural cursor calibration method and system under three coordinates

The invention provides a structured light calibration method and system under three coordinates, and relates to the technical field of structured light imaging, and the method comprises the steps: 1, collecting calibration plate images of a plurality of angles, extracting sub-pixel angular point data, constructing a camera parameter description mode including radial and tangential distortion, and based on a pinhole imaging principle, carrying out the calibration of a plurality of angles; solving an internal parameter matrix and a distortion coefficient of the camera by establishing a linear equation set so as to obtain calibrated camera parameters; 2, projecting a composite coding pattern to the calibration plate based on the calibrated camera parameters, synchronously acquiring related images, decoding phase and coding data, establishing a corresponding relation between projector pixels and three-dimensional space points of the calibration plate by using the calibrated camera parameters, and solving an internal and external parameter matrix of the projector to obtain initial projector parameters; according to the invention, through multi-view image acquisition, composite coding projection and three-dimensional target matching, structured light collaborative calibration of camera and projector parameters is realized.
Owner:XIAN HIGH TECH AEH INDAL METROLOGY

Artificial intelligence chatbot

Methods and systems for interacting with users via a chatbot. A natural language query is received and processed by submitting a search query to a search engine. The search engine identifies relevant information including textual information and images for formulating a response. The identified information and query are submitted to a Large Language Model which generates a response displayed via the chatbot. The response may include textual information and relevant images. The system can extract text from images of documents and convert textual information into numerical vector representations for processing. Selectable options based on clustered relevant information can be provided to users for query refinement when appropriate. The chatbot interface enables natural language interactions while leveraging search capabilities and Artificial Intelligence to provide informative and helpful responses with both text and visual elements.
Owner:HONEYWELL INTERNATIONAL INC

Geosynchronization of an aerial image using localizing multiple features

A georegistration (a.k.a. georectification) of an image captured by a camera in an aerial vehicle, such as a satellite, is based on identifying multiple features using descriptor sets, and sending to a ground station only the descriptors of the identified features and the associated locations in the captured image, without sending of the captured image itself, thus requiring a low communication bandwidth. Using a database of geosynchronized reference images, the ground station uses the received descriptors sets and the associated image locations to localize the features on a selected geosynchronized reference image from the database, and forms a mapping function that map any locations in the captured image to geographical coordinates on Earth. The mapping may be used to geosynchronize an additional feature identified in the aerial vehicle, or to geo synchronize a region that may be cropped from the captured image and sent to the ground station.
Owner:EDGY BEES LTD

Explanatable farmland image enhancement method and system based on physical perception and reinforcement learning

ActiveCN120876343AImage enhancementImage analysisImaging processingIncrement threshold
The invention relates to the technical field of image processing, in particular to an interpretable farmland image enhancement method based on physical perception and reinforcement learning, and the method comprises the steps: obtaining an original farmland image, carrying out the global perception analysis, extracting multi-scale features, and recognizing a degradation region and feature distribution; on the basis of the global perception analysis result, constructing a semantic enhancement blueprint for quantifying the physical attributes and optimization requirements of the degradation area; initializing a reinforcement learning agent according to the semantic enhancement blueprint, and selecting an image processing operation sequence through a physical constraint reward function; executing the image operation sequence, and terminating enhancement processing according to the quality evaluation index increment threshold and the physical consistency of the semantic enhancement blueprint; and outputting an interpretability report of the enhanced image and associated image operation sequence physical basis traceability. The objective of the invention is to solve the technical problems of lack of interpretability, insufficient environmental adaptability and rigid decision-making mechanism of a farmland image enhancement technology.
Owner:CHINA TOWER CO LTD

Multi-modal sentiment analysis method and system based on thinking chain and background knowledge

The invention relates to a multi-modal sentiment analysis method and system based on a thinking chain and background knowledge. The method comprises the following steps: constructing a sample set containing original texts, related images, aspect items and sentiment polarities thereof; constructing a multi-modal sentiment analysis model, and generating background knowledge information by utilizing a visual language large model and combining an input text and an image; the generated background knowledge is screened, optimized and fused; for each sample in the sample set, in combination with the text data and the emotion polarity, generating a thinking chain thinking process and adding the thinking chain thinking process into the sample set, and then finely adjusting the visual language large model through the sample set added with the thinking chain thinking process; finally, the fused background knowledge, the text data and the picture data are sequentially input into the trained visual language large model, aspects in the text data are extracted, and the emotion polarity corresponding to the aspects is predicted. The method and the system are beneficial to improving the accuracy of multi-modal emotion polarity analysis.
Owner:FUZHOU UNIV

Tear total IGE detection method and system based on multi-modal data fusion

The invention discloses a tear total IGE detection method and system based on multi-modal data fusion. The method comprises the following steps: acquiring ocular surface images, physiological parameters and environmental allergen data, and acquiring historical medical data from a hospital information system; preprocessing and feature extraction are performed, patient symptom description is analyzed in combination with a natural language processing technology, and related image block features are recognized; adjusting a space-time attention weight based on frequency domain analysis and transfer learning, inputting a lightweight Transform model, and predicting a tear total IgE concentration and an allergy risk score; whether a trace tear sampling program is started or not is judged, and the actually measured concentration of total tear IgE is detected through a disposable micro-fluidic chip; and integrating all information by using a Bayesian adaptive filtering algorithm to determine the total tear IgE level. By implementing the method provided by the invention, the sampling amount is extremely small, the detection time is short, the environmental interference can be dynamically corrected, and the absolute concentration value is output, so that the dual requirements of clinical precise diagnosis and treatment and home monitoring are met.
Owner:SHANGHAI LIANGXIN TECHNOLOGY CO LTD

Extracting images and determining their meaning for semantic image retrieval and training a transformer-based multi-modal large language model to generate domain-aware images based on image meanings

The disclosure relates to systems and methods automatically extracting an image and related image components, computationally determining an understanding of the image, and generating mathematical vector embeddings via sentence encoders based on the computationally determined understanding. The mathematical vector embeddings may be used for semantic image retrieval that enables image searching based on a semantic understanding of input images and / or input text. The mathematical vector embeddings may be used for training and executing generative Artificial Intelligence (AI) models to create new content that includes retrieved images and / or generate new images.
Owner:ROHIRRIM INC

Deep learning technique for automated radiological image analysis and disease detection

A real-time artificial intelligence (AI) framework is provided for the automated analysis of radiological images and detection of disease, such as extracapsular extension (ECE) in prostate cancer. The system includes a dual deep learning architecture comprising a first convolutional neural network (CNN) for identifying diagnostically relevant image slices from three-dimensional MRI data, and a second CNN for classifying disease presence based on those slices. A preprocessing pipeline standardizes and harmonizes image input, and cropping algorithms isolate the region of interest for enhanced model performance. This framework enables scalable, high-accuracy diagnosis across various imaging modalities including but not limited to MRI, CT, PET, ultrasound, and diverse disease types, improving clinical decision-making and supporting integration into real-time radiology workflows.
Owner:RES FOUND THE CITY UNIV OF NEW YORK

Historical record backtracking method based on LLM

The invention discloses a historical record backtracking method based on LLM, and relates to the field of data retrieval and natural language processing, and the method comprises the steps: collecting various operation data of a user in a local computer, extracting text information in pictures and voices in combination with an OCR technology and an ASR technology, and storing the text information and related information in a database in a unified format; the method comprises the following steps of: processing collected multi-modal data information, firstly performing de-duplication and cleaning on the collected data, then converting a collected screenshot into a video, and finally performing vector embedding on a video image frame, an OCR (Optical Character Recognition) result and a voice transcription text and storing the video image frame, the OCR result and the voice transcription text into a vector database; optimizing the pre-trained LLM in a specific direction by adopting a fine tuning technology; according to the method, conversation with a historical timeline is realized, user input is captured and stored in a query queue, personalized response is generated through LLM interaction in combination with a memory module and an RAG technology, and answers and related image or audio links are displayed.
Owner:HUNAN UNIV

Large-view-field high-resolution compound eye camera array system based on bionic human eye vision

The invention provides a large-view-field high-resolution compound eye camera array system based on bionic human eye vision. The system is composed of a curved surface support and an industrial camera array arranged on the curved surface support. Reverse extension lines of optical axes of all the cameras intersect at the same point to form a concave surface structure similar to the retina, and the fixed view field overlapping rate of the adjacent cameras is kept. The industrial camera array is used for collecting a target scene image, large-view-field information is obtained by fusing and splicing all camera images, and a high-resolution image is obtained by conducting super-resolution reconstruction on a plurality of related images in a view field center interested area. The system provides three working modes, namely a low-resolution mode, a conventional-resolution mode and a super-resolution mode, which can be switched by adjusting camera parameters. Finally, the system can perceive the dynamic change of the edge of the field of view in a high-frame-rate and low-resolution mode, magnify the details in the center of the field of view in a super-resolution mode, realize efficient non-uniform imaging, and optimize the information acquisition and processing efficiency and the structure compactness.
Owner:SICHUAN UNIV

Parking lot-based lidar monitoring method, system, device, and storage medium

The application provides a parking lot-based laser radar monitoring method, system, device and storage medium, wherein the method comprises the following steps: collecting point cloud data in a parking lot and dividing the point cloud data into a static point cloud set and a dynamic point cloud set, dividing target dynamic point cloud sets to be tracked from the static point cloud set and the dynamic point cloud set, obtaining a coincidence coefficient of a first projection area of a preset parking space and a second projection area of the target dynamic point cloud set, and updating the target dynamic point cloud set to a static point cloud when the target dynamic point cloud set is static in the range of a parking space point cloud and the coincidence coefficient meets a preset threshold, and determining that a collision is sent when the nearest distance between the second contour of the target dynamic point cloud set and the third contour of the static point cloud representing all vehicles parked on the parking space is 0. The application can solve the problem of car safety protection after parking, confirm the danger and capture the relevant image, and provide the greatest vehicle safety guarantee for the owner.
Owner:CHINA TELECOM CORP LTD

Monitoring video viewing method and device, computer equipment and medium

The invention relates to a monitoring video viewing method and device, computer equipment and a medium, and the method comprises the steps: determining a source camera of a target monitoring video needing to be viewed and corresponding viewing time information in response to a monitoring video viewing instruction; searching a camera to which an external monitoring video having a view overlapping area with the target monitoring video within the same time according to the viewing time information, and taking the camera as a part of neighbor cameras of the source camera; determining the optimal privacy metadata from the privacy metadata of the corresponding viewing time information of the partial neighbor cameras, wherein the privacy metadata comprises frame associated image blocks corresponding to the shielded areas of the privacy sensitive targets in the external monitoring videos of the corresponding neighbor cameras; and repairing a covered area of a corresponding privacy sensitive target in the target monitoring video according to the frame associated image block of the optimal privacy metadata, and then playing the target monitoring video. According to the invention, on the premise of default privacy protection, key details of the monitoring picture are intelligently restored and enhanced as required.
Owner:深圳市灵智无界科技有限公司

Intelligent identification method and device for personnel tumbling in logistics center

The invention relates to the technical field of logistics personnel safety protection methods, and discloses an intelligent identification method for personnel tumbling in a logistics center. Comprising the following steps: acquiring a picture sample monitored by a logistics transfer center camera, pre-processing the picture sample to generate a pre-processed picture sample, and labeling the pre-processed picture sample through LabelImg; the method comprises the following steps: training a YOLOv11 model through an IoU-based losses loss function; training a human body posture estimation OpenPose model, and adjusting the architecture of the OpenPose model according to personnel characteristics; the trained YOLOv11 model and the trained OpenPose model are fused, and a personnel wrestling identification model is obtained; pictures monitored by a logistics transfer center camera are input into a personnel wrestling recognition model frame by frame, whether personnel wrestling exists in the pictures or not is judged through the personnel wrestling recognition model, if yes, an alarm is triggered, monitoring personnel are notified through an audible and visual alarm, and related images and video clip storage records are stored; the system has the advantages that the working state of the employee is monitored in real time, wrestling events are found in time, and therefore the safety management level and the response speed are improved.
Owner:SHANGHAI DONGPU INFORMATION TECH CO LTD

Applications for gain curves in imaging and video

Techniques are disclosed relating to exchange of images in networked computing applications. In particular, the disclosure relates to exchange of gain curves that are used to represent imaging and / or video in such applications. A gain curve may define a mathematical transformation that relates values from a source image domain to a destination image domain. The image and its associated gain curve(s) may be published to destination devices for consumption. When a destination device consumes the image, the destination device may apply a transform to source image content according to the gain curve(s) published with the image. For example, the destination device may apply a gain curve to an associated image directly, or it may derive another transform from the gain curve and additional information known to the destination device.
Owner:APPLE INC

Physician-guided machine learning system for assessing medical images to facilitate locating of a historical twin

A computer-implemented method of evaluating a user image of a patient to enable identification of a historical twin of the patient. The method includes organizing a plurality of medical images in an archive and receiving from a medical professional each of: (i) a region of interest; (ii) a textual description; (iii) selections for binary criteria; and (iv) weights of weighable criteria. The method comprises using a natural language search to create a relevant set of medical images and creating an optimal set from the relevant set of medical images by discarding medical images from the relevant set based at least on the selections for binary criteria. The method includes image processing medical images in the optimal set using the weight of the features of the region of interest to create medical image results. The relevant set comprises less than ten percent of the medical images in the archive.
Owner:IRANI NEVILLE

Region-text caption generation using global caption information

Approaches presented herein may be used to generate captions using raw caption information. Raw caption information may be used, with an associated image, to generate a detailed image caption. Object lists may then be generated from the image and / or the detailed image caption to produce an image including boxing box proposals for objects within the image. One or more trained machine learning systems may then be used to generate region of interest captions that infuse the global caption context associated with the raw caption information.
Owner:NVIDIA CORP

Systems and methods for estimating parking spot availability

Methods and systems for assisting drivers in finding available parking spots by estimating parking spot availabilities. Vehicle image sensors, and associated image processing techniques, detect whether or not other vehicles are located in parking spaces. An occupancy status of each parking spot is determined over time based on these image processing results, and the occupancy status is stored, for example as historical data. The occupancy statuses can be stored or retrieved based on a particular geographical region. An estimated occupancy status of a first parking spot can be generated based upon a last-known occupancy status of that parking spot and the average parking spot availabilities of nearby parking spots in that geographic region. The estimated occupancy status can be displayed, for example as overlaid into a map or navigation application.
Owner:VALEO SCHALTER & SENSOREN GMBH

Ecological regreening effect detection system and method

The invention relates to the field of ecological regreening detection, and particularly discloses an ecological regreening effect detection system and method.The system comprises an information acquisition module, an image analysis module, a regreening effect judgment module, a regreening analysis module and a regreening early warning module; the image analysis module determines a target regreening area and divides the target regreening area into a plurality of regreening analysis areas, the vertical distribution image extracts regreening analysis features of each area, and the regreening effect judgment module determines the effect judgment type of each regreening analysis area and judges whether dominant anomaly exists or not. The regreening analysis module further judges whether ecological regreening abnormity exists or not, the regreening early warning module sends out an early warning signal according to the abnormal judgment result, the accuracy of regreening effect detection is improved, and data support is provided for improving regreening efficiency.
Owner:TIANJIN GEOLOGICAL ENG INVESTIGATION INST

Aviation oil pipeline unmanned aerial vehicle intelligent inspection method and system

The invention discloses an aviation oil pipeline unmanned aerial vehicle intelligent inspection method and system, and the method comprises the steps: collecting visual image data along a pipeline through an inspection terminal carried by an unmanned aerial vehicle, covering a pipeline body and a surrounding environment, and recognizing an abnormal scene endangering the safety of the pipeline; inputting the image data into an abnormal scene recognition model deployed at an unmanned aerial vehicle end, and judging whether a preset type of abnormality is included; if at least one type of abnormity is identified, generating an alarm signal; alarm and related images are uploaded to a remote monitoring center server through wireless communication, and intelligent unmanned inspection of the running state of the pipeline is achieved. An image acquisition and recognition model is integrated at an unmanned aerial vehicle end, a pipeline and an environment are sensed in real time, and key abnormity is automatically recognized; the edge deployment model reduces invalid return, and only triggers alarm uploading when a risk is detected; and in combination with wireless communication return alarms, high-reliability monitoring is realized, and intelligent and refined guarantee is provided for safe operation of aviation oil pipelines.
Owner:CHINA AVIATION OIL PENGZHOU PIPELINE TRANSPORTATION CO LTD

Multi-modal map enhanced retrieval method and dialogue system based on feature fusion optimization

The invention discloses a feature fusion optimization-based multi-modal map enhancement retrieval method and a dialogue system. The method comprises the following steps of: respectively carrying out pre-training and fine tuning on a visual model and a language model by utilizing a domain image and text data; constructing a knowledge graph based on the text data in the knowledge base and constructing a vector database containing associated image data; performing semantic analysis and optimization on the original query of the user by using the language model and forming a structured retrieval intention; searching related sub-graphs, text semantic vector information and associated image data based on the search intention; encoding the sub-images into knowledge contexts, inputting the knowledge contexts into a dynamic prompt generator to generate visual prompts, and extracting enhanced visual features from the associated image data through a visual model; and inputting the subgraph, the text semantic vector information and the enhanced visual features into a language model for collaborative reasoning, and generating and outputting a final answer. According to the method, deep fusion and accurate retrieval of multi-modal knowledge can be realized, and the accuracy and efficiency are remarkably improved.
Owner:ZHEJIANG UNIV

Cow accurate feeding method and system based on vision and artificial intelligence

The invention provides an accurate cattle feeding method and system based on vision and artificial intelligence. The method comprises the steps that S1, related images of a trough of a cattle pen are collected; s2, performing histogram equalization and image defogging operation on the acquired image; s3, detecting and positioning a trough in the image based on a YOLOv5 target detection algorithm, and obtaining and extracting a trough region image; s4, carrying out classification processing on the extracted trough image through a ViT classification algorithm, and obtaining the scoring condition of the trough; s5, comprehensively analyzing the grading condition of the trough, and automatically adjusting and modifying the next feeding plan in combination with the growth condition of the cattle in the barn; and S6, storing historical data, feeding records and algorithm models of the cattle, and performing data analysis and algorithm optimization. Real-time trough score detection can be realized under different illumination conditions and complex environments, so that a cattle feeding plan is debugged in time, the method is suitable for a complex pen scene, and the labor cost and subjective errors can be remarkably reduced.
Owner:OPTICAL VALLEY JINXIN (WUHAN) TECH CO LTD

Method, computer device, and computer-readable recording medium to provide message summary and associated image

A method of providing a message summary and an associated image may include requesting an image search in relation to a message summary created based on a message in a chatroom; receiving an image bundle that includes at least one image in response to an image search request; and displaying the image bundle in association with the message summary.
Owner:LINE PLUS

Power system-oriented dynamic knowledge base driven dialogue generation system, method, equipment and medium

The invention discloses a power system-oriented dynamic knowledge base driven dialogue generation system, method, equipment and medium, and the system comprises a dialect collection and recognition module which is used for collecting dialect voice data under different regional power scenes, carrying out the noise reduction of the dialect voice data, carrying out the dialect recognition through a deep learning model, and obtaining a dialect recognition result; converting the dialect voice data into a standard text; a dynamic knowledge base module; the language processing and image reasoning module is used for receiving the standard text, performing semantic understanding, generating semantic representation in combination with a knowledge base in the dynamic knowledge base module, and performing target detection and recognition on a power equipment fault related image input by a user to obtain an image recognition result; and a dialogue generation module. According to the invention, a cooperative system of four modules of dialect acquisition and identification, a dynamic knowledge base, language processing and image reasoning and dialogue generation is constructed, so that intelligent dialogue service oriented to the power industry is realized.
Owner:GUIZHOU POWER GRID CO LTD

Associated imaging impurity detection method and system for drug production

The invention discloses a correlated imaging impurity detection method and system for drug production, and relates to the technical field of drug impurity detection. The method comprises the following steps: loading a first binary mask matrix, projecting a target product area, controlling a single-pixel detector to detect, determining a first measurement value, and performing lightweight reconstruction to obtain a first preview; generating a second binary mask matrix and performing projection and detection reconstruction by pre-checking the first preview, and performing multi-round iteration until an Nth preview is determined; calling the first preview to the Nth preview from a temporary database of an online detection platform, executing multi-layer compressed sensing and reconstruction, and determining a reconstruction result; and aiming at a reconstruction result, matching in an impurity feature library, and determining an impurity verification result. The technical problems of low impurity detection efficiency and insufficient detection accuracy in the medicine production process in the prior art are solved, and the technical effect of efficient and accurate online detection of medicine impurities in the production link is achieved.
Owner:NANTONG MEDICAL DEVICES

CoT-based multi-source remote sensing image ship target real-time identification and retrieval method

The invention discloses a CoT-based ship target real-time identification and retrieval method for a multi-source remote sensing image. The method comprises the following steps: collecting and preprocessing multi-source remote sensing image data in the ocean field; a ship target detection model is trained, an RT-DETR model is used for target detection and labeling, and a ship target area is segmented; generating a ship target description text through the large language model thinking chain; finely adjusting the CLIP model based on the segmented ship target area; and completing cross-modal retrieval according to the text retrieval image, marking a specific position in the multi-source remote sensing image, and performing real-time positioning. According to the method, through the fine-tuned CLIP model, accurate semantic association is established between the image and the text, the accuracy of cross-modal retrieval is improved, the method better adapts to special requirements and data distribution in the ocean field, related image blocks are rapidly matched, when a new remote sensing image arrives, a detection result is automatically updated, and real-time positioning and labeling are carried out, so that the accuracy of cross-modal retrieval is improved. And real-time monitoring of the ship target is realized.
Owner:SICHUAN UNIV