Cross-modal retrieval method and device, equipment and medium
By analyzing negative semantic texts with a large semantic model and obtaining difference results in the image database, the problem of negative semantic texts being unusable in cross-modal retrieval is solved, thereby improving retrieval efficiency and applicability.
Patent Information
- Application Number
- CN202411865420.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-09-12
AI Technical Summary
Existing technologies cannot effectively utilize text with negative semantics for cross-modal retrieval, resulting in retrieval results that cannot meet user needs in specific scenarios.
The retrieved text is analyzed through a pre-set semantic large model to obtain prompt sentences with positive and negative semantics, and the corresponding picture sets are obtained from the image database respectively, and the cross-modal retrieval results are obtained by using difference set calculation.
It realizes cross-modal retrieval of text with negative semantics, improves retrieval efficiency and quality, increases the applicability of cross-modal retrieval, and can adapt to the retrieval needs of various application fields.
Smart Images

Figure CN120632137A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to data processing technology, and specifically to a cross-modal retrieval method, apparatus, device, and medium. Background Art
[0002] Cross-modal retrieval refers to the task of retrieving relevant information from one modality (e.g., text, image) based on requests from different modalities. It involves processing the similarity between the content of different forms of data.
[0003] Using semantic methods to retrieve images with the same semantic meaning within large-scale image databases is a part of cross-modal retrieval. This retrieval method has achieved significant success in areas such as urban governance and smart transportation. For example, it can search for phrases such as "elderly person wearing a white mask" or "woman wearing red clothes driving a convertible" in image databases.
[0004] However, due to semantic gaps, differences in feature extraction, cross-modal matching difficulties, and technical and dataset limitations, existing technologies cannot perform cross-modal search using text with negative semantics. This results in search results that fail to meet user needs in specific scenarios, such as urban governance and smart transportation. Summary of the Invention
[0005] The embodiments of the present application provide a cross-modal retrieval method, apparatus, device, and medium to solve the problem that the prior art cannot perform cross-modal retrieval using text with negative semantics.
[0006] In a first aspect, an embodiment of the present application provides a cross-modal retrieval method, the method comprising:
[0007] receiving a search text;
[0008] Performing semantic analysis on the search text using a pre-set semantic model to obtain prompt sentences with positive semantics and prompt sentences with negative semantics in the search text, wherein the prompt sentences with negative semantics are expressed in positive form;
[0009] Obtaining, from an image database, a first picture set corresponding to the prompt sentence with positive semantics and a second picture set corresponding to the prompt sentence with negative semantics;
[0010] A cross-modal retrieval result is obtained according to a difference set between the first picture set and the second picture set.
[0011] In a second aspect, an embodiment of the present application provides a cross-modal retrieval device, the device comprising:
[0012] A receiving module, used for receiving a search text;
[0013] a semantic analysis module, configured to perform semantic analysis on the search text using a preset semantic model, and obtain prompt sentences with positive semantics and prompt sentences with negative semantics in the search text, wherein the prompt sentences with negative semantics are expressed in positive form;
[0014] A retrieval module is used to obtain, from an image database, a first image set corresponding to the prompt sentence with positive semantics and a second image set corresponding to the prompt sentence with negative semantics;
[0015] The retrieval result generation module is used to obtain a cross-modal retrieval result based on the difference between the first image set and the second image set.
[0016] In a third aspect, an embodiment of the present application provides an electronic device comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any embodiment of the first aspect.
[0017] In a fourth aspect, an embodiment of the present application provides a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the method described in any embodiment of the first aspect.
[0018] The cross-modal retrieval method, apparatus, equipment and medium provided by the embodiment of the present application, since the prompt statement with negative semantics is expressed in an affirmative form, the technical solution provided by the embodiment of the present application can not only retrieve the first picture set corresponding to the prompt statement with positive semantics in the image database, but also retrieve the second picture set corresponding to the prompt statement with negative semantics, and obtain the cross-modal retrieval results based on the difference between the first picture set and the second picture set, thereby achieving the purpose of cross-modal retrieval of the retrieval text containing negative speech. The problem that the technology cannot use text with negative semantics for cross-modal retrieval is solved. Since the embodiment of the present application can perform cross-modal retrieval on text with negative semantics, the cross-modal retrieval can adapt to the retrieval needs of various application fields, improve the efficiency and quality of cross-modal retrieval, and increase the applicability of cross-modal retrieval. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Other features, objects and advantages of the present application will become more apparent upon reading the detailed description of non-limiting embodiments made with reference to the following drawings:
[0020] Figure 1 is a flowchart of an embodiment of the cross-modal retrieval method of the present application;
[0021] Figure 2 is a flowchart of another embodiment of the cross-modal retrieval method of the present application;
[0022] Figure 3 yes Figure 2 The flowchart of step 201 in the cross-modal retrieval method of the present application is shown;
[0023] Figure 4 This is a flowchart of an application example of the cross-modal retrieval method of the present application;
[0024] Figure 5 This is a schematic structural diagram of an embodiment of a cross-modal retrieval device of the present application;
[0025] Figure 6 It is a structural diagram of an electronic device used to implement an embodiment of the present application. DETAILED DESCRIPTION
[0026] All actions of acquiring signals, information or data in this application are carried out in compliance with the relevant data protection laws and policies of the country where they are located and with the authorization given by the owner of the corresponding device.
[0027] The present application will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the relevant invention and are not intended to limit the invention. It should also be noted that, for ease of description, only portions relevant to the relevant invention are shown in the accompanying drawings.
[0028] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0029] Please refer to Figure 1 , which illustrates a process 100 according to one embodiment of a cross-modal search method of the present application. This cross-modal search method can be applied to various electronic devices with data processing capabilities. For example, these electronic devices may include, but are not limited to, cloud servers, physical servers, etc. The execution entity of this cross-modal search method may be a processor in these electronic devices.
[0030] like Figure 1 As shown, the cross-modal retrieval method includes the following steps:
[0031] Step 101: Receive search text.
[0032] The search text is text information input by the user for cross-modal search. In this embodiment, the search text contains negative words. For example, the search text can be: searching for "an old man without a white mask" in the image database.
[0033] Step 102: Perform semantic analysis on the search text using a pre-set semantic model to obtain prompt sentences with positive semantics and prompt sentences with negative semantics in the search text, wherein prompt sentences with negative semantics are expressed in positive form.
[0034] This embodiment does not specifically limit the semantic big model. In actual use, any model that can perform semantic analysis can be used.
[0035] To enable the big model to output prompt sentences with positive and negative semantics in the search text, and to express negative semantics in positive form, in this embodiment, after the search text is input into the semantic big model, output result condition information can also be input into the semantic big model. For example, the output result condition information can be: Please extract prompt sentences with positive and negative semantics from this sentence, and answer them in a positive form.
[0036] Step 103: Obtain from the image database a first picture set corresponding to the prompt sentence with positive semantics and a second picture set corresponding to the prompt sentence with negative semantics.
[0037] In this embodiment, step 103 can obtain the first picture set and the second picture set from the image database through any semantic-based cross-modal retrieval method, such as a deep learning-based method and a hash representation-based method, which will not be described in detail here.
[0038] Step 104: Obtain cross-modal retrieval results based on the difference between the first image set and the second image set.
[0039] In this embodiment, step 104 can specifically obtain the cross-modal search result target_res through the following formula (1):
[0040] target_res=p_res-n_res(1)
[0041] Among them, p_res is the first picture set, and n_res is the second picture set.
[0042] In this embodiment, both the first picture set and the second picture set may include multiple pictures.
[0043] The cross-modal retrieval method provided by the embodiment of the present application, since the prompt sentence with negative semantics is expressed in an affirmative form, enables the technical solution provided by the embodiment of the present application to not only retrieve the first picture set corresponding to the prompt sentence with positive semantics in the image database, but also retrieve the second picture set corresponding to the prompt sentence with negative semantics, and obtain the cross-modal retrieval result based on the difference between the first picture set and the second picture set, thereby achieving the purpose of cross-modal retrieval of the retrieval text containing negative speech. The problem that the technology cannot use text with negative semantics for cross-modal retrieval is solved. Since the embodiment of the present application can perform cross-modal retrieval on text with negative semantics, the cross-modal retrieval can adapt to the retrieval needs of various application fields, improve the efficiency and quality of cross-modal retrieval, and increase the applicability of cross-modal retrieval.
[0044] Please refer to Figure 2 , which illustrates process 200 of another embodiment of a cross-modal search method according to the present application. This cross-modal search method can be applied to various electronic devices with data processing capabilities. For example, these electronic devices may include, but are not limited to, cloud servers, physical servers, etc. The execution entity of this cross-modal search method may be a processor in these electronic devices.
[0045] like Figure 2 As shown, the cross-modal retrieval method includes the following steps:
[0046] Step 201: Segment the image to be retrieved to obtain images with instance targets as units.
[0047] In this embodiment, if Figure 3 As shown, step 201 may include:
[0048] Step 301: Use the fast-SAM model to segment the image used for retrieval to obtain a segmented image.
[0049] In this embodiment, the fast-SAM model is a lightweight image segmentation model that uses YOLOv8-seg as a base model for full instance segmentation. The model generates segmentation masks for all instances in the image.
[0050] Step 302: Use the monitoring data and segmentation data of the field where the image to be retrieved is located to fine-tune the segmented image, fit the segmented area, and obtain an image with instance targets as units.
[0051] In this embodiment, the fields in which the images used for retrieval are located may include multiple fields, such as the smart transportation field and the smart city field, etc., and no specific limitation is made here on the fields in which the images used for retrieval are located.
[0052] Step 201 utilizes the large, universal visual segmentation model and the image towers in the multimodal model to achieve fine-grained image storage. This process, leveraging universal segmentation methods, eliminates the need to train the model in a closed domain and successfully parses instantiated objects within the image.
[0053] Step 202: Obtain image feature representations of instance targets in the image.
[0054] In this embodiment, step 202 may include: using a pre-set image feature extraction module to obtain an image feature representation of the instance target; wherein the image feature extraction module is based on a contrastive language-image pretraining (CLIP) model and uses a low-rank adaptation (LongRange, LoRA) of a large language model for fine-tuning.
[0055] In this example, CLIP is a cross-modal retrieval model that learns the connection between vision and language by jointly training images and text. The CLIP model enables bidirectional retrieval between images and text.
[0056] In this example, LoRA is a low-rank adaptation method for large language models that can fine-tune the model by introducing trainable parameters without changing the model weights. In this example, LoRA can be used to fine-tune the CLIP model to better suit specific tasks, such as instance-based object recognition in smart transportation and smart cities.
[0057] In this embodiment, the image feature extraction module uses image-text pairs for model training, wherein the text in the image-text pair is generated by an image description (caption) generation model.
[0058] In this embodiment, once the image features of an instance object are extracted, a caption model can be used to generate a text description. This description can be a detailed explanation of the instance object in the image or a summary of the entire scene. By associating the text description with its corresponding image, an image-text pair is obtained.
[0059] Step 203: Establish an image database based on the image feature representation of the instance target in the picture.
[0060] In this embodiment, the image database includes not only the pictures used for retrieval, but also the image feature representations of instance targets in the pictures and image-text pairs, so that subsequent steps can obtain more accurate retrieval results when performing cross-modal retrieval.
[0061] Step 204: Input the image feature representation of the instance target into the Milvus vector database, and establish an offline retrieval index for the image database.
[0062] In this embodiment, Milvus is a vector database specifically designed to support AI applications. It efficiently stores, indexes, and manages large-scale embedding vectors generated by deep neural networks and other machine learning models. Milvus offers high performance, high availability, and easy scalability, easily handling trillion-level vector indexing tasks.
[0063] Step 205: Receive the search text.
[0064] The search text is text information input by the user for cross-modal search. In this embodiment, the search text contains negative words. For example, the search text can be: searching for "an old man without a white mask" in the image database.
[0065] Step 206: Perform semantic analysis on the search text using a preset semantic model to obtain prompt sentences with positive semantics and prompt sentences with negative semantics in the search text, wherein prompt sentences with negative semantics are expressed in positive form.
[0066] This embodiment does not specifically limit the semantic big model. In actual use, any model that can perform semantic analysis can be used.
[0067] To enable the big model to output prompt sentences with positive and negative semantics in the search text, and to express negative semantics in positive form, in this embodiment, after the search text is input into the semantic big model, output result condition information can also be input into the semantic big model. For example, the output result condition information can be: Please extract prompt sentences with positive and negative semantics from this sentence, and answer them in a positive form.
[0068] Step 206 analyzes the input semantic prompts using a large semantic model to obtain prompts with positive semantics and prompts with negative semantics that describe the content. This method cleverly analyzes prompts with negative semantics into prompts with positive and negative labels, and uses a multimodal text tower to represent these features. This method is advantageous in leveraging existing cross-modal frameworks, transforming prompts into a foundation for cross-modal search with negative semantics.
[0069] Step 207 : perform similarity calculations on the feature representations of the prompt sentences with positive semantics and the prompt sentences with negative semantics with the feature data in the Milvus vector database to obtain a first picture set and a second picture set.
[0070] Step 208: Obtain cross-modal retrieval results based on the difference between the first image set and the second image set.
[0071] In this embodiment, step 208 can specifically obtain the cross-modal retrieval result through the above formula (1).
[0072] The cross-modal retrieval method provided by the embodiment of the present application, since the prompt sentence with negative semantics is expressed in an affirmative form, enables the technical solution provided by the embodiment of the present application to not only retrieve the first picture set corresponding to the prompt sentence with positive semantics in the image database, but also retrieve the second picture set corresponding to the prompt sentence with negative semantics, and obtain the cross-modal retrieval result based on the difference between the first picture set and the second picture set, thereby achieving the purpose of cross-modal retrieval of the retrieval text containing negative speech. The problem that the technology cannot use text with negative semantics for cross-modal retrieval is solved. Since the embodiment of the present application can perform cross-modal retrieval on text with negative semantics, the cross-modal retrieval can adapt to the retrieval needs of various application fields, improve the efficiency and quality of cross-modal retrieval, and increase the applicability of cross-modal retrieval.
[0073] In order to enable those skilled in the art to more clearly understand the cross-modal retrieval method provided in the above embodiment, Figure 4 The specific example shown is used for explanation.
[0074] See for example Figure 4 As shown, the cross-modal retrieval method provided in the embodiment of the present application may include:
[0075] 1. Offline warehousing process
[0076] First, input the image;
[0077] Secondly, the image is segmented through the detection / segmentation module;
[0078] Third, the segmented image is passed through the image-feature module to extract image feature representation and then stored in the database to establish a database.
[0079] 2. Online search process
[0080] First, receive the search text entered by the user.
[0081] like Figure 4 As shown, in this embodiment, there are two search texts, namely: 1. An old man without a mask; 2. A non-red fire extinguisher.
[0082] Secondly, the search text is semantically analyzed through a pre-set semantic model to obtain prompt sentences with positive semantics and prompt sentences with negative semantics in positive expressions.
[0083] In this embodiment, the search text is: an old man without a mask, the prompt statement with positive semantics is: an old man, and the prompt statement with negative semantics in a positive expression is: an old man wearing a mask; the search text is: a non-red fire extinguisher, the prompt statement with positive semantics is: fire extinguisher, and the prompt statement with negative semantics in a positive expression is: a red fire extinguisher.
[0084] Third, the prompt sentences with positive semantics and the prompt sentences with negative semantics in positive expressions are extracted with text feature representations through the text-feature module, and the database is searched according to the text feature representations to obtain the forward search results corresponding to the prompt sentences with positive semantics and the reverse search results corresponding to the prompt sentences with negative semantics in positive expressions.
[0085] Fourth, based on the forward search results and the reverse search results, the final search results are obtained: an old man without a mask and a non-red fire extinguisher.
[0086] Please refer to Figure 5 As an implementation of the methods shown in the figures, the present application provides an embodiment of a cross-modal retrieval device. The device embodiment corresponds to the method shown in the above embodiments, and the device can be specifically applied to various electronic devices.
[0087] like Figure 5 As shown, the cross-modal search device 500 of this embodiment includes:
[0088] Receiving module 501, for receiving a search text;
[0089] The search text is text information input by the user for cross-modal search. In this embodiment, the search text contains negative words. For example, the search text can be: searching for "an old man without a white mask" in the image database.
[0090] Semantic analysis module 502, configured to perform semantic analysis on the search text using a preset semantic model to obtain prompt sentences with positive semantics and prompt sentences with negative semantics in the search text, wherein the prompt sentences with negative semantics are expressed in positive form;
[0091] This embodiment does not specifically limit the semantic big model. In actual use, any model that can perform semantic analysis can be used.
[0092] To enable the big model to output prompt sentences with positive and negative semantics in the search text, and to express negative semantics in positive form, in this embodiment, after the search text is input into the semantic big model, output result condition information can also be input into the semantic big model. For example, the output result condition information can be: Please extract prompt sentences with positive and negative semantics from this sentence, and answer them in a positive form.
[0093] A retrieval module 503 is configured to obtain, from an image database, a first image set corresponding to the prompt sentence with positive semantics and a second image set corresponding to the prompt sentence with negative semantics;
[0094] In this embodiment, the retrieval module 503 can obtain the first picture set and the second picture set from the image database through any semantic-based cross-modal retrieval method, such as: a deep learning-based method and a hash representation-based method, etc., which will not be repeated here.
[0095] The retrieval result generating module 504 is configured to obtain a cross-modal retrieval result based on the difference between the first image set and the second image set.
[0096] In this embodiment, the search result generation module 504 can obtain the cross-modal search result through the above-mentioned formula (1).
[0097] Optionally, the cross-modal search apparatus 500 of this embodiment may further include:
[0098] The image database creation module is used to segment the pictures used for retrieval to obtain pictures with instance targets as units; obtain image feature representations of the instance targets in the pictures; and establish the image database based on the image feature representations of the instance targets in the pictures.
[0099] Optionally, the image database creation module is also used to use the fast-SAM model to segment the pictures used for retrieval to obtain segmented pictures; use the monitoring data and segmentation data of the field where the pictures used for retrieval are located to fine-tune the segmented pictures, fit the segmented areas, and obtain pictures with instance targets as units.
[0100] Optionally, the image database creation module is also used to obtain the image feature representation of the instance target using a pre-set image feature extraction module; wherein the image feature extraction module is based on the comparative language-image pre-trained CLIP model and uses the low-rank adaptation LoRA of the large language model for fine-tuning.
[0101] Optionally, the image feature extraction module uses image-text pairs for model training, wherein the text in the image-text pair is generated by an image description caption generation model.
[0102] Optionally, the image database creation module is further configured to input the image feature representation of the instance target into a Milvus vector database, and to establish an offline retrieval index for the image database.
[0103] Optionally, the retrieval module 403 is further configured to perform similarity calculations on the feature representations of the prompt sentence with positive semantics and the prompt sentence with negative semantics with feature data in the Milvus vector database to obtain the first picture set and the second picture set.
[0104] The cross-modal retrieval device provided by the embodiment of the present application, since the prompt sentence with negative semantics is expressed in an affirmative form, enables the technical solution provided by the embodiment of the present application to not only retrieve the first picture set corresponding to the prompt sentence with positive semantics in the image database, but also retrieve the second picture set corresponding to the prompt sentence with negative semantics, and obtain the cross-modal retrieval result based on the difference between the first picture set and the second picture set, thereby achieving the purpose of cross-modal retrieval of the retrieval text containing negative speech. The problem that the technology cannot use text with negative semantics for cross-modal retrieval is solved. Since the embodiment of the present application can perform cross-modal retrieval on text with negative semantics, the cross-modal retrieval can adapt to the retrieval needs of various application fields, improve the efficiency and quality of cross-modal retrieval, and increase the applicability of cross-modal retrieval.
[0105] Reference below Figure 6 , which shows a structural schematic diagram of an electronic device for implementing some embodiments of the present application. Figure 6 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0106] like Figure 6 As shown, the electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. Various programs and data required for the operation of the electronic device 600 are also stored in the RAM 603. The processing device 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0107] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic disk, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 6 The electronic device 600 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead. Figure 6 Each block shown in the figure may represent one device, or may represent multiple devices as needed.
[0108] In particular, according to some embodiments of the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In some such embodiments, the computer program can be downloaded and installed from a network via the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the method of some embodiments of the present application are performed.
[0109] It should be noted that the computer-readable medium described in some embodiments of the present application may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present application, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or device. In some embodiments of the present application, the computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0110] In some embodiments, the client and server can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0111] The computer-readable medium may be included in the electronic device, or may exist independently and not be incorporated into the electronic device. The computer-readable medium carries one or more programs. When executed by the electronic device, the electronic device: receives a search text; performs semantic analysis on the search text using a pre-set semantic model to obtain prompt statements with positive semantics and prompt statements with negative semantics from the search text, wherein the prompt statements with negative semantics are expressed in a positive form; obtains a first set of images corresponding to the prompt statements with positive semantics and a second set of images corresponding to the prompt statements with negative semantics from an image database; and obtains cross-modal search results based on the difference between the first set of images and the second set of images.
[0112] Computer program code for performing the operations of some embodiments of the present application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++; and also conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, or can be connected to an external computer (for example, through the Internet using an Internet service provider). The above network includes a local area network (LAN) or a wide area network (WAN).
[0113] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0114] The units described in some embodiments of this application may be implemented in software or hardware. The units described may also be provided in a processor. For example, a processor may be described as comprising a first determination unit, a second determination unit, a selection unit, and a third determination unit. The names of these units do not, in some cases, limit the units themselves.
[0115] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0116] The above description is only an illustration of some preferred embodiments of the present application and the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present application is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features and the technical features with similar functions disclosed in the embodiments of the present application (but not limited to) are replaced with each other to form a technical solution.
Claims
1. A cross-modal retrieval method, characterized in that: The method comprises: receiving a search text; Performing semantic analysis on the search text using a pre-set semantic model to obtain prompt sentences with positive semantics and prompt sentences with negative semantics in the search text, wherein the prompt sentences with negative semantics are expressed in positive form; Obtaining, from an image database, a first picture set corresponding to the prompt sentence with positive semantics and a second picture set corresponding to the prompt sentence with negative semantics; A cross-modal retrieval result is obtained according to a difference set between the first picture set and the second picture set.
2. The method according to claim 1, characterized in that The method further comprises: Segment the image used for retrieval to obtain images with instance targets as units; Obtaining an image feature representation of the instance target in the image; The image database is established based on the image feature representation of the instance target in the picture.
3. The method according to claim 2, characterized in that The segmentation process of the image for retrieval to obtain the image with instance targets as units includes: Use the fast-SAM model to segment the image used for retrieval and obtain the segmented image; The segmented image is fine-tuned using the monitoring data and segmentation data of the field where the image for retrieval is located, the segmented area is fitted, and an image with instance targets as units is obtained.
4. The method according to claim 2, characterized in that The obtaining of the image feature representation of the instance target in the picture includes: A pre-set image feature extraction module is used to obtain the image feature representation of the instance target; wherein, the image feature extraction module is based on the comparative language-image pre-trained CLIP model and is fine-tuned using the low-rank adaptation LoRA of a large language model.
5. The method according to claim 4, characterized in that The image feature extraction module uses image-text pairs for model training, wherein the text in the image-text pairs is generated by an image description generation caption model.
6. The method according to claim 2, characterized in that After characterizing the image features of the instance target in the picture, the method further includes: The image feature representation of the instance target is input into the Milvus vector database, and a retrieval index of the image database is established offline.
7. The method according to claim 2, characterized in that The acquiring, from the image database, a first picture set corresponding to the prompt statement with positive semantics and a second picture set corresponding to the prompt statement with negative semantics comprises: The feature representations of the prompt sentence with positive semantics and the prompt sentence with negative semantics are respectively subjected to similarity calculations with the feature data in the Milvus vector database to obtain a first picture set and a second picture set.
8. A cross-modal retrieval device, characterized in that: The device comprises: A receiving module, used for receiving a search text; a semantic analysis module, configured to perform semantic analysis on the search text using a preset semantic model, and obtain prompt sentences with positive semantics and prompt sentences with negative semantics in the search text, wherein the prompt sentences with negative semantics are expressed in positive form; A retrieval module is used to obtain, from an image database, a first image set corresponding to the prompt sentence with positive semantics and a second image set corresponding to the prompt sentence with negative semantics; The retrieval result generation module is used to obtain a cross-modal retrieval result based on the difference between the first image set and the second image set.
9. An electronic device, characterized in that: include: one or more processors; a storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.
10. A computer-readable medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.