Image search method, related device and computer program product

By extracting target images from video frames as search results based on semantic matching, the problem of limited image sources in image databases is solved, an efficient and resource-saving image search service is implemented, and the quality and recall capability of image search are improved.

CN119166850BActive Publication Date: 2025-09-09SHANGHAI BILIBILI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411304229.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-18
Publication Date
2025-09-09
Estimated Expiration
2044-09-18

AI Technical Summary

Technical Problem

Existing technologies make it difficult to provide users with image search services efficiently and save computing resources while ensuring search quality, especially when the image sources in the image database are limited, making it difficult to provide enough images of the same type for training image classification models.

Method used

Through content-based semantic matching, after determining that the semantic similarity of the target video meets the requirements, the target image is extracted from the video frame as the search result, and the video is used as the source of image material to enrich the search and recall capabilities of the image.

Benefits of technology

While ensuring search quality, it improves the efficiency of image search, saves computing resources, and enhances the source richness of image materials and the ability to recall similar images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119166850B_ABST
    Figure CN119166850B_ABST
Patent Text Reader

Abstract

The present application provides a method, related apparatus, and computer program product for searching images. The method determines a target video based on the first content of a sample image used to search for semantically similar images, wherein the semantic similarity between the second content of the target video and the first content is greater than or equal to a first similarity threshold; in response to being able to determine the target image from a video frame of the target video, the target image is used as a semantically similar image search result for the sample image, wherein the semantic similarity between the third content of the sample image and the first content is greater than or equal to a second similarity threshold, and the value of the second similarity threshold is greater than or equal to the first similarity threshold. In this way, the search capability and recall capability for similar images can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method and device for searching images, an electronic device, a computer-readable medium, and a computer program product. Background Art

[0002] As society develops, computer technology is also making progress. At the same time, image search technology is an important branch of computer technology in the field of computer vision. Image search technology can identify, match and retrieve related and similar images by analyzing image content.

[0003] Image search technology allows users to enter a "reference" image and then provide search results with other images that meet a certain degree of similarity to the reference image. Therefore, how to provide search services more efficiently and with less computational resources while ensuring search quality is a pressing need. Summary of the Invention

[0004] Various aspects of the present application provide a method, apparatus, electronic device, computer-readable storage medium, and computer program product for searching images, which can determine, based on semantic matching of content, videos whose content meets the semantic difference between that included and that included in sample images. Then, based on the method of "splitting" the video, the video is used as a source to provide and mine search results for semantically similar images that can serve as sample images. In this way, videos can be used to enrich the source of image materials, improve the search and recall capabilities for similar images, and provide search services to users more efficiently and with less computing resources while ensuring search quality.

[0005] In one aspect of the present application, a method for searching images is provided, comprising: determining a target video based on a first content of a sample image used to search for semantically similar images, wherein the semantic similarity between the second content of the target video and the first content is greater than or equal to a first similarity threshold; in response to being able to determine the target image from a video frame of the target video, using the target image as a semantically similar image search result for the sample image, wherein the semantic similarity between the third content of the sample image and the first content is greater than or equal to a second similarity threshold, and the value of the second similarity threshold is greater than or equal to the first similarity threshold.

[0006] In another aspect of the present application, a device for searching for images is provided, comprising: a target video determination module, configured to determine a target video based on a first content of a sample image for searching for semantically similar images, wherein the semantic similarity between the second content of the target video and the first content is greater than or equal to a first similarity threshold; a search result determination module, configured to, in response to being able to determine a target image from a video frame of the target video, use the target image as a semantically similar image search result for the sample image, wherein the semantic similarity between the third content of the sample image and the first content is greater than or equal to a second similarity threshold, and the value of the second similarity threshold is greater than or equal to the first similarity threshold.

[0007] Another aspect of the present application provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method for searching images provided above.

[0008] In another aspect of the present application, a computer-readable storage medium is provided, on which computer program instructions are stored. The computer program instructions can be executed by a processor to implement the image searching method provided above.

[0009] In another aspect of the present application, a computer program product includes a computer program having computer program instructions stored thereon. When the computer program is executed by a processor, the method for searching images as provided above can be implemented.

[0010] In the solution provided by the embodiment of the present application, a target video is determined based on the first content of a sample image used to search for semantically similar images, wherein the semantic similarity between the second content of the target video and the first content is greater than or equal to a first similarity threshold; in response to being able to determine the target image from the video frame of the target video, the target image is used as a semantically similar image search result for the sample image, wherein the semantic similarity between the third content of the sample image and the first content is greater than or equal to a second similarity threshold, and the value of the second similarity threshold is greater than or equal to the first similarity threshold. This method can determine, based on the semantic matching of content, a video whose content meets the semantic difference requirements with the content included in the sample image. Then, based on the method of "splitting" the video, the video is used as a source to provide and mine search results for semantically similar images that can serve as sample images. In this way, video can be used to enrich the source of image material, improve the search and recall capabilities for similar images, and provide search services to users more efficiently and save computing resources while ensuring search quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application, a brief introduction will be given below to the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0012] Other features, objects and advantages of the present application will become more apparent upon reading the detailed description of non-limiting embodiments made with reference to the following drawings:

[0013] Figure 1 A schematic diagram of an image search process provided in an embodiment of the present application;

[0014] Figure 2 A schematic diagram of an architecture of a process that can be used for searching images in an application scenario provided by an embodiment of the present application;

[0015] Figure 3 A schematic diagram of the structure of an apparatus for searching images provided in an embodiment of the present application;

[0016] Figure 4 The figure is a schematic diagram of the structure of an electronic device suitable for implementing the solution in the embodiment of the present application.

[0017] The same or similar reference numerals in the drawings represent the same or similar components. DETAILED DESCRIPTION

[0018] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0019] In a typical configuration of the present application, the terminal and the equipment of the service network each include one or more processors (CPUs), input / output interfaces, network interfaces and memories.

[0020] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0021] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology for information storage. The information can be computer program instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc-read only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission medium that can be used to store information that can be accessed by a computing device.

[0022] As discussed above, how to provide search services to users more efficiently and save computing resources while ensuring search quality is worthy of attention and an urgent need.

[0023] In some solutions, the choice is to configure as many candidate images as possible in, for example, image databases or image datasets, and to add more dimensional annotation information and auxiliary information to the candidate images in the image database to expand the candidate images and enhance their usability. This is done in the hope of providing a higher-quality search service.

[0024] However, in this way, due to the limited image sources and the difficulty in pre-maintaining annotation information and auxiliary information for each image as the number of images in the image database increases and expands, the number of images that can actually be maintained in some scenes may not meet the needs and cannot provide users with similar pictures in "sufficient" quantities.

[0025] For example, in the scenario of training an image classification model, a large number of images of the "same type" may be required for training the image classification model. To this end, it may be desirable to use, for example, a "basic image" to search and recall a large number of images similar to the "basic image" in order to construct a sample image set as model input for training the image classification model. However, in the strategy provided by the above-described method, it may be difficult to maintain a large number of images of the same type in advance in the image database or image dataset, resulting in difficulty in providing sufficient images to construct a training sample set. This, in turn, affects the training effect of the classification model due to a lack of samples or difficulty in obtaining training samples during the training process of the image classification model.

[0026] In this regard, an embodiment of the present application provides a method for searching images. The method determines a target video based on the first content of a sample image used to search for semantically similar images, wherein the semantic similarity between the second content of the target video and the first content is greater than or equal to a first similarity threshold. In response to being able to determine the target image from the video frame of the target video, the target image is used as a semantically similar image search result for the sample image, wherein the semantic similarity between the third content of the sample image and the first content is greater than or equal to a second similarity threshold, and the value of the second similarity threshold is greater than or equal to the first similarity threshold. This method can determine, based on semantic matching of content, a video whose content meets the semantic difference requirements with the content included in the sample image. Then, based on a "split" video method, the video is used as a source to provide and mine search results for semantically similar images that can serve as sample images. In this way, video can be used to enrich the source of image material, improve the search and recall capabilities for similar images, and provide search services to users more efficiently and save computing resources while ensuring search quality.

[0027] In practical scenarios, the execution subject of this method can be a user device, or a device formed by integrating a user device and a network device via a network, or an application running on such a device. User devices include but are not limited to computers, mobile phones, tablets, smart watches, wristbands, and other terminal devices. Network devices include but are not limited to network hosts, single network servers, multiple network server clusters, or cloud computing-based computer clusters. Here, the cloud is composed of a large number of hosts or network servers based on cloud computing. Cloud computing is a type of distributed computing, consisting of a group of loosely coupled computers forming a virtual computer.

[0028] When the execution subject is software, it can be installed in the electronic devices listed above. It can be implemented as multiple software or software modules, or as a single software or software module, and is not specifically limited here.

[0029] Figure 1 A process 100 for searching an image according to an embodiment of the present application is shown. The process 100 includes at least the following processing steps:

[0030] (Step) S101 : Determine a target video based on first content of a sample image for searching for semantically similar images.

[0031] In the embodiments of the present application, in combination with different application scenarios, the execution subject may obtain sample images used as a "benchmark" for searching for semantically similar images in different ways. For example, in the case where the execution subject is a server of a platform capable of providing image search services, the sample image may be communicated to the execution subject by a user through a terminal device used by the user (for example, when the user requests the execution subject to search and provides other images with "similar content" to the sample image).

[0032] For another example, in the scenario described above where it is desired to construct an input image sample set for training an image classification model, the execution entity may also, based on the training purpose, determine images that can be used to train the image classification model from, for example, a locally maintained (sample) image database or image dataset, as "benchmark" sample images for expanding the images in the image database or image dataset. For example, these may include "images" with clearly defined classification labels that meet the training requirements.

[0033] Accordingly, the execution entity can search and determine the target video based on the content of the sample image (described as the first content for ease of understanding). The semantic similarity between the content of the target video (described as the second content for ease of understanding) and the first content is greater than a (preset) first similarity threshold. In some embodiments, the "content" of the image or video can be the content of the image or video expressed in natural language.

[0034] In some embodiments, the value of the first similarity threshold can be determined based on the standard of considering that two contents simultaneously include the "same main content" (i.e., based on the value corresponding to the minimum semantic similarity allowed when it can be reliably considered that two contents simultaneously point to the "same content" in terms of main content).

[0035] For example, if the sample image indicated by the first content contains "Subject A did XXX," and the video content indicated by the second content corresponding to video I contains "The video tells the story of Subject A doing XXX in the past," the execution entity can determine, based on a semantic comparison of the two contents, that the semantic similarity between the two contents meets the requirements (for example, this may be due to the fact that both "Subject A did XXX" is included, leading to the assumption that the first content and the second content at least refer to the "same content"). Accordingly, in this case, if the semantic difference between the two contents (or, in other words, between the two contents) meets the requirements of the first similarity threshold, that is, the semantic similarity between the two contents is greater than or equal to the first similarity threshold, the execution entity can determine that video I is the "target video."

[0036] Typically, both the first content and the second content may be in text form. This facilitates direct language processing and text processing to identify semantics, perform semantic comparison, and determine semantic similarity between the two.

[0037] In some embodiments, the first content may be generated instantly by the execution entity locally or using another device after acquiring the sample image (for example, by processing the sample image using an image content generation model deployed locally or on another device, for example, the image content generation model may determine the content of the image based on semantic analysis of the image). For example, the sample image may be processed using a deployed convolutional neural network, recurrent neural network, generative adversarial network, etc. to obtain its "content."

[0038] In some embodiments, the first content may also be obtained in advance by, for example, processing the sample image in a similar manner by another device. Thus, after the execution entity determines and obtains the sample image, it can directly read the first content (for example, the first content may be communicated to the execution entity together with the sample image).

[0039] In some embodiments, the second content can be generated by processing the title information and content introduction information of the target video using a content generation model. For example, the content generation model can be a recursive neural network (RNN), a convolutional neural network, or other models.

[0040] The content generation model can process the target video's title information and content introduction information (for example, combining the target video's title information and content introduction information to obtain a combined result, parsing the "semantics" of the "combined result" using the content generation model, and using the textual form of the "semantics" as the second content indicating the content of the target video). This allows the "target video's title information and content introduction information" to be used as a reference to simply and efficiently "understand" the content of the video and generate second content corresponding to the video's content.

[0041] Similarly, in some embodiments, if the first content needs to be generated by an execution entity, the execution entity may also similarly utilize such a content generation model to generate the first content, which will not be repeated here.

[0042] In some embodiments, to more accurately determine the second content of a (target) video, the second content corresponding to the video can be determined based on a combination of the fourth content of each video frame included in the video, that is, the combined image content of the target video. This allows for a more accurate determination of the second content of the video by actually "analyzing" the content of each video frame (or, in other words, the individual video images comprising the video).

[0043] Specifically, the content generation model can be configured to first process each video frame in the video and identify the fourth content of each video frame (for example, the fourth content can also be in text form). Then, the fourth content is combined to obtain the image composite content for the target video. For example, the fourth content in text form is spliced ​​together to obtain the image composite content expressed in text form.

[0044] Finally, similarly, by processing the semantics of the image combination (for example, analyzing the textual content of the image combination to determine the textual "semantics" of the "image combination content"), this "semantics" is used as the "secondary content" of the video. Thus, the execution entity can use the content generation model to determine the secondary content of the video from the perspective of "the content included in each video frame," thereby improving the quality of the secondary content.

[0045] In some embodiments, given that a video may actually be composed of modal data in multiple modalities, such as text, audio, and images, in order to further and more accurately determine the "content" of the video and determine higher-quality "secondary content," the process of generating the secondary content may include simultaneously referencing modal data in various modalities to determine the "secondary content."

[0046] Specifically, the target video can be split into a set of unimodal data, where each unimodal data in the set corresponds to only one data modality. For example, a set of modal data can consist of unimodal data Q, unimodal data W, and unimodal data E. Unimodal data Q, for example, corresponds to the text-based content of the video; unimodal data W, for example, corresponds to the audio-based content of the video; and unimodal data E, for example, corresponds to the image-based content of the video.

[0047] Taking into account the differences in recognition capabilities and performance of different modalities, in this method, "semantics" can be selected as a substitute for "content" for combination. In this way, interference with the combination result caused by differences in the way "content" is determined can be avoided. Accordingly, it is possible to choose to generate target modal data for each single modal data under the target modality, or in other words, "unify each single modal data into the target modality". In this way, subsequent analysis of modal data and determination of semantics can be performed in the modal space of the same "target modality", so as to avoid the processing differences caused by different modal spaces affecting the "semantic" recognition quality.

[0048] For example, when "text modality is used as the target modality," the target modal data corresponding to the unimodal data W and the unimodal data E in the "text modality" can be determined separately. For example, after converting image, audio, or other data into embedding vectors, the embedding vectors can be mapped from the original modal space to the target modal space of the text modality based on modal alignment to determine the target modal data corresponding to the unimodal data W and the unimodal data E in the "text modality."

[0049] Furthermore, the execution subject can similarly process the target modal data to obtain the corresponding target modal data semantics, and then combine the target modal data semantics of each target modal data to obtain the target modal data semantic combination result (for example, the target modal data semantics and the target modal data semantic combination result can also be in text form). Then, the execution subject uses the content generation model to process the target modal data semantic combination result to generate the second content (for example, the textual target modal data semantic combination result is subjected to "semantic recognition" again to finally determine the second content corresponding to the video). In this way, in the process of generating the second content, the modal data under each modality can be "fully" referenced, and the "content" of the video can be more accurately "analyzed and determined", thereby accurately generating the "second content".

[0050] Next, after the execution entity determines the target video, the execution entity may similarly detect whether the video frame of the target video includes a target image having a semantic similarity with the sample image greater than or equal to a second similarity threshold by comparing semantic similarities.

[0051] That is, the execution subject determines whether there is a “target image” in the video frame of the target video whose image semantic similarity with the sample image satisfies the requirement.

[0052] In an embodiment of the present disclosure, the value of the second similarity threshold is greater than or equal to the first similarity threshold. Thus, based at least on the screening criteria for the target video (e.g., the values ​​of the two similarity thresholds are equal), a "target image" is searched through a "fine screening" that is stricter than the "coarse screening" of the target video (e.g., the value of the second similarity threshold is greater than or equal to the first similarity threshold).

[0053] Therefore, after determining the target video whose semantic similarity between the second content and the first content is greater than or equal to the first similarity threshold, the determined "target video" can be used as a carrier to "mine" the target image from the video frame of the target video that meets the requirement that the semantic similarity between the second content and the first content is greater than or equal to the second similarity threshold.

[0054] Accordingly, this approach not only allows the "unsplit video" to be directly used as a source of material for image search and matching services, enriching the source of "candidate images", but also eliminates the need to pre-split the video, which can save the "waste of resources" that exists in some solutions that construct image datasets by pre-splitting the video (for example, pre-splitting a video that may not be needed into multiple images). Moreover, this approach does not require or require a large amount of image content annotation (for example, during preliminary screening, the content of the video can be referenced to "limit" the image range, and subsequently, images can be annotated and searched based on the "limited" range without having to annotate all images in the image database), which can further save computing and storage resources.

[0055] S102 : In response to being able to determine a target image from the video frame of the target video, using the target image as a semantically similar image search result for the sample image.

[0056] In the embodiment of the present application, based on the above S101, if the target image can be determined from the video frame of the target video, the execution entity responds by using the target image as the search result of semantically similar images to the sample image. This completes the "video-based" search and mining process for semantically similar images to the sample image.

[0057] Then, the image search method provided by the present application determines a target video based on the first content of a sample image used to search for semantically similar images, wherein the semantic similarity between the second content of the target video and the first content is greater than or equal to a first similarity threshold; in response to being able to determine the target image from the video frame of the target video, the target image is used as the search result for semantically similar images for the sample image, wherein the semantic similarity between the third content of the sample image and the first content is greater than or equal to a second similarity threshold, and the value of the second similarity threshold is greater than or equal to the first similarity threshold. This method can determine a video that meets the semantic similarity requirements with the sample image based on semantic matching. Then, based on the method of "splitting" the video, semantically similar images of the sample image are provided and mined. In this way, videos can be used to enrich the source of image materials, improve the search and recall capabilities for similar images, and provide users with search services more efficiently and save computing resources while ensuring search quality.

[0058] In some embodiments, images and videos may be further classified based on the perspective of "type". This allows for searching for images of the same type or different types that have a semantic similarity that meets the requirements with respect to the sample image, based on different usage requirements. For example, in the scenario of training an image classification model as described above, if the image classification model is expected to have the ability to classify images of the same type, the search target for the sample image may be "same type", while if the image classification model is expected to have the ability to classify images of multiple scenes and different types, the search target for the sample image may be "different type".

[0059] The "type" of images and videos can be exemplarily classified as "real scene", "cartoon scene" and so on. Therefore, different search needs can be further met by refining the "type". For example, when the above search target is "different type", or in other words, when the user expects to search "across types" (for example, hoping to search for the visual effect of the image content of the sample image of "real scene" under the "cartoon scene"), the user can indicate that the search target for the sample image is a semantically similar image search result of a different image type from the sample image when providing a sample image.

[0060] In some embodiments, the "type" of an image or video can be represented using a type label, and the type label for the sample image can be uploaded together with the sample image, or generated locally by the executing entity based on the user's instructions and search target (for example, based on an image category recognition model, the sample image is processed to determine its type label such as "real scene" or "cartoon scene").

[0061] Then, based on the search target of semantically similar image search results of different image types from the sample image, the execution entity determines a set of candidate videos based on the type label of the sample image. The type label associated with each candidate video in the set of candidate videos is different from the type label associated with the sample image. For example, if the sample image is a "real scene," each candidate video in the set of candidate videos can be a "cartoon scene" video. This enables the search to find a "target video" that meets the requirements among the various "cartoon scene" videos.

[0062] For example, as an alternative or alternative, the executing subject may select the first content based on the sample image during the execution of S101 to determine the target video from a set of candidate videos. Thus, in the scenario of training an image classification model, for example, "semantic similarity" can be used to provide and mine "cross-type" images as training samples, thereby enriching the training sample set and improving the classification effect of the training image classification model.

[0063] In some embodiments, the target video that can be determined from the search results can be provided to the user based on the user's needs for reference and decision-making assistance. For example, the user can communicate with the execution entity through the terminal device used by the user to request the target video by sending a video provision request for the target video to the execution entity.

[0064] Accordingly, when the execution subject receives a video provision request for a target video from a terminal device used by a user as a target device, for example, the execution subject can respond and provide the target video to the target device based on the communication path between the execution subject and the target device, thereby providing the target video to the user.

[0065] On the basis of any of the above embodiments, in order to improve the content generation capability and efficiency, a large language model may be used as a content generation model to generate the first content and / or the second content.

[0066] A Large Language Model (LLM) is an artificial intelligence model designed to understand and generate human language. Based on its understanding, a generative language model can perform processing operations to produce corresponding results.

[0067] LLMs are characterized by their large size and can typically include a large number of parameters to help them learn complex patterns in language data. These models are often based on deep learning architectures such as transformers, which helps them provide better processing performance on various NLP tasks.

[0068] As discussed above, in embodiments of the present application, the execution entity may also choose to utilize LLM to improve processing performance. For example, in the aforementioned example of generating the secondary content of a video, LLM can be used to process the video's title information and content summary information to generate the secondary content. Alternatively, LLM can be used to generate the secondary content based on modal data from multiple video modes.

[0069] Accordingly, the execution subject can instruct the LLM to generate the second content of the video based on the title information and content introduction information of the video based on pre-configured guide words, guide tags, etc. For example, in this process, the LLM can be based on, for example, "Please summarize the content of the video based on the title information and content introduction information of the video, and use the results of the summary as the second content of the video." Or, for the scenario of generating the second content with reference to video content in different modalities, the guide words can be, for example, "Please refer to the text, audio and image content of the video to generate the text summary corresponding to the video text content, audio and image content, and then summarize the content of the video based on the combination of the three summaries, and use the results of the summary as the second content of the video" and so on.

[0070] Similarly, LLM can be configured by default to omit "guide words." For example, to generate secondary content based on a video's title and synopsis, the LLM can, based on the default configuration, automatically understand the need to generate secondary content based on the video's title and synopsis. This default configuration allows the generative language model to stably and consistently determine the need to generate secondary content based on the video's title and synopsis.

[0071] Therefore, the LLM can be used to generate the second content more efficiently and with higher quality.

[0072] Similarly, in step S102 discussed above, the LLM can also be used to search for a target image within the target video's frames. For example, the LLM can be instructed to search for the target image within the target video's frames by constructing a prompt: "Search within the target video's frames for a target image whose semantic similarity to the sample image is greater than or equal to a second similarity threshold." This description will not be repeated here. Thus, the LLM can be directly used to fully implement the entire process of generating content, searching for images and videos, and finding the target image, simplifying the overall configuration.

[0073] To deepen understanding, this application also combines a specific application scenario and exemplarily gives a specific implementation solution, please refer to Figure 2 The architecture 200 of a process that can be used for searching images is shown by way of example.

[0074] Exemplarily, in the architecture, for video 220 and video 230, the execution entity can use LLM 240 to process video 220 and video 230 to obtain content 225 corresponding to video 220 (used to indicate the content of video 220) and content 235 corresponding to video 230 (used to indicate the content of video 230).

[0075] Then, after determining sample image 210 (e.g., as discussed above, this may be a user-provided image intended to serve as a "benchmark" for searching for semantically similar images), the execution entity uses content 215 of sample image 210 to determine whether the "target image" discussed above exists in videos 220 and 230. That is, the execution entity may determine the semantic difference between content 215 and content 225, and the semantic difference between content 215 and content 235, respectively.

[0076] For example, the semantic similarity between content 215 and content 225 can meet the requirements of the first similarity threshold and the semantic similarity is greater than or equal to the first similarity threshold. In this case, the execution entity searches the video frames of video 220 (e.g., video frames 220-1, 220-2, ..., 220-N, where N is a positive integer) to find a "target image" whose content has a semantic similarity greater than or equal to the second similarity threshold with the content of sample image 210.

[0077] Exemplarily, if the semantic similarity between the contents of video frame 220-1 and sample image 210 is greater than or equal to the second similarity threshold, the execution entity can determine that video frame 220-1 is the "target image" and use video frame 220-1 as the semantically similar image search result for sample image 210.

[0078] The present application also provides a device for searching images, the structure of which is as follows: Figure 3 The device 300 shown. The device 300 includes: a target video determination module 310, configured to determine a target video based on a first content of a sample image for searching for semantically similar images, wherein the semantic similarity between the second content of the target video and the first content is greater than or equal to a first similarity threshold; a search result determination module 320, configured to, in response to being able to determine a target image from a video frame of the target video, use the target image as a semantically similar image search result for the sample image, wherein the semantic similarity between the third content of the sample image and the first content is greater than or equal to a second similarity threshold, and the value of the second similarity threshold is greater than or equal to the first similarity threshold.

[0079] In some embodiments, the second content is generated by a content generation model by processing the title information and content introduction information of the target video.

[0080] In some embodiments, the second content is generated by processing an image combination content of the target video by a content generation model, where the image combination content is obtained by combining fourth content of video frames of the target video.

[0081] In some embodiments, the second content is generated based on the following method: splitting the target video into a set of unimodal data, wherein each unimodal data in a set of modal data corresponds to only one data modality; generating target modal data of each unimodal data under the target modality; combining the target modal data semantics of each target modal data to obtain a target modal data semantic combination result; and using a content generation model to process the target modal data semantic combination result to generate the second content.

[0082] In some embodiments, the content generation model includes a large language model.

[0083] In some embodiments, in response to the search results for semantically similar images whose search target for the sample image is a different image type from the sample image, the device 300 also includes: a candidate video determination module, configured to determine a group of candidate videos based on the type label of the sample image, wherein the type label associated with each candidate video in the group of candidate videos is different from the type label associated with the sample image; and the target video determination module 310 is further configured to determine the target video from the group of candidate videos based on the first content of the sample image.

[0084] In some embodiments, the apparatus 300 further includes a target video providing module configured to provide the target video to the target device based on a communication path with the target device in response to receiving a video providing request for the target video sent by the target device.

[0085] Based on the same inventive concept, embodiments of the present application also provide an electronic device, a readable storage medium, and a computer program product. The method corresponding to the electronic device may be the method for searching for images in the aforementioned embodiments, and its principle of solving the problem is similar to that of the method. The electronic device provided in embodiments of the present application includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the methods and / or technical solutions of the aforementioned multiple embodiments of the present application.

[0086] An electronic device can be a user device, or a device formed by integrating a user device and a network device via a network, or an application running on any of the above devices. User devices include, but are not limited to, computers, mobile phones, tablets, smart watches, wristbands, and other terminal devices. Network devices include, but are not limited to, network hosts, single network servers, multiple network server clusters, or cloud computing-based computer clusters, and can be used to implement some of the processing functions required for setting an alarm. Here, the cloud is composed of a large number of hosts or network servers based on cloud computing. Cloud computing is a type of distributed computing, consisting of a virtual computer composed of a group of loosely coupled computers.

[0087] Figure 4 The structure of an electronic device suitable for implementing the method and / or technical solution in the embodiment of the present application is shown. The electronic device 400 includes a central processing unit (CPU) 401, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 402 or the program loaded from the storage part 408 into the random access memory (RAM) 403. In RAM 403, various programs and data required for system operation are also stored. CPU 401, ROM 402 and RAM 403 are connected to each other through bus 405. Input / output (I / O) interface 404 is also connected to bus 405.

[0088] The following components are connected to the I / O interface 404: an input section 406 including a keyboard, a mouse, a touch screen, a microphone, an infrared sensor, etc.; an output section 407 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), an LED display, an OLED display, etc., and a speaker; a storage section 408 including one or more computer-readable media such as a hard disk, an optical disk, a magnetic disk, a semiconductor memory, etc.; and a communication section 409 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 409 performs communication processing via a network such as the Internet.

[0089] In particular, the methods and / or embodiments in the embodiments of the present application can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. When the computer program is executed by the central processing unit (CPU) 401, the above-mentioned functions defined in the method of the present application are performed.

[0090] Another embodiment of the present application further provides a computer-readable storage medium and a computer program product, on which computer program instructions are stored. The computer program instructions can be executed by a processor to implement the methods and / or technical solutions of any one or more embodiments of the present application.

[0091] Specifically, the present embodiment can adopt any combination of one or more computer-readable media.Computer-readable media can be computer-readable signal media or computer-readable storage media.Computer-readable storage media can be, for example, systems, devices or components including but not limited to electricity, magnetism, light, electromagnetic, infrared, or semiconductors, or any combination thereof.More specific examples (non-exhaustive list) of computer-readable storage media include: electrical connections with one or more wires, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination thereof.In this document, computer-readable storage media can be any tangible medium containing or storing a program that can be used by an instruction execution system, device or device or used in combination with it.

[0092] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0093] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0094] The computer program code for performing the operations of the present application can be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0095] The flow chart or block diagram in the accompanying drawings illustrate the possible architecture, functions and operations of the equipment, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code include one or more executable instructions for realizing the logical function of the specification. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented with a dedicated system for hardware that performs the function or operation of the specification, or can be implemented with a combination of dedicated hardware and computer instructions.

[0096] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0097] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of modules and units is only a logical function division. There may be other division methods in actual implementation. For example, with units as an example, for example, multiple units or page components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0098] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0099] In addition, the functional modules and units in the various embodiments of the present application may be integrated into a single processing module or unit, or each module or unit may exist physically separately, or two or more units may be integrated into a single module or unit. The aforementioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional modules or units.

[0100] The above-mentioned integrated modules and units implemented in the form of software functional modules and units can be stored in a computer-readable storage medium. The above-mentioned software functional modules and units are stored in a storage medium, including a number of instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute some steps of the methods of each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., various media that can store program code.

[0101] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

[0102] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in a device claim may also be implemented by a single unit or device through software or hardware. Terms such as "first" and "second" are used to indicate names and do not imply any particular order.

Claims

1. A method for searching an image, characterized in that: include: determining a target video based on a first content of a sample image for searching for semantically similar images, wherein a semantic similarity between a second content of the target video and the first content is greater than or equal to a first similarity threshold; In response to being able to determine a target image from the video frame of the target video, the target image is used as a semantically similar image search result for the sample image, wherein the semantic similarity between the third content of the sample image and the first content is greater than or equal to a second similarity threshold, and the value of the second similarity threshold is greater than or equal to the first similarity threshold.

2. The method according to claim 1, characterized in that The second content is generated by a content generation model by processing the title information and content introduction information of the target video.

3. The method according to claim 1, characterized in that The second content is generated by processing the image combination content of the target video by a content generation model, and the image combination content is obtained by combining the fourth content of the video frames of the target video.

4. The method according to claim 1, wherein The second content is generated based on the following method: Splitting the target video into a set of unimodal data, wherein each unimodal data in the set of unimodal data corresponds to only one data modality; generating target modality data of each of the single modality data in the target modality; Combining the target modal data semantics of each of the target modal data to obtain a target modal data semantic combination result; The content generation model is used to process the semantic combination result of the target modal data to generate the second content.

5. The method according to any one of claims 2 to 4, characterized in that The content generation model includes a large language model.

6. The method according to claim 1, characterized in that In response to a search result of a semantically similar image search target for the sample image being a different image type from the sample image, the method further includes: determining a group of candidate videos based on the type label of the sample image, wherein the type label associated with each candidate video in the group of candidate videos is different from the type label associated with the sample image; and The determining of the target video based on the first content of the sample image includes: Based on the first content of the sample image, a target video is determined from the set of candidate videos.

7. The method according to claim 1, characterized in that Also includes: In response to receiving a video providing request for the target video sent by the target device, the target video is provided to the target device based on a communication path with the target device.

8. A device for searching images, characterized in that: include: a target video determination module configured to determine a target video based on a first content of a sample image used to search for semantically similar images, wherein a semantic similarity between a second content of the target video and the first content is greater than or equal to a first similarity threshold; A search result determination module is configured to, in response to being able to determine a target image from the video frame of the target video, use the target image as a semantically similar image search result for the sample image, wherein the semantic similarity between the third content of the sample image and the first content is greater than or equal to a second similarity threshold, and the value of the second similarity threshold is greater than or equal to the first similarity threshold.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.

10. A computer-readable medium, characterized in that Computer program instructions are stored thereon, and the computer program instructions can be executed by a processor to implement the method according to any one of claims 1 to 7.

11. A computer program product, characterized in that The invention comprises a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 7.