Visual media searching method and device and storage medium

By analyzing the semantic proportion related to visual content in search statements and using different recall methods, the problem of poor processing of complex search statements in the existing technology is solved, and the completeness and accuracy of search results are improved.

CN120067369APending Publication Date: 2025-05-30HONOR DEVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311564011.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-21
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, it is difficult to effectively process complex search statements when conducting visual media searches, resulting in incomplete or inaccurate search results.

Method used

By analyzing the semantic proportion related to visual content in search statements, different recall methods are used to improve the integrity and accuracy of search results. The specific method includes recalling based on the matching degree of the visual content of the visual media to the search statement, and recalling based on the matching degree of the attributes of the visual media to the search statement.

Benefits of technology

Improves search accuracy and user search experience, ensuring the completeness and accuracy of search results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067369A_ABST
    Figure CN120067369A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a visual media searching method and device and a storage medium. The method is suitable for the electronic equipment and comprises the following steps: if a search statement belongs to a high-vision semantic search statement, determining a search result of the search statement according to a plurality of first vision media files; the plurality of first visual media files are determined from the plurality of visual media according to the matching condition of the search statement and the visual contents of the plurality of visual media; if the search statement belongs to the low visual semantic search statement, determining a search result of the search statement according to a plurality of second visual media; the plurality of second visual media files are determined from the plurality of visual media according to the matching condition of the semantic main body irrelevant to the visual content in the search statement and the attributes of the plurality of visual media. According to the scheme, for high and low visual semantic search statements, the candidate sets are recalled by adopting the adaptive recall modes respectively, so that the integrity and accuracy of the search result are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of terminals, and in particular, to a visual media search method, device, and storage medium. Background Art

[0002] With the popularization of intelligent terminals, more and more users use intelligent terminals, such as mobile phones, to take pictures and videos, and store the taken pictures and videos in the picture gallery of the electronic device, so as to record every bit of life. In addition, users can also download pictures, take screenshots of the mobile phone interface, and store the downloaded pictures and screenshots in the picture gallery of the electronic device.

[0003] In order to facilitate users to manage and view pictures in the terminal, a picture management function and a picture search function are configured in the picture gallery application or other similar applications of the terminal device. For example: the picture gallery application in the terminal can classify the pictures in the terminal according to information such as the shooting time and location of the pictures to generate corresponding albums, and users can view relevant pictures by searching for information such as time and location. Summary of the Invention

[0004] Multiple aspects of this application provide a visual media search method, device, and storage medium, which can improve the search accuracy and the user search experience.

[0005] In a first aspect, a visual media search method applicable to an electronic device is provided, including:

[0006] Display a first interface; the first interface includes a search box;

[0007] Receive a search operation on a search statement input into the search box;

[0008] If the semantic proportion related to visual content in the search statement is greater than or equal to a preset proportion threshold, determine the search result of the search statement according to multiple first visual media files; the multiple first visual media files are determined from the multiple visual media according to the matching situation between the search statement and the visual content of the multiple visual media; the visual content is data that needs to be obtained through a natural picture understanding model;

[0009] If the semantic proportion related to visual content in the search statement is less than the preset proportion threshold, determine the search result of the search statement according to multiple second visual media; the multiple second visual media files are determined from the multiple visual media according to the matching situation between the semantic subject unrelated to visual content in the search statement and the attributes of the multiple visual media;

[0010] Display the search result.

[0011] Specifically, the matching situation between the search statement and the visual content of multiple visual media may include: the matching situation (e.g., matching degree) between the first sentence semantic vector of the search statement and the visual semantic vectors of multiple visual media, or the matching situation (e.g., matching degree) between the subject semantic vector of the semantic subject related to the visual content in the search statement and the visual semantic vectors of multiple visual media.

[0012] In this solution, the semantic proportion related to the visual content in the search statement is greater than or equal to a preset proportion threshold, which can be understood as a high-visual-semantic search statement; the semantic proportion related to the visual content in the search statement is less than the preset proportion threshold, which can be understood as a low-visual-semantic search statement. For high-visual-semantic search statements and low-visual-semantic search statements, their respective adapted recall methods are used to recall the candidate set, thereby improving the integrity and accuracy of the search results and avoiding the incompleteness or inaccuracy of the search results caused by using the same recall method.

[0013] For example: when the search statement belongs to a low-visual-semantic search statement (e.g., photos taken in Beijing this year), if recalled according to the matching situation between the search statement and the visual content of multiple visual media, then many photos taken by users in Beijing this year will be filtered out, resulting in the incompleteness of the search results. When the search statement belongs to a high-visual-semantic search statement (e.g., the sky taken in Beijing this year), if recalled according to the matching situation between the time or location in the search statement and the attributes of multiple visual media, then many photos unrelated to "sky" will be recalled, resulting in the inaccuracy of the search results.

[0014] In a possible implementation manner, the above method further includes:

[0015] Remove the semantic subjects in the search statement that are irrelevant to the visual content to obtain a rewritten search statement;

[0016] Determine multiple reference visual media from the multiple visual media according to the matching situation between the rewritten search statement and the visual content of the multiple visual media;

[0017] Determine the first representative visual semantic vector corresponding to the multiple first visual media files and the second representative visual semantic vector corresponding to the multiple reference visual media;

[0018] Determine the semantic proportion related to the visual content in the search statement according to the vector similarity between the first difference vector and the second difference vector;

[0019] The first difference vector represents the difference between the first sentence semantic vector and the first representative visual semantic vector; the second difference vector represents the difference between the second sentence semantic vector and the second representative visual semantic vector.

[0020] The greater the vector similarity between the first difference vector and the second difference vector, the greater the proportion of semantics related to visual content in the search statement; the smaller the vector similarity between the first difference vector and the second difference vector, the smaller the proportion of semantics related to visual content in the search statement. That is: the proportion of semantics related to visual content in the search statement is positively correlated with the vector similarity between the first difference vector and the second difference vector.

[0021] In practical applications, the composition of a sentence is relatively complex. Therefore, it is difficult to simply calculate the proportion of semantics related to visual content in this sentence based on the number of semantic entities related to visual content and the number of semantic entities unrelated to visual content it contains. The above method utilizes the word analogy property of the distributed representation vector and can accurately determine the proportion of semantics related to visual content in the search statement.

[0022] In a possible implementation, removing the semantic entities unrelated to visual content in the search statement to obtain the rewritten search statement includes:

[0023] Removing the semantic entities related to time and / or the semantic entities related to location in the search statement to obtain the rewritten search statement.

[0024] In a possible implementation, the method further includes:

[0025] Performing semantic understanding on the search statement to obtain the first sentence semantic vector;

[0026] Obtaining the visual semantic vectors of the multiple visual media files; the visual semantic vector of each visual media file is obtained by performing semantic understanding on the image or image frame of the visual media file using a natural picture understanding model;

[0027] Determining M candidate visual media files whose visual semantic vectors match the first sentence semantic vector from the multiple visual media files; where M is an integer greater than 1;

[0028] Determining the multiple first visual media files according to the M candidate visual media files.

[0029] That is to say, the above-mentioned multiple first visual media files are recalled by matching the visual semantic vectors of the visual media files with the sentence semantic vectors of the search statements (hereinafter referred to as the first branch recall). In this way, it can be ensured that the visual content of the visual media in the finally obtained search results is semantically matched with the search statements.

[0030] In a possible implementation manner, determining the multiple first visual media files according to the M candidate visual media files includes:

[0031] Determining N to-be-matched dimensions corresponding to the search statement according to the semantic entities related to visual content in the search statement; N is greater than 1 and is an integer;

[0032] Obtaining the weights of the N to-be-matched dimensions; wherein, the weight of the jth to-be-matched dimension is positively correlated with the degree of variation of the matching degree of the visual content of the M candidate visual media files on the jth to-be-matched dimension; j is an integer, and the value of j ranges from 1 to N in sequence;

[0033] Performing weighted summation on the matching degrees of the visual content of the ith candidate visual media file on each of the N to-be-matched dimensions according to the weights of the N to-be-matched dimensions to obtain the comprehensive matching degree of the ith candidate visual media file; i is an integer, and the value of i ranges from 1 to M in sequence;

[0034] Determining the multiple first visual media files whose comprehensive matching degrees meet the preset requirements from the M candidate visual media files.

[0035] The above-mentioned M candidate visual media files are obtained by matching the visual semantic vectors of the visual media files with the sentence semantic vectors of the search statements. That is to say, when recalling the above-mentioned M candidate visual media files, the matching degree between the visual semantics of the visual media files and the entire search statement of the user is considered, rather than the matching degree between the visual semantics of the visual media files and the semantic entities that the user focuses on visually in the search statement, resulting in the recall of some inaccurate photos and videos. Therefore, in this solution, the result of the above first branch recall is fine-tuned by means of the matching degree between the visual semantics of the visual media and the "semantic entities related to visual content" to filter out inaccurate visual media.

[0036] Specifically, each semantic entity related to visual content in the search statement and / or the search statement itself can be used as different dimensions to be matched. Different dimensions to be matched play different roles in fine-tuning, that is, the weights of different dimensions to be matched are different. The weight of each dimension to be matched is determined according to the degree of variation in the matching degree of the visual content of M candidate visual media files in this dimension to be matched. The greater the degree of variation, the greater the role played by the dimension to be matched in fine-tuning. Therefore, its weight is greater. In this way, some inaccurate visual media such as photos and videos can be effectively excluded.

[0037] In a possible implementation manner, the method further includes:

[0038] Match the search terms in the search statement with a plurality of preset tags to determine whether the search terms belong to semantic entities related to the tags; the tags are used to describe visual content;

[0039] Determine the semantic entities related to the tags in the search statement as the semantic entities related to visual content in the search statement.

[0040] That is, the search terms that match a certain tag among the plurality of preset tags belong to the semantic entities related to the tag. Since the tag is related to visual content, the semantic entities related to the tag also belong to the semantic entities related to visual content. In this solution, by presetting a plurality of tags, the semantic entities that the user is more visually concerned about can be relatively simply determined from the search statement.

[0041] In a possible implementation manner, the method further includes:

[0042] Obtain the attributes of the plurality of visual media files;

[0043] Determine the plurality of second visual media among the plurality of visual media files whose attributes match the semantic entities unrelated to visual content in the search statement.

[0044] That is, recall is performed through the attributes of the visual media files (for example: time, location).

[0045] In a possible implementation manner, determining the search result of the search statement according to a plurality of first visual media files includes:

[0046] If the search statement includes a semantic entity related to time, filter the plurality of first visual media files according to the semantic entity related to time and the acquisition time attribute of the plurality of first visual media files;

[0047] If the search statement includes a semantic entity related to a location, filter the multiple first visual media files according to the semantic entity related to the location and the collection location attribute of the multiple first visual media files;

[0048] If the search statement includes a semantic entity related to a person relationship, filter the multiple first visual media files according to the semantic entity related to the person relationship and the person relationship attribute of the multiple first visual media files; and / or

[0049] If the search statement includes a semantic entity related to a person name, filter the multiple first visual media files according to the semantic entity related to the person name and the person name attribute of the multiple first visual media files.

[0050] In this solution, time filtering, location filtering, person relationship filtering, and person name filtering are performed on the recalled visual media.

[0051] In a second aspect, the present application provides an electronic device, including: a memory, a processor, and a display, where

[0052] The memory is used to store a program;

[0053] The processor is coupled to the memory and the display, and is configured to execute the program stored in the memory to implement any one of the methods.

[0054] In a third aspect, the present application provides a computer-readable storage medium storing a computer program, where the computer program, when executed by a computer, is capable of implementing any one of the methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation of the present application. In the drawings:

[0056] Figure 1A A set of interface diagrams of the search interface for a mobile phone to enter the gallery application provided by an embodiment of the present application;

[0057] Figure 1B A schematic diagram of the search interface after clearing the search history provided by an embodiment of the present application;

[0058] Figure 1C A set of interface diagrams related to the search in the gallery application provided by an embodiment of the present application;

[0059] Figure 1D A schematic diagram of the search failure interface provided by an embodiment of the present application;

[0060] Figure 2A Schematic diagram of the structure of the electronic device provided in another embodiment of the present application;

[0061] Figure 2B Software structure block diagram of the electronic device provided in another embodiment of the present application;

[0062] Figure 3A Another set of interface diagrams involved in searching in the gallery application provided in an embodiment of the present application;

[0063] Figure 3B A set of interface diagrams involved in searching on the negative first screen provided in an embodiment of the present application;

[0064] Figure 3C Schematic diagram one of the search result interface provided in an embodiment of the present application;

[0065] Figure 3D Schematic diagram two of the search result interface provided in an embodiment of the present application;

[0066] Figure 3E Schematic diagram three of the search result interface provided in an embodiment of the present application;

[0067] Figure 3F Schematic diagram of the search result interface provided in an embodiment of the present application Figure Four ;

[0068] Figure 4 Interaction diagram of the visual media search method provided in an embodiment of the present application;

[0069] Figure 5 Flow schematic diagram of the visual media search method provided in an embodiment of the present application;

[0070] Figure 6 Schematic diagram of the search result interface provided in an embodiment of the present application Figure Five . Detailed implementation manners

[0071] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application. Among them, in the description of the embodiments of the present application, unless otherwise specified, " / " means "or". For example, A / B may represent A or B; herein, "and / or" is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of the present application, "a plurality of" means two or more than two.

[0072] Hereinafter, the terms "first", "second", and "third" are only used for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first", "second", and "third" may explicitly or implicitly include one or more of such features.

[0073] First, the vocabulary involved in the embodiments of this application will be described. It can be understood that this description is for a clearer understanding of the embodiments of this application and does not necessarily constitute a limitation on the embodiments of this application.

[0074] Visual media: refers to pictures or videos.

[0075] Semantic entity: The named entity recognition technology can identify text and recognize entities with specific meanings in the text, such as person names (PER), place names (LOC), etc. In this solution, the entities with specific meanings identified are called semantic entities.

[0076] Visual content related and visual content unrelated: Visual content refers to the objects presented by visual media and their interrelationships, etc. Computer vision can endow a computer with capabilities similar to human vision, including perceiving, understanding, analyzing, and interpreting visual content. Currently, the Generative Pre-trained Transformer 4 (GPT-4) can support inputting an image into the model and then outputting human natural language describing the important information in the picture.

[0077] In the context of image search in this solution, the data that a visual media file needs to go through the natural picture understanding of the model to obtain is called "visual content related". That is to say, visual content is the data that needs to be obtained through the natural picture understanding model. This solution refers to the data that is related to the visual media file and can be obtained without going through the picture understanding ability of the model as "visual content unrelated", such as the shooting location, shooting time, name, file attributes, etc. that can be obtained and saved when the terminal device collects the visual media file.

[0078] For example, in the sentence "a photo taken in Beijing this year", "this year" (shooting time), "Beijing" (shooting location), and "photo" (file attribute) are all data that can be obtained and saved when the terminal device collects the visual media file. Therefore, "this year", "Beijing", and "photo" are unrelated to visual content; in the sentence "the sky taken in Beijing this year", "the sky" needs to be obtained through the picture understanding ability of the model to understand the image or image frame of the visual media file. Therefore, "the sky" is visual content related.

[0079] Text semantic vector: It can be obtained by feeding text into a text encoder, which is a vector capable of representing the semantic features of an entire sentence. The text encoder can adopt models such as Transformer commonly used in Natural Language Processing (NLP), and this solution places no restrictions here. In this solution, the text semantic vector obtained for a sentence is called a sentence semantic vector, the text semantic vector obtained for the semantic subject in a sentence is called a subject semantic vector, and the text semantic vector obtained for a label is called a label semantic vector.

[0080] Visual semantic vector: It can be obtained by feeding an image or image frame of a visual media file into an image encoder. Commonly used CNN (Convolutional Neural Network) models or VIT (Vision Transformer) models can be adopted, and this solution places no restrictions here.

[0081] Density-based clustering algorithm: It is based on a set of neighborhoods to describe the tightness of a sample set, and (the first parameter ∈, the second parameter MinPts) is used to describe the tightness of the sample distribution in the neighborhood. The first parameter ∈ is used to describe the neighborhood radius of a data point; the second parameter MinPts is used to describe the minimum number of data points in the neighborhood of a data point. Its representative algorithms include: DBSCAN (Density-Based Spatial Clustering of Application with Noise); the DBSCAN algorithm is a relatively representative density-based clustering algorithm that can divide regions with sufficient high density into clusters and can discover clusters of arbitrary shapes in a spatial database with noise;

[0082] Vector similarity: It is used to describe the similarity degree between two vectors (for example: between a sentence semantic vector and a visual semantic vector). In the embodiments of this application, the visual media that matches the search statement can be determined by comparing the similarity between the sentence semantic vector and the visual semantic vector. Generally, the vector similarity can be calculated through the cosine similarity calculation formula. Of course, it can also be calculated through other methods.

[0083] In the prior art, a mobile phone manages visual media files such as pictures and videos of users through a gallery application (hereinafter referred to as: visual media). Taking the example of a mobile phone taking a photo, after the mobile phone takes a photo, the gallery application can obtain and save attributes unrelated to the visual content such as the shooting location, shooting time, and photo name of the photo.

[0084] In practical applications, when the mobile phone is charging and the screen is off, photos can also be input into the natural image understanding model to generate and save tags for the photos by the natural image understanding model. These tags can be regarded as attributes related to the visual content of the photos, and the tags can be: "sky", "cat", "dog", and so on. The gallery application can establish an index for the photos based on their attributes. After the index is established, the gallery application can provide corresponding search services to users. Specifically, users can search for pictures or videos in the gallery application by entering keywords. Exemplarily, users can enter keywords such as "Beijing", "sky", "National Day" in the search box provided by the gallery application. The gallery application matches the keywords entered by the user with the indexes of visual media such as pictures and videos in the gallery application to obtain search results.

[0085] The following describes the interfaces involved in the search process of the gallery application in the prior art with reference to the accompanying drawings:

[0086] As Figure 1A shown in (a) of [Figure], the mobile phone can display the main interface 101, which can also be called the desktop. The main interface 101 can include the icon 102 of the gallery application. The mobile phone can receive the operation of the user clicking the icon 102. In response to this operation, the mobile phone can start the gallery application and display the interface 103 as shown in Figure 1A shown in (b) of [Figure], where the interface 103 can be an album interface. It should be noted that in response to the operation of the user clicking the icon 102, the mobile phone can start the gallery application and display the photo interface of the gallery. The photo interface includes thumbnails of the photos (i.e., pictures) in the gallery or a large picture of a certain photo. In the photo interface, in response to the operation of the user on the "Album" control, the above-mentioned album interface 103 is displayed.

[0087] As Figure 1A shown in (b) of [Figure], the interface 103 includes multiple albums. Among them, the "All Photos" album includes 2,023 photos, the "Camera" album includes 1,502 photos and videos, the "Screenshots & Screen Recordings" album includes 102 photos and videos, the "My Favorites" album includes 48 photos and videos, the "One-Tap Multi-Gains" album has 34 photos and videos, the "Video Editing" album has 65 videos, the "Self-created Album" has 57 photos and videos, and the "Shared Album" has 100 photos and videos.

[0088] As Figure 1A shown in (b) of [Figure], the interface 103 can include a search box 104. The mobile phone can receive the operation of the user clicking the search box 104. In response to this operation, the mobile phone can display as shown in Figure 1AThe interface 105 shown in (c) therein can be called a search interface. Among them, the interface 105 can display the classification information of photos to the user. For example, in the interface 105, the mobile phone classifies the photos of the local machine according to time, portraits, and things, etc. For example, in the dimension of time, the mobile phone classifies the photos of the local machine according to the three time periods of "this month", "last month", and "this year". Among them, the "this month" album includes the photos or videos taken by the mobile phone this month, the "last month" album includes the photos or videos taken by the mobile phone last month, and the "this year" album includes the photos or videos taken by the mobile phone this year. In the dimension of portraits, the mobile phone classifies the photos of the local machine according to different people, such as the four different people in the interface 105. In the dimension of things, the mobile phone classifies and displays the photos of the local machine according to "scenery", "animals", "documents", and "buildings". It should be noted that the above classification dimensions can also be others, and no specific restrictions are made here. In the interface 105, the user can see this classification information without entering keywords.

[0089] Optionally, the interface 105 may further include a search history 107 and an option of "clear" 108. The search history includes the keywords that the user has entered, such as "flowers", "coffee", "cat", etc. The mobile phone can receive the operation of the user clicking "clear" 108, and in response to this operation, the mobile phone can clear the search history. After the mobile phone clears the search history, the keywords that the user has entered are no longer displayed on the search interface 105. For example, in response to the operation of the user clicking "clear" 108, as Figure 1B shown, the search history 107 and the option of "clear" 108 are no longer displayed on the search interface 105, and the content displayed below moves up.

[0090] In response to the operation of the user entering the keyword "sky" in the interface 105, the mobile phone displays the interface 109 shown in (a) of Figure 1C . As shown in (a) of Figure 1C , 100 photos related to "sky" and 32 photos related to the photos containing the word "sky". Among them, the 100 photos related to "sky" can be recalled because the tags of these 100 photos match "sky"; the 32 photos related to the photos containing the word "sky" can be recalled because through the OCR (Optical Character Recognition) technology, it is recognized that these 32 photos contain the word "sky". In practical applications, the mobile phone can also associate the keywords entered by the user to obtain associated words and perform searches based on the associated words.

[0091] The interface 109 also displays some search results related to the keyword "sky" and a "More" option 110 corresponding to the search results of the keyword "sky". The mobile phone receives the user's click operation on the "More" option 110 and displays the interface 111 as shown in Figure 1C (b) in the figure. Among them, the interface 111 is used to display photos and videos in the search results of the keyword "sky". Optionally, the photos and videos can be classified and displayed according to time. In addition, the interface 111 also includes a return key 112 and a title 113. In response to the user's operation on the return key 112, the mobile phone can redisplay the interface 109. The title 113 may include the keyword "sky".

[0092] That is to say, in the existing gallery applications, when users enter simple keywords in the search box, such as: sky, Beijing, National Day, etc., corresponding search results can be obtained. However, because the number of photo attributes is simple and limited, and the mobile phone's ability to understand and associate search statements is also limited. When users enter a more complex search statement in the search box, if the keywords in the search statement cannot match the attributes of the pictures or the text in the pictures, no photos can be found. That is to say, the existing gallery applications do not support the search function based on complex search statements. As Figure 1D shown, when the user enters a more complex search statement "warming oneself by the fire and brewing tea" in the search box of the interface 114, the mobile phone cannot understand the associated words of "warming oneself by the fire and brewing tea". Since the photos do not have attributes that can match "warming oneself by the fire and brewing tea" or its associated words, the search result shows "no pictures".

[0093] However, in practical applications, users have a strong demand for the function of searching for pictures based on complex search statements. This is because users can describe the pictures or videos they want more comprehensively through complex search statements, thereby achieving accurate search. To meet this demand of users, an embodiment of the present application provides a visual media search method. This method can be applied to electronic devices, and the electronic devices can be terminal devices such as mobile phones, tablet computers, wearable devices, in-vehicle devices, augmented reality (AR) / virtual reality (VR) devices, laptop computers, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), etc.

[0094] Exemplarily, Figure 2AThe structural schematic diagram of the electronic device 200 is shown. The electronic device 100 may include a processor 210, an external memory interface 220, an internal memory 221, a universal serial bus (USB) interface 230, a charging management module 240, a power management module 241, a battery 242, an antenna 1, an antenna 2, a mobile communication module 250, a wireless communication module 260, an audio module 270, a speaker 270A, a receiver 270B, a microphone 270C, a headphone interface 270D, a sensor module 280, a key 290, a motor 291, an indicator 292, a camera 293, a display screen 294, and a subscriber identification module (SIM) card interface 295, etc.

[0095] Among them, the sensor module 280 may include a pressure sensor 280A, a gyroscope sensor 280B, a barometric pressure sensor 280C, a magnetic sensor 280D, an acceleration sensor 280E, a distance sensor 280F, a proximity light sensor 280G, a fingerprint sensor 280H, a temperature sensor 280J, a touch sensor 280K, an ambient light sensor 280L, a bone conduction sensor 280M, etc.

[0096] It can be understood that the structure schematically shown in the embodiments of the present application does not constitute a specific limitation on the electronic device 200. In other embodiments of the present application, the electronic device 200 may include more or fewer components than those shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0097] The processor 210 may include one or more processing units. For example, the processor 210 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.

[0098] The controller may generate operation control signals according to the instruction operation code and the timing signal to complete the control of fetching instructions and executing instructions.

[0099] A memory can also be provided in the processor 210 for storing instructions and data. In some embodiments, the memory in the processor 210 is a cache memory. This memory can store the instructions or data that the processor 210 has just used or recycled. If the processor 210 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 210, and thus improves the efficiency of the system.

[0100] In some embodiments, the processor 210 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0101] It can be understood that the interface connection relationships between the modules illustrated in the embodiments of the present application are only illustrative descriptions and do not constitute a structural limitation on the electronic device 200. In other embodiments of the present application, the electronic device 200 may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.

[0102] The electronic device 200 implements the display function through a GPU (Graphics Processing Unit), a display screen 294, and an application processor, etc. The GPU is connected to the display screen 294 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 210 may include one or more GPUs, which execute program instructions to generate or change display information.

[0103] The display screen 294 is used to display images, videos, etc. The display screen 294 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include one or N display screens 294, where N is a positive integer greater than 1.

[0104] The electronic device 200 can implement the shooting function through an ISP, a camera 293, a video codec, a GPU, a display screen 294, an application processor, etc.

[0105] The camera 293 is used to capture still images or videos. An object generates an optical image through a lens and projects it onto a photosensitive element. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the ISP to convert it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in standard formats such as RGB and YUV. In some embodiments, the electronic device 100 may include one or N cameras 293, where N is a positive integer greater than 1.

[0106] The video codec is used to compress or decompress digital videos. The electronic device 200 can support one or more video codecs. In this way, the electronic device 200 can play or record videos in multiple coding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0107] The NPU is a neural-network (NN) computing processor. By drawing on the structure of biological neural networks, such as the transmission pattern between human brain neurons, it can quickly process input information and can also continuously learn on its own. Through the NPU, applications such as intelligent cognition of the electronic device 200 can be realized, such as: image recognition, face recognition, voice recognition, text understanding, etc.

[0108] The external memory interface 220 can be used to connect to an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 200. The external memory card communicates with the processor 210 through the external memory interface 220 to implement the data storage function. For example, files such as music and videos are saved in the external memory card.

[0109] The internal memory 221 can be used to store computer-executable program code, and the executable program code includes instructions. The internal memory 221 can include a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, image playback function, etc.). The data storage area can store data created during the use of the electronic device 200 (such as audio data, phone book, etc.). In addition, the internal memory 221 can include high-speed random access memory and can also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. The processor 210 executes various functional applications and data processing of the electronic device 200 by running the instructions stored in the internal memory 221 and / or the instructions stored in the memory provided in the processor.

[0110] Figure 2B It is the software structure block diagram of the electronic device 200 in the embodiment of the present application. The software system of the electronic device 200 can adopt a layered architecture. The layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. Taking the Android system as an example, in some embodiments, the Android system can be divided into four layers, from top to bottom are the application layer, application framework layer, Android runtime, and system library, and the kernel layer.

[0111] As Figure 2B shown, the application layer can include application programs such as a gallery service module, a search module, a multimodal understanding module, a natural language understanding module, and a camera application.

[0112] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions.

[0113] The system libraries can include multiple functional modules. For example: surface manager, Media Libraries, 3D graphics processing libraries (such as OpenGL ES), 2D graphics engines (such as SGL), etc. Among them, the Media Libraries support the playback and recording of multiple common audio and video formats, as well as static image files, etc. The Media Libraries can support multiple audio and video coding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.

[0114] The kernel layer is the layer between the hardware and the software.

[0115] The following takes the scenario of capturing and taking pictures as an example to exemplarily illustrate the working processes of the software and hardware of the electronic device 200.

[0116] When the touch sensor 280K receives a touch operation, the corresponding hardware interrupt is sent to the kernel layer. The kernel layer processes the touch operation into a raw input event (including information such as touch coordinates and the timestamp of the touch operation). The raw input event is stored in the kernel layer. The application framework layer obtains the raw input event from the kernel layer and identifies the control corresponding to the input event. Taking the touch operation as a touch click operation and the control corresponding to the click operation being the control of the camera application icon as an example, the camera application calls the interface of the application framework layer to start the camera application, and then starts the camera driver by calling the kernel layer to capture a static image or video through the camera 293.

[0117] Taking the electronic device as a mobile phone as an example, the visual media search method provided by the embodiments of the present application is introduced. The visual media search method provided by the embodiments of the present application can be applied to application programs such as the gallery application and the file management application.

[0118] The following will describe the interfaces and search logics involved in the visual media search method provided by the embodiments of the present application with reference to the accompanying drawings.

[0119] As Figure 3A shown in (a) of [reference], the interface 301 displays the search history 303 and the "Clear" option 304. The search history includes the search statements that the user has entered, such as: "Watching the sunrise on the mountain top", "The Great Wall photographed in Beijing this year". Other contents that can be displayed on the interface 301 can refer to the relevant display contents of the above interface 305 and will not be elaborated here. Among them, the interface 301 can be called the search interface. The mobile phone can respond to the user's operation onFigure 1A By clicking the search box 104 in the interface 103 shown in (b), the interface 301 is displayed.

[0120] like Figure 3A In the interface 305 (i.e., the search result interface) shown in (b) of FIG. 305, the user enters the search sentence "cooking tea around the fire" in the search box 306 of the interface 305, and the mobile phone searches for 239 pictures. The mobile phone displays part of the search results for the search sentence "cooking tea around the fire" (e.g., thumbnails of 8 pictures) and the "more" option 307 corresponding to the search results for the search sentence "cooking tea around the fire" on the interface 305. In response to the user's operation on the "more" option 307, the mobile phone displays the following Figure 3A Interface 308 shown in (c) of FIG. Interface 308 can display the images in the search results of the search sentence "cooking tea around the fire" in descending order according to the matching degree between the visual content of the images and the search sentence "cooking tea around the fire" (specifically, the similarity between the visual semantic vector of the image and the sentence semantic vector of the search sentence). The user can also perform an upward sliding operation on interface 308 to view the images that are not displayed.

[0121] At present, mobile phones are also equipped with a negative screen, a drop-down search interface, etc. It can be understood that the negative screen can be the leftmost split screen of the electronic device, which is used to provide users with functions such as search and quick services. Among them, the negative screen can also be used to display notification messages that need to be pushed to users, such as application messages subscribed by users, real-time hot search messages, segment selection, itinerary information, etc. The drop-down search interface is an interface displayed in response to the user's drop-down operation performed on the main interface. This interface is used to provide users with functions such as search and application suggestions. This interface is similar to Figure 3B The interface 315 in can be the same interface.

[0122] The following will take the negative one screen as an example for introduction. When the user needs to view the negative one screen of the mobile phone, the electronic device can display the negative one screen by sliding the screen of the mobile phone.

[0123] For example, refer to Figure 3B As shown in (a) of FIG. 1 , the mobile phone can receive a first operation performed by the user on the interface 309 (which can be called the desktop) of the mobile phone. For example, the first operation can be as follows: Figure 3B In response to the first operation, the mobile phone may display the following Figure 3B The negative first screen 310 shown in (b) of FIG. The negative first screen 310 may include: a search box 311, a quick service 312, a default card 313, a recommended card 314, etc. The quick service 312 may be a quick entry to a page or function of an application, such as: scan, payment code, ride code, etc.; the default card may be: a gallery card, a remaining battery card, etc.; the recommended card may be a recommended application card.

[0124] The mobile phone receives a click operation by the user on the search box 311 on the minus one screen 310, and displays the interface 315 as shown in Figure 3B (c). The interface 315 may include: a search box 316, application suggestions, and a search history 217. The application suggestions include: icons of each application recommended for use. The interface 215 may also include: a search history 217 and its corresponding "Clear" option 318. In response to a trigger operation by the user on the "Clear" option 318, the search history 317 and the "Clear" option 318 are no longer displayed on the interface 315. In addition, a hot news title, such as: "Tianjin Marathon", may be displayed in the search box 316.

[0125] As Figure 3B (d) shows the interface 319. A search statement "Great Wall photographed in Beijing during the National Day" input by the user is displayed in the search box of the interface 319; a preview area 322 of the search results of the picture library for the search statement "Great Wall photographed in Beijing during the National Day" and a "Search in App" option 323 corresponding to the picture library are also displayed in the interface 319. In response to a trigger operation by the user on the preview area 322, the mobile phone enters a photo details interface provided by the picture library application for the user to flip through the search results of the search statement "Great Wall photographed in Beijing during the National Day". In response to a trigger operation by the user on the "Search in App" option 323, the mobile phone displays the interface 324 provided by the picture library application as shown in Figure 3B (e). The interface 324 displays some search results of the search statement "Great Wall photographed in Beijing during the National Day" and a "More" option corresponding to the search results of the search statement "Great Wall photographed in Beijing during the National Day". In response to a trigger operation by the user on the "More" option, the mobile phone may display a search result details interface, and the search result details interface shows the pictures in the search results of the search statement "Great Wall photographed in Beijing during the National Day". An online search option 321 may also be displayed in the interface 319. In response to a trigger operation by the user on the online search option 321, the mobile phone displays a search web page and shows online search results in the search web page.

[0126] As Figure 3CThe interface 325 shown. The user enters the search statement "The Great Wall taken during last year's National Day" in the search box of interface 325, and the mobile phone searches for 419 pictures. The mobile phone displays the thumbnails of each picture or video in the search results of the search statement "The Great Wall taken during last year's National Day" on interface 325. Among them, the shooting time of the picture or video corresponding to thumbnail A is 23:22 on September 30, 2022; the shooting time of the picture or video corresponding to thumbnail B is 00:19 on October 1, 2022. Among them, the time search range corresponding to National Day is from 00:00 on October 1, 2022 (i.e., the first time point) to 00:00 on October 8, 2022 (i.e., the second time point). The shooting time of the picture or video corresponding to thumbnail A is before 00:00 on October 1, 2022. The shooting time of the picture or video corresponding to thumbnail B is 00:19 on October 1, 2022, which is between 00:00 on October 1, 2022 and 00:00 on October 8, 2022. Among them, the shooting times of the pictures or videos corresponding to thumbnail A and thumbnail B are relatively close.

[0127] As Figure 3D Shown in the interface 326, for the search statement "The sky taken during last year's National Day", the shooting time of the picture or video corresponding to the displayed thumbnail C is 22:19 on October 7, 2022; the shooting time of the picture or video corresponding to the displayed picture D is 01:24 on October 8, 2022. Among them, the time search range corresponding to National Day is from 00:00 on October 1, 2022 (i.e., the first time point) to 00:00 on October 8, 2022 (i.e., the second time point). The shooting time of the picture or video corresponding to thumbnail D is after 00:00 on October 8, 2022. The shooting time of the picture or video corresponding to thumbnail C is 00:19 on October 1, 2022, which is between 00:00 on October 1, 2022 and 00:00 on October 8, 2022. Among them, the shooting times of the pictures or videos corresponding to thumbnail C and thumbnail D are relatively close.

[0128] In practical applications, the search statement may include, in addition to the time search range (such as National Day), a first keyword (such as The Great Wall or the sky). Then, for the search statement, the visual media file corresponding to its search results matches the first keyword.

[0129] When the first keyword belongs to a semantic entity unrelated to visual content (such as: location), the first keyword can be matched with the attributes of the visual media to determine the visual media file that matches the first keyword.

[0130] When the first keyword belongs to a semantic entity related to visual content (such as the Great Wall or the sky), match the first keyword with the tags of the visual media file to determine the visual media file whose visual content matches the first keyword; or, match the text semantic vector of the first keyword with the visual semantic vector of the visual media file to determine the visual media file whose visual content matches the first keyword (including the above-mentioned first, second, and third visual media files).

[0131] The above thumbnail is a scaled thumbnail of the corresponding visual media file or a scaled thumbnail of any image frame.

[0132] As Figure 3E Shown in the interface 327, in which, for the search statement "the sky photographed during last year's National Day", the two thumbnails E and F shown are sorted and displayed in descending order according to the matching degree of the visual content of their respective visual media files with the first keyword "sky". Among them, the matching degree of the visual content of the visual media file corresponding to the thumbnail E with "sky" is 0.87; the matching degree of the visual content of the visual media file corresponding to the thumbnail F with "sky" is 0.76.

[0133] As Figure 3F Shown in the interface 328, in which, for the search statement "photos taken in Nankai District, Tianjin during last year's National Day", the two thumbnails J and H shown are sorted and displayed in descending order according to the matching degree of the location attributes of their respective visual media files with the first keyword "Nankai District, Tianjin". Among them, the matching degree of the location attribute of the visual media file corresponding to the thumbnail J with "Nankai District, Tianjin" is 0.87; the matching degree of the location attribute of the visual media file corresponding to the thumbnail H with "Nankai District, Tianjin" is 0.8.

[0134] Figure 4 This is the interaction diagram of the visual media search method provided by the embodiments of the present application. As Figure 4 Shown, on the mobile phone, there are provided: a gallery service module (i.e., a gallery application) 41, a search module 42, a multimodal understanding module 43, and a natural language understanding module 44.

[0135] As Figure 4 Shown, the visual media search method provided by the embodiments of the present application can be divided into two stages: an index construction stage and a search stage.

[0136] In the index construction stage, the following steps are included:

[0137] S401, Add and / or modify visual media and its attributes.

[0138] In the above S401, for the newly added visual media, the attributes that the gallery application can automatically generate and are irrelevant to the visual content may include but are not limited to: collection location, collection time, and visual media name. Taking a captured video or picture as an example, the collection location refers to the shooting location, and the collection time refers to the shooting time; taking a screenshot as an example, the collection location refers to the screenshot location, and the collection time refers to the screenshot time; taking a downloaded video or picture as an example, the collection location refers to the download location, and the collection time refers to the download time.

[0139] Users can add new visual media by means such as shooting, downloading, and taking screenshots. In addition, users can also modify existing visual media. Such modifications include but are not limited to: beautification, custom naming, adding watermarks, etc.

[0140] S402. Store the visual media and its attributes.

[0141] The gallery service module 41 can, in response to the above-mentioned addition or modification operations, store the visual media and its attributes locally on the mobile phone. In actual applications, with the user's authorization, the mobile phone can store the locally stored visual media and its attributes in the cloud to relieve the storage pressure on the local mobile phone.

[0142] S403. Request visual semantic understanding of the visual media.

[0143] Since visual semantic understanding requires a large amount of computing resources, in order not to affect the user's use, the above step S403 can be executed when the mobile phone is in the state of being charged and the screen is off.

[0144] The gallery service module 41 can request the multimodal understanding module 43 to perform visual semantic understanding on the newly added or modified visual media to obtain the visual semantic vector of the visual media.

[0145] Among them, the multimodal understanding module 43 can perform visual semantic understanding on the visual media based on the multimodal model to obtain the visual semantic vector of the visual media.

[0146] Among them, the multimodal model is not only used for: performing visual semantic understanding on the visual media to obtain the visual semantic vector of the visual media; but also used for: performing semantic understanding on the search statement to obtain the sentence semantic vector of the search statement; performing semantic understanding on the rewritten search statement in the following text to obtain the sentence semantic vector of the rewritten search statement; performing semantic understanding on the semantic subject in the search statement to obtain the subject semantic vector of the semantic subject; performing semantic understanding on the label of the visual media to obtain the label semantic vector of the label. The multimodal model can be trained based on training samples.

[0147] Exemplarily, the multimodal model can specifically be CLIP (Contrastive Language-Image Pre-training), a pre-training model based on contrastive text-image pairs. The CLIP model can map visual media and text (i.e., search statements, rewritten search statements, semantic entities, tags) into a unified vector space to understand the relationships between different modal resources visually and textually, and then be used for image retrieval. That is, in the embodiments of the present application, the visual media file and the text can be specifically matched through the CLIP model.

[0148] The multimodal model can map visual media and text into vectors of the same dimension. That is to say, the dimension of the visual semantic vector of the visual media is the same as that of the semantic vector of the text (for example: the sentence semantic vector of the search statement). The multimodal model includes: the above-mentioned image encoder and text encoder.

[0149] S404. Return the visual semantic vector of the visual media.

[0150] The multimodal understanding module 44 returns the visual semantic vector of the visual media to the gallery service module 41.

[0151] S405. Store the visual semantic vector of the visual media.

[0152] The gallery service module 41 can locally store the visual semantic vector of the visual media.

[0153] S406. Send the attribute information of the visual media and its visual semantic vector.

[0154] Exemplarily, the gallery service module 41 can store the visual semantic vector of the visual media returned by the multimodal understanding module 44, and then batch send the attributes of the visual media and its visual semantic vector to the search module 42 for the search module 42 to build an index of the visual media.

[0155] S407. Build an index

[0156] The index of the visual media built by the search module 42 can include: the attributes of the visual media, the visual semantic vector of the visual media.

[0157] In the search stage, the following steps are included:

[0158] S408. Input a search statement.

[0159] The user can input a search statement through the search interface provided by the gallery service module 41, for example: Figure 3A the interface 301 shown in (a) of Figure 3AFor the interface 305 shown in (b) in [reference], enter "warming the tea around the stove" in the search box 306.

[0160] S409. Send the search statement.

[0161] After the gallery service module 41 receives the search statement input by the user, it sends the search statement to the search module 42 for searching.

[0162] S410. Request semantic subject recognition for the search statement.

[0163] The search module 42 requests the natural language understanding module 44 to perform semantic subject recognition on the search statement to obtain the semantic subjects included in the search statement. Among them, the natural language understanding module 44 performs semantic subject recognition based on a natural language understanding model. Specifically, named entity recognition technology (NER) can be used to perform semantic subject recognition on the search statement to obtain the semantic subjects included in the search statement. In the embodiments of the present application, the semantic subject can also be referred to as an entity.

[0164] Using named entity recognition technology, semantic subjects related to time, location, and tags in the search statement can be recognized. Among them, semantic subjects related to time and location are semantic subjects unrelated to visual content; the tags of visual media files are data that can only be obtained through the natural picture understanding of the model. Therefore, semantic subjects related to tags are semantic subjects related to visual content. In practical applications, based on actual experience, multiple tags that users are more concerned about can be statistically obtained, such as: "sky", "cat", "dog", "birthday", "child", etc. These tags are used to describe visual content. Subsequently, named entity recognition technology can match the keywords (or search terms) in the search statement with the pre-set multiple tags to determine whether the keyword belongs to a semantic subject related to the tag.

[0165] Exemplarily, using named entity recognition technology to perform semantic subject recognition on the search statement "the sky photographed in Beijing during the National Day", it is determined that "National Day" belongs to a semantic subject related to time, "Beijing" belongs to a semantic subject related to location, and "sky" belongs to a semantic subject related to tags.

[0166] S411. Return the semantic subject.

[0167] The natural language understanding module 44 returns the recognized semantic subjects to the search module 42.

[0168] S412. Request semantic understanding of the search statement.

[0169] The search module 42 can send the search statement to the multimodal understanding module 43, and the multimodal understanding module 43 performs semantic understanding on the search statement to obtain the sentence semantic vector of the search statement (i.e., the first sentence semantic vector). The specific semantic understanding process can refer to the corresponding content in the above embodiments and will not be elaborated here.

[0170] It should be additionally supplemented that when the search statement includes a semantic entity related to visual content (i.e., a semantic entity related to a label), the search module 42 can also send the semantic entity related to visual content to the multimodal understanding module 43, so that the multimodal understanding module 43 performs semantic understanding on the semantic entity related to visual content to obtain the entity semantic vector of this semantic entity. Continuing with the above example, "sky" belongs to the semantic entity related to visual content, and the multimodal understanding module 43 can perform semantic understanding on "sky" to obtain the entity semantic vector corresponding to "sky".

[0171] S413. Return the text vector.

[0172] The multimodal understanding module 43 can return the sentence semantic vector of the search statement to the search module 42.

[0173] S414. Perform recall respectively based on the attributes of the visual media and the visual semantic vector of the visual media.

[0174] The search module 42 includes different branches of search methods:

[0175] Exemplarily, the first branch is: perform recall based on the visual semantic vector of the visual media. Specifically, obtain the visual semantic vectors of each visual media among multiple visual media stored in the mobile phone; calculate the vector similarity between the sentence semantic vector of the search statement and the visual semantic vectors of each visual media; according to the vector similarity, determine M candidate visual media (i.e., M candidate visual media files) whose visual semantic vectors match the sentence semantic vector from the multiple visual media stored in the mobile phone; where M is an integer greater than or equal to 1; these M visual media can be used as the visual media set recalled by the first branch.

[0176] Exemplarily, the visual media with a vector similarity greater than the preset similarity threshold is used as the visual media whose visual semantic vector matches the sentence semantic vector.

[0177] Exemplarily, sort the multiple visual media according to the vector similarity from high to low, and use the top F (F≥1) visual media as the visual media whose visual semantic vector matches the sentence semantic vector.

[0178] Exemplarily, Q (Q≥1) visual media with vector similarity greater than a preset similarity threshold are determined from multiple visual media stored via the mobile phone; if Q is greater than or equal to a preset quantity threshold D, these Q visual media can be sorted in descending order of vector similarity; the top D visual media in the sorting are used as the visual media whose visual semantic vectors match the semantic vector of the sentence; if Q is less than the preset quantity threshold D, these Q visual media can be directly used as the visual media whose visual semantic vectors match the semantic vector of the sentence.

[0179] It should be noted that the multiple visual media stored via the mobile phone can include: visual media stored locally by the mobile phone and / or visual media stored in the cloud by the mobile phone (for example: the cloud storage space applied for by the mobile phone). To protect user privacy, all the multiple visual media stored via the mobile phone are stored in the mobile phone.

[0180] It should be noted that during the recall process of the first branch, the attributes of the visual media are not understood, and only the overall visual semantic information of the visual media can be understood.

[0181] Exemplarily, the second branch is: recall based on the attributes of the visual media.

[0182] Specifically, obtain the attributes of each visual media among the multiple visual media stored via the mobile phone; match the semantic subject in the search statement with the attributes of each visual media to determine the visual media (i.e., the second visual media) that matches the semantic subject; use the visual media that matches the semantic subject as the set of visual media recalled by the second branch. When the number of semantic subjects in the search statement is one, the set of visual media recalled by the second branch includes the visual media that matches this one semantic subject; when the number of semantic subjects in the search statement is multiple, the set of visual media recalled by the second branch includes the visual media that each of these multiple semantic subjects matches.

[0183] Exemplarily, for a semantic subject related to time, the corresponding time range (i.e., the time search range) of the semantic subject can be determined; match the time range with the acquisition time of each visual media as an attribute to determine the visual media whose acquisition time is within this time range; use the visual media whose acquisition time is within this time range as the visual media that matches the semantic subject. For example: the semantic subject is "National Day", its corresponding time range is "October 1st to October 7th", the acquisition time of Picture 1 is "October 2nd", and the acquisition time of Picture 2 is "October 8th", then, according to the above matching method, Picture 1 matches the semantic subject "National Day", and Picture 2 does not match the semantic subject "National Day".

[0184] Exemplarily, for a semantic entity related to a location, the geographical range corresponding to the semantic entity (i.e., the geographical search range) can be determined; the geographical range is matched with the attribute of the collection location of each visual medium to determine the visual media whose collection location is within the geographical range; the visual media whose collection location is within the geographical range is used as the visual media matched by the semantic entity. For example: the semantic entity is "Beijing", its corresponding geographical range is the whole of Beijing, the collection location of Picture 3 is "Xicheng District, Beijing", and the collection location of Picture 4 is "Nankai District, Tianjin", then according to the above matching method, Picture 3 matches the semantic entity "Beijing", and Picture 4 does not match the semantic entity "Beijing".

[0185] Exemplarily, for a semantic entity related to a label, the main semantic vector of the semantic entity and the label semantic vectors of the labels of each visual medium can be obtained; the vector similarity between the main semantic vector of the semantic entity and the label semantic vectors of the labels of each visual medium is calculated; according to the vector similarity, the visual media matched by the semantic entity is determined. For example: the semantic entity is "human cub", and Picture 5 has a label of "child". Through calculation, it is found that the main semantic vector of "human cub" is similar to the label semantic vector of "child", that is, Picture 5 matches the semantic entity "human cub". In practical applications, after obtaining the set of visual media recalled by the first branch and the set of visual media recalled by the second branch, the candidate set of visual media can be determined according to the set of visual media recalled by the first branch and the set of visual media recalled by the second branch. In an optional implementation manner, the union or intersection of the set of visual media recalled by the first branch and the set of visual media recalled by the second branch can be used as the candidate set of visual media.

[0186] In practical applications, when a user searches for pictures on a mobile phone, sometimes the user focuses on the visual semantic information of the pictures, sometimes on the attribute information such as the shooting location and shooting time of the pictures, and sometimes on both. Exemplarily, when the user searches for "pictures taken today", the user focuses on the shooting time of the pictures; when the user searches for "the sky taken today", the user focuses not only on the shooting time of the pictures, but also on the visual semantics of the pictures, that is, whether the picture content is the sky; when the user searches for "pictures taken while walking in Beijing", the user focuses on the shooting location of the pictures; when the user searches for "pictures of walking taken in Beijing", the user focuses not only on the shooting location attribute of the pictures, but also on the visual semantics of the pictures, that is, whether the picture content is a walking picture.

[0187] Taking the two search statements of "photos taken in Beijing this year" and "the sky taken in Beijing this year" as examples, referring to the foregoing introduction, the semantic proportion of visual content in the search statement of "the sky taken in Beijing this year" is greater than that in the search statement of "photos taken in Beijing this year". Obviously, for the search statement of "photos taken in Beijing this year", it is more appropriate to use the visual media set recalled by the second branch as the candidate visual media set. Subsequently, the intersection of the visual media matching "this year" and the visual media matching "Beijing" in the candidate visual media set can be used as the final search result. The collection locations of the visual media in the final search result are all Beijing, and the collection time is this year, which meets the user's search requirements. If the visual media set recalled by the first branch is used as the candidate visual media set, the following situation may occur: There is a picture of "a girl holding a camera to take a photo" in the visual media set recalled by the first branch, but the shooting location of this photo is Shanghai and the shooting time is last year. Since there is the word "take" in the search statement of "photos taken in Beijing this year" and there is a "taking" action in the picture of "a girl holding a camera to take a photo", there is a certain similarity between the sentence semantic vector of the search statement of "photos taken in Beijing this year" and the visual semantic vector of the picture of "a girl holding a camera to take a photo". That is to say, the picture of "a girl holding a camera to take a photo" may be recalled. That is, when the semantic proportion of visual content is relatively low, it is not appropriate to use the visual media set recalled by the first branch as the candidate visual media set.

[0188] To solve the above problems, in an optional implementation manner, the semantic proportion related to visual content in the search statement can be determined; according to the semantic proportion related to visual content in the search statement, the first extraction ratio corresponding to the visual media set recalled by the first branch and the second extraction ratio corresponding to the visual media set recalled by the second branch are determined; when the semantic proportion is greater than or equal to the preset proportion threshold, the first extraction ratio is greater than the second extraction ratio. Exemplarily, if the semantic proportion is S, the first extraction ratio is: β*S, and the second extraction ratio is: 1-β*S, where the value of β can be set according to actual needs, and the embodiments of the present application do not make specific limitations on this.

[0189] In another alternative embodiment, the semantic proportion related to visual content in the search statement can be determined; when the semantic proportion related to visual content in the search statement is greater than or equal to a preset proportion threshold, the visual media set recalled by the first branch is used as the candidate visual media set; when the semantic proportion related to visual content in the search statement is less than the preset proportion threshold, the visual media set recalled by the second branch is used as the candidate visual media set. That the semantic proportion related to visual content in the search statement is greater than or equal to the preset proportion threshold indicates that the search statement belongs to a high-semantic search statement; that the semantic proportion related to visual content in the search statement is less than the preset proportion threshold indicates that the search statement belongs to a low-semantic search statement. The specific process will be described in detail in the embodiments below in combination with Figure 5 will be introduced in detail.

[0190] S415. Visual media filtering.

[0191] The search module 42 can perform semantic entity filtering, spatio-temporal filtering, and / or person relationship filtering on the candidate visual media set. The specific filtering method will be described in detail in the embodiments below in combination with Figure 5 will be introduced in detail.

[0192] S416. Visual media ranking.

[0193] The search module 42 ranks the candidate visual media set.

[0194] The specific ranking method will also be described in detail in the following embodiments.

[0195] S417. Return search results.

[0196] The search module 42 sends the search results obtained after ranking to the gallery service module 41.

[0197] Specifically, the search module 42 can select the top P (P≥1) visual media and their ranking information as the final search results.

[0198] S418. Display search results.

[0199] The gallery service module 41 can display the search results to the user. For example, through Figure 3A the interface 305 in Figure 3A the interface 308 in to display the search results.

[0200] Such as Figure 3A in the interface 305, the search results include 239 pictures, and the interface 305 only displays the thumbnails of 8 pictures in the search results; if the user wants to view these 239 pictures, the user can click the "More" option in the interface 305, and in response to this operation, the mobile phone displays as Figure 3AThe interface 308 therein. In one example, the display order of the pictures in the search results is related to the matching degree between the pictures and the search results. For example, the matching degree between the pictures displayed earlier and the search statement is greater than or equal to the matching degree between the pictures displayed later and the search statement. The calculation of the matching degree and the sorting method will be introduced in detail in the following embodiments.

[0201] Next, the search process executed by the search module 42 of the present application will be introduced in detail in conjunction with Figure 5 :

[0202] 501. Receive the search statement.

[0203] 502. Execute the search process corresponding to the search statement based on the visual semantic vector of the visual media.

[0204] The visual media set recalled by the first branch obtained by executing step 502 can be specifically referred to the search process corresponding to the first branch above.

[0205] Exemplarily, assume that the multimodal model can encode the visual media and the search statement into a k-dimensional vector space. The search statement is Q, and the corresponding sentence semantic vector is V Q ={α z}, z = 1, 2,..., Z, and a total of M picture sets R are recalled, and their vector representations are:

[0206]

[0207] Among them, the i-th row corresponds to the visual semantic vector of the i-th picture, and the value range of i is [1, Mm].

[0208] 503. Semantic entity recognition.

[0209] For the specific process of semantic entity recognition of the search statement, reference can be made to the corresponding content in the above embodiments, and details will not be repeated here.

[0210] 504. Perform a search based on the attributes of the visual media.

[0211] Specifically, match the semantic entities irrelevant to the time content in the search statement with the attributes of the visual media to obtain the visual media set recalled by the second branch. For details, reference can be made to the search process corresponding to the second branch above.

[0212] 505. Rewrite the search statement.

[0213] Specifically, delete the semantic entities irrelevant to the visual content in the search statement to obtain the rewritten search statement. The semantic entities irrelevant to the visual content specifically refer to: semantic entities related to time and semantic entities related to location.

[0214] Exemplary: The search statement is "the sky photographed this year", where "today" is the semantic subject related to time, and the rewritten search statement is "the photographed sky".

[0215] In practical applications, after deleting the semantic subjects irrelevant to visual content, there may be some redundant stop words. For example, for the search statement "the sky photographed in Beijing this year", where "this year" is the semantic subject related to time and "Beijing" is the semantic subject related to location, after deleting "this year" and "Beijing", the stop word "in" becomes redundant and thus also needs to be deleted. Specifically, delete the semantic subjects irrelevant to visual content in the search statement and their related stop words to obtain the rewritten search statement. Exemplarily, the rewritten search statement corresponding to the search statement "the sky photographed in Beijing this year" is "the photographed sky".

[0216] It should be noted that there is no order restriction in the execution of the above steps 502, 503, and 505. In an optional example, to improve efficiency, these three steps can be executed simultaneously.

[0217] 506. Perform the search process corresponding to the rewritten search statement based on the visual semantic vector of the visual media.

[0218] Performing step 506 obtains the set of visual media recalled by the third branch.

[0219] Specifically, obtain the visual semantic vectors of each visual media among the multiple visual media stored in the mobile phone; calculate the vector similarity between the sentence semantic vector of the rewritten search statement (i.e., the second sentence semantic vector) and the visual semantic vectors of each visual media; determine, based on the vector similarity, multiple visual media (i.e., reference visual media) whose visual semantic vectors match the sentence semantic vector from the multiple visual media stored in the mobile phone; these multiple visual media can be used as the set of visual media recalled by the third branch.

[0220] Among them, the specific implementation process of the step of "determining, based on the vector similarity, multiple visual media whose visual semantic vectors match the sentence semantic vector from the multiple visual media stored in the mobile phone" can refer to the corresponding content in the above embodiments and will not be elaborated here.

[0221] Exemplarily, the rewritten search statement is Q′, and the corresponding vector V Q′ ={α′ z}, z = 1, 2,..., Z, and a total of M′ picture sets R′ are recalled, and its vector representation is:

[0222]

[0223] Among them, the i-th row corresponds to the visual semantic vector of the i-th picture, and the value range of i is [1, M′].

[0224] 507. Calculate the semantic proportion related to visual content in the search statement.

[0225] The following will introduce a method for determining the semantic proportion related to visual content:

[0226] 5071. Determine the representative visual semantic vector V I (i.e., the first representative visual semantic vector) corresponding to the visual media set recalled by the first branch and the representative visual semantic vector V I′ (i.e., the second representative visual semantic vector) corresponding to the visual media set recalled by the third branch.

[0227] 5072. Determine the sentence semantic vector V Q and the difference from the representative visual semantic vector V I to obtain a difference vector (V Q - V I ) (i.e., the first difference vector).

[0228] 5073. Determine the sentence semantic vector V Q′ and the difference from the representative visual semantic vector V I′ to obtain a difference vector (V Q′ - V I′ ) (i.e., the second difference vector).

[0229] 5074. Determine the semantic proportion related to visual content in the search statement according to the vector similarity between the difference vector (V Q - V I ) and the difference vector (V Q′ - V I′ ).

[0230] Among them, the semantic proportion related to visual content is positively correlated with this vector similarity.

[0231] In the above 5071, in one example, the vector similarity between the visual semantic vector of each visual media in the visual media set recalled by the first branch and the sentence semantic vector of the search statement can be obtained; the visual media in the visual media set recalled by the first branch are sorted according to the vector similarity from high to low; the average vector of the visual semantic vectors of the top T (T≥1) visual media in the sorting is used as the representative visual semantic vector V I .

[0232] Continuing with the above example, the representative visual semantic vector V I is:

[0233]

[0234] The vector similarity between each visual semantic vector of the visual media recalled by the third branch and the sentence semantic vector of the rewritten search statement can be obtained; the visual media in the visual media set recalled by the third branch are sorted according to the vector similarity from high to low; the average vector of the visual semantic vectors of the top H (H≥1) visual media in the sorting is used as the representative visual semantic vector V I′ 。

[0235] Continuing with the above example: The representative visual semantic vector V I′ is:

[0236]

[0237] The values of H and T above can be the same or different, and the embodiments of the present application do not make specific limitations on this. In another example, according to the vector similarity between the visual semantic vector of each visual media in the visual media set recalled by the first branch and the sentence semantic vector of the search statement, the visual media in the visual media set recalled by the first branch can be clustered using a clustering algorithm to obtain the clustering center point; the visual semantic vector of the clustering center point is used as the first representative visual semantic vector.

[0238] According to the vector similarity between the visual semantic vector of each visual media in the visual media set recalled by the third branch and the second sentence semantic vector of the rewritten search statement, the visual media in the visual media set recalled by the third branch can be clustered using a clustering algorithm to obtain the clustering center point; the visual semantic vector of the clustering center point is used as the second representative visual semantic vector.

[0239] In the embodiments of the present application, the representative visual semantic vector is used to represent the visual semantics of the entire set. The specific clustering algorithm can be selected according to actual needs, and the embodiments of the present application do not make any limitations on this.

[0240] In the above 5074, in an optional embodiment, the vector similarity between the difference vector (V Q -V I ) and the difference vector (V Q′ -V I′ ) can be directly used as the semantic proportion related to the visual content in the search statement.

[0241] Among them, the larger the vector similarity between the difference vector (V Q -V I ) and the difference vector (V Q′ -V I′ ), the greater the semantic proportion of the visual content in the search statement; conversely, the smaller the semantic proportion of the visual content in the search statement.

[0242] It should be noted that in the embodiments of the present application, the calculation method of the semantic proportion related to visual content in the search statement draws on the word analogy characteristics of the distribution representation vector, that is, the additivity of word meanings is directly reflected in the additivity of the distribution representation vector.

[0243] When the semantic proportion related to visual content in the search statement is less than the preset proportion threshold, step 508 is executed subsequently.

[0244] When the semantic proportion related to visual content in the search statement is greater than or equal to the preset proportion threshold, steps 509 and 510 are executed subsequently.

[0245] 508. Determine the visual media recalled by the second branch as the candidate set.

[0246] 509. Determine the visual media recalled by the first branch as the candidate set.

[0247] 510. Visual media filtering.

[0248] To improve the search accuracy, one or more of the following processes can also be performed on the visual media recalled by the first branch: semantic entity filtering, spatio-temporal filtering, person name filtering, and person relationship filtering. Among them, semantic entity filtering refers to filtering the visual media recalled by the first branch based on the matching degree between the visual media recalled by the first branch and the semantic entity related to visual content in the search statement. Spatio-temporal filtering includes: time filtering and space (i.e., location) filtering. Person name filtering refers to filtering the visual media recalled by the first branch based on the matching degree between the person name attribute of the visual media recalled by the first branch and the semantic entity related to the person name in the search statement. Person relationship filtering refers to filtering the visual media recalled by the first branch based on the matching degree between the person relationship attribute of the visual media recalled by the first branch and the semantic entity related to the person relationship in the search statement. In practical applications, in addition to the above attributes such as the collection location, collection time, and visual media name stored in the mobile phone, the visual media may also include the person name attribute and person relationship attribute manually input by the user. Therefore, in practical applications, the named entity recognition technology can also be used to match the keywords in the search statement with a variety of preset person relationships to determine whether the keyword belongs to the semantic entity related to the person relationship.

[0249] When multiple processes such as semantic entity filtering, spatio-temporal filtering, person name filtering, and person relationship filtering need to be performed on the candidate set recalled by the first branch, the execution order between these multiple processes can be set according to actual needs, and the embodiments of the present application do not make specific limitations on this.

[0250] Exemplarily, as Figure 5 shown, the visual media filtering includes:

[0251] 5101. Semantic entity filtering.

[0252] 5102. Spatiotemporal filtering.

[0253] 5103. Character relationship filtering.

[0254] 5104. Person name filtering.

[0255] In the above 5101, generally, when a user searches, they hope that the recalled images contain some visual content described in their search query, such as: sky, puppy, child, etc. That is to say, the semantic entities related to visual content in the search query represent the visual focus that the user pays attention to during the search.

[0256] However, when the first branch conducts the recall, it considers the matching degree between the visual content of visual media such as images and videos and the entire search query of the user, without fully considering the role of certain specific semantic entities (i.e., the semantic entities related to visual content) in the search query in the mobile phone gallery search scenario, resulting in the recall of some inaccurate photos.

[0257] For example: when a user searches for "photos of the Great Wall taken on National Day", the semantic entity related to visual content among them includes "the Great Wall". When the first branch performs the matching, it considers the matching degree between the entire search query and the visual content of the image, which may lead to the first branch recalling images that contain the "taking" action but do not contain the "Great Wall". This is because the image contains the "taking" action and the search query contains the word "taking", that is, there is a certain similarity between the visual semantic vector of the image and the sentence semantic vector of the search query. Then, this image has the possibility of being recalled.

[0258] Another example: when a user searches for "last year, the child had a birthday holding a cake", the semantic entities related to visual content among them include "child", "birthday", and "cake". When the first branch performs the matching, it considers the matching degree between the entire search query and the image, which may lead to the model recalling images of an adult having a birthday holding a cake. This is because the similarity between the visual semantic vector of the image of an adult having a birthday holding a cake and the sentence semantic vector of "last year, the child had a birthday holding a cake" is very high.

[0259] Therefore, in order to further improve the accuracy of searching for visual media on mobile phones, on the basis of the first branch recalling visual media, the "semantic entity" can be strengthened, that is, the results recalled by the first branch can be fine-tuned by means of the matching degree between the visual media and the "semantic entity".

[0260] Specifically, the "semantic entity filtering" in the above 5101 may include the following steps:

[0261] 5101a. Determine the dimensions to be matched according to the semantic entities related to the visual content in the search statement.

[0262] 5101b. Filter the M candidate visual media according to the matching degrees of the visual content of the M candidate visual media in the dimensions to be matched.

[0263] In the above 5101a, in one example, the semantic entity related to the visual content is used as the dimension to be matched. When there are multiple semantic entities related to the visual content in the search statement, these multiple semantic entities are respectively used as different dimensions to be matched to obtain multiple dimensions to be matched. Among them, the number of multiple dimensions to be matched is the same as the number of semantic entities related to the visual content.

[0264] Exemplarily, for the search statement "Last year, the child held a cake on his / her birthday", the semantic entities related to the visual content are "child", "birthday", and "cake". Then, the corresponding three dimensions to be matched are: "child", "birthday", and "cake".

[0265] In practical applications, in addition to using the semantic entity as the dimension to be matched, the search statement itself can also be used as the dimension to be matched. In this way, when filtering the semantic entity, the matching situation between the visual content of the candidate visual media and the entire search statement can be taken into account to improve the rationality of filtering. Specifically, the semantic entity related to the visual content and the search statement can be respectively used as different dimensions to be matched to obtain multiple dimensions to be matched. When there are multiple semantic entities related to the visual content in the search statement, these multiple semantic entities and the search statement are respectively used as different dimensions to be matched. Among them, the number of multiple dimensions to be matched is one more than the number of semantic entities related to the visual content.

[0266] Exemplarily, for the search statement "Last year, the child held a cake on his / her birthday", the semantic entities related to the visual content are "child", "birthday", and "cake". Then, the corresponding four dimensions to be matched are: "child", "birthday", "cake", and "Last year, the child held a cake on his / her birthday".

[0267] In the above 5101b, the matching degree of the visual content of the candidate visual media in the dimension to be matched refers to the matching degree between the visual content of the candidate visual media and the dimension to be matched. The vector similarity between the visual semantic vector of the candidate visual media and the semantic vector to be matched in the dimension to be matched can be determined; among them, when the dimension to be matched is the semantic subject related to the visual content, the semantic vector to be matched in the dimension to be matched is the subject semantic vector of the semantic subject; when the dimension to be matched is the search statement, the semantic vector to be matched in the dimension to be matched is the sentence semantic vector of the search statement; according to the vector similarity, the matching degree of the visual content of the candidate visual media in the dimension to be matched is determined. Among them, the matching degree is positively correlated with the vector similarity.

[0268] When the number of dimensions to be matched is one, the candidate visual media with a matching degree less than or equal to the preset matching degree threshold can be filtered out according to the matching degrees of the visual contents of multiple candidate visual media in the dimension to be matched.

[0269] For example: the search statement is "photos of the sky taken on National Day", and the only semantic subject related to the visual content is "sky"; then, the recalled pictures that contain the "taking" behavior but do not contain "sky" have a relatively low matching degree with "sky" and will be filtered out.

[0270] When the number of dimensions to be matched is multiple, for each candidate visual media, the comprehensive matching degree of the candidate visual media is determined according to the matching degrees of the visual content of the candidate visual media in multiple dimensions to be matched; multiple candidate visual media are filtered according to the comprehensive matching degrees of each candidate visual media. Specifically, from M candidate visual media, multiple first visual media with a comprehensive matching degree greater than or equal to the preset matching degree threshold (that is, meeting the preset requirements) are determined, which is equivalent to filtering out the candidate visual media with a comprehensive matching degree less than the preset matching degree threshold; or, the M candidate visual media are sorted in descending order of the comprehensive matching degree, and the top Z′ (Z′≥1) candidate visual media (that is, meeting the preset requirements) are used as multiple first visual media, which is equivalent to filtering out the (M-Z′) candidate visual media at the back.

[0271] In an optional implementation manner, any one of the following three methods can be used to determine the comprehensive matching degree of the candidate visual media:

[0272] Method 1: Sum the matching degrees of the visual content of the candidate visual media in multiple dimensions to be matched to obtain the comprehensive matching degree of the candidate visual media.

[0273] Method 2: Perform a weighted sum of the matching degrees of the visual content of the candidate visual media in multiple dimensions to be matched to obtain the comprehensive matching degree of the candidate visual media.

[0274] Among them, the weights of multiple dimensions to be matched can be configured in advance by the user.

[0275] Method 3: Use a machine learning model to determine the comprehensive matching degree of candidate visual media according to the matching degree of the visual content of the candidate visual media on multiple dimensions to be matched.

[0276] Among them, the machine learning model needs to be trained based on a data set, and the purpose of its training is essentially to learn the weights of each dimension to be matched.

[0277] In the above Method 1, the contribution degree of the matching degree on different dimensions to the comprehensive matching degree is not distinguished, which may lead to the numerical values of the comprehensive matching degrees of multiple candidate visual media calculated finally being relatively close or equal, and then it is impossible to screen multiple candidate visual media.

[0278] Exemplarily, assume that there are Picture 1, Picture 2, and Picture 3 among multiple candidate visual media, and there are Dimension A, Dimension B, and Dimension C among multiple dimensions to be matched. Calculate the matching degree of the visual content of each picture on each dimension respectively, and the results are shown in Table 1.

[0279] Table 1:

[0280] Dimension A Dimension B Dimension C Picture 1 0.30 0.35 0.55 Picture 2 0.40 0.38 0.42 Picture 3 0.38 0.37 0.45

[0281] If calculated according to Method 1, the comprehensive matching scores of Picture 1, Picture 2, and Picture 3 are all 1.2, which will lead to the inability to screen the three pictures.

[0282] In the above Method 2, when there are too many semantic entities related to visual content to be concerned about (that is, the number of preset multiple tags is too large), it is difficult to accurately configure the weights of different dimensions.

[0283] In the above Method 3, the construction of the data set is inseparable from user data. However, user data belongs to user privacy content, and users do not want their data to be reported to the cloud side.

[0284] In an optional implementation manner, to solve the above problems, the following steps can be adopted to determine the comprehensive matching degree:

[0285] S51: Determine the weights of each of the N dimensions to be matched.

[0286] Among them, the weight of the jth dimension to be matched is positively correlated with the degree of variation of the matching degree of the visual content of the M candidate visual media files on the jth dimension to be matched; j is an integer, and the value of j ranges from 1 to N in sequence;

[0287] S52. According to the weights of the N dimensions to be matched, perform a weighted sum of the degrees of match of the visual content of the i-th candidate visual media file on each dimension to be matched among the N dimensions to be matched, so as to obtain the comprehensive degree of match of the i-th candidate visual media file.

[0288] Wherein, i is an integer, and the values of i sequentially range from 1 to M.

[0289] Taking the search statement "Last year, the child held a cake on his / her birthday" as an example, the semantic entities related to the visual content include: child, birthday, and cake. If multiple recalled photos all contain a cake, then the degree of variation of the degrees of match of the visual content of the multiple recalled photos on the dimension of "cake" will be relatively small, and the weight corresponding to the dimension of "cake" will be relatively small; if some of the multiple recalled photos contain a child and some do not, then the degree of variation of the degrees of match of the visual content of the multiple recalled photos on the dimension of "child" will be relatively large, and the weight corresponding to the dimension of "child" will be relatively large.

[0290] In one example, for each dimension to be matched, determine the degree of variation of the degrees of match of the visual content of multiple candidate visual media on this dimension to be matched according to the information entropy of the degrees of match of the visual content of the multiple candidate visual media on this dimension to be matched; wherein, the degree of variation is inversely proportional to the information entropy. The calculation method of the information entropy will be introduced in detail in the following embodiments.

[0291] In this embodiment, the information entropy is used to measure the degree of variation of the degrees of match under each dimension to be matched, and based on this, the weight corresponding to this dimension to be matched is determined.

[0292] In order to ensure that the degrees of match of the visual content of the candidate visual media on different dimensions to be matched have a unified dimension, a normalization processing step can be performed. Specifically, for each candidate visual media, determine the initial degree of match of the candidate visual media on the dimension to be matched according to the vector similarity between the visual semantic vector of the candidate visual media and the semantic vector to be matched of the dimension to be matched. Exemplarily, the vector similarity between the visual semantic vector of the candidate visual media and the semantic vector to be matched of the dimension to be matched can be used as the initial degree of match of the candidate visual media on the dimension to be matched. Perform normalization processing on the initial degrees of match of the visual content of multiple candidate visual media on the dimension to be matched to obtain the degrees of match of the visual content of multiple candidate visual media on this dimension to be matched.

[0293] The normalization process and the calculation processes of the weights and the comprehensive degree of match will be introduced in detail below:

[0294] Suppose there are m candidate visual media and n dimensions to be matched. The initial matching degrees of the visual contents of the m candidate visual media on each dimension to be matched among the n dimensions to be matched can be regarded as a data matrix:

[0295] X = (x ij ) m*n , i = 1, 2, …, m; j = 1, 2, …, n (5)

[0296] where x ij is the initial matching degree of the visual content of the i-th candidate visual media on the j-th dimension to be matched.

[0297] Continuing with the above example, there are multiple candidate visual media including Picture 1, Picture 2, and Picture 3, and multiple dimensions to be matched including Dimension A, Dimension B, and Dimension C. That is, m is 3 and n is 3.

[0298] Step 1: Perform normalization processing on the above data matrix.

[0299] The normalized matrix is:

[0300] R = (r ij ) m*n , i = 1, 2, …, m; j = 1, 2, …, n (6)

[0301] where

[0302]

[0303] where max(x j ) refers to the maximum value among the initial matching degrees of the visual contents of multiple candidate visual media on the j-th dimension to be matched; min(x j ) refers to the minimum value among the initial matching degrees of the visual contents of multiple candidate visual media on the j-th dimension to be matched.

[0304] In practical applications, other normalization methods can also be used, and the embodiments of this application do not make specific limitations on this.

[0305] Exemplarily, the results obtained by normalizing the example data in Table 1 above are shown in Table 2.

[0306] Table 2:

[0307] Dimension A Dimension B Dimension C Picture 1 0.00 0.00 1.00 Picture 2 1.00 1.00 0.00 Picture 3 0.80 0.67 0.23

[0308] Step 2: Calculate the information entropy corresponding to each dimension to be matched.

[0309] Formulas (8) and (9) can be used for calculation:

[0310]

[0311] Among them,

[0312]

[0313] Among them, e j refers to the information entropy corresponding to the j-th dimension to be matched.

[0314] Among them, since the domain of the ln(x) function is x > 0. In actual calculation, to avoid the situation where p ij in ln(p ij ) takes 0, ln(p ij ) in formula (8) can be replaced by ln(p ij +α), where α << 0.001.

[0315] The smaller the information entropy corresponding to the dimension to be matched, the greater the degree of variation in the matching degree of the visual content of multiple candidate visual media in this dimension to be matched, and the greater the amount of information provided. It can be considered that the role played by this dimension to be matched in the comprehensive evaluation is also greater.

[0316] Exemplarily, for the example data in Table 2 above, the information entropy calculated according to the above formula (8) and formula (9) is shown in Table 3:

[0317] Table 3:

[0318] Dimension A Dimension B Dimension C Information Entropy 0.69 0.67 0.48

[0319] Step 3: Calculate the weights corresponding to each dimension to be matched.

[0320] The weights corresponding to each dimension to be matched can be calculated using formula (10):

[0321]

[0322] Among them, d j refers to the weight corresponding to the j-th dimension to be matched.

[0323] It can be seen that the above formula (10) is a monotonically decreasing function of information entropy. In this embodiment, the smaller the information entropy, the greater the degree of variation; the greater the degree of variation, the greater the weight. That is to say, the weight is negatively correlated with the information entropy.

[0324] To ensure that the sum of the weights corresponding to multiple dimensions to be matched is 1, formula (11) can be used for the following calculation to obtain the final weight w j :

[0325]

[0326] Among them, w jRefers to the final weight corresponding to the j-th dimension to be matched.

[0327] In an alternative embodiment, other monotonically decreasing functions may also be used to calculate the above weights, and the embodiments of the present application do not make specific limitations thereto.

[0328] Exemplarily, for the example data in Table 3 above, the weights of different dimensions to be matched can be calculated according to the above formulas (10) and (11), as shown in Table 4:

[0329] Table 4:

[0330] Dimension A Dimension B Dimension C Weight 0.28 0.29 0.42

[0331] It should be noted that since rounding is introduced in the process of calculating the information entropy, the sum of the three dimensions in Table 4 above is not 1.

[0332] Step 4: Weighted summation.

[0333] Perform weighted summation on the normalized matching degrees of any candidate visual media to obtain the comprehensive matching degree of the candidate visual media, which can be specifically calculated using formula (12):

[0334]

[0335] where s i Refers to the comprehensive matching degree of the i-th visual media.

[0336] Exemplarily, for the example data in Tables 2 and 4 above, the comprehensive matching degrees of different pictures are calculated using the above formula (12), as shown in Table 5:

[0337] Table 5:

[0338] Picture 1 Picture 2 Picture 3 Comprehensive Matching Degree 0.42 0.58 0.52

[0339] In the above 5102, when the first branch recalls visual media, it only considers the visual semantic information of the visual media, without considering the attribute information such as the time and location of the visual media. Therefore, it is necessary to perform time filtering, location filtering, etc. on the visual media recalled by the first branch. Specifically, when the user performs semantic search, if the search statement contains time information, pictures that do not meet the time limit need to be filtered out.

[0340] However, the user's description form of time information is rich and diverse and is fuzzy. For example, when the user searches for "the sky photographed in the afternoon", the "afternoon" in the user's search statement has no clear and standardized definition. When the user uses a fuzzy time expression in the search statement, the time window used for time filtering will affect the user experience.

[0341] Exemplary: The user started traveling to other places on September 30, 2022 and returned home on October 9, 2022. During this period, the user took a lot of photos. One day in 2023, the user wanted to view the beautiful scenery taken during this trip. Then the user was likely to enter the search statement "Scenery taken during last year's National Day holiday". If directly based on "last year's National Day" in the search statement, the time window used for time filtering was set to "from October 1, 2022 to October 7, 2022", then the scenic photos taken by the user on September 30, 2022, October 8, 2022, and October 9, 2022 would be filtered out, which obviously did not meet the user's expectations. To improve the rationality of time filtering, the embodiments of the present application provide a new time filtering method. Specifically, using a clustering algorithm, multiple candidate visual media are clustered according to the acquisition time of each of the multiple candidate visual media, and K (K≥1) clustering clusters are obtained.

[0342] In this way, pictures of the same series with relatively close acquisition times can be grouped into the same clustering cluster.

[0343] Continuing with the above example, the natural scenery picture taken by the user on September 30, 2022 and the natural scenery picture taken by the user on October 1, 2022 have relatively close shooting times. Through the above clustering algorithm, they are grouped into the same clustering cluster.

[0344] The above clustering algorithm may include but is not limited to: K-Means clustering algorithm, Mean shift clustering algorithm, and density-based clustering algorithm.

[0345] Taking the density-based clustering algorithm as an example, in the scenario of semantic search including time information, it is unreasonable to set fixed first parameter ∈ and second parameter MinPts. For example, when the user searches for "Photos taken during the outing in 2022", the time range is 1 whole year; while when the user searches for "Photos taken during the morning outing", the time range is several hours. The same first parameter ∈ and second parameter MinPts should not be set in these two cases. To improve the rationality of clustering, the following steps can be used to determine the first parameter ∈ and the second parameter MinPts:

[0346] 51021. Determine the first parameter involved in the density-based clustering algorithm according to the time search range included in the search statement.

[0347] Among them, the first parameter is positively correlated with the duration corresponding to the time search range.

[0348] The time search range can be determined according to the time-related semantic entities in the search statement. Specifically, the time range corresponding to the time-related semantic entity can be returned through a mapping table (i.e., the time search range). Among them, the mapping table can be constructed in advance as needed, and the specific form is not specifically limited in the embodiments of the present application.

[0349] Exemplarily, the time-related semantic entity extracted from the search statement "photos taken in spring" is "spring". Querying the mapping table, the corresponding time range is obtained as: February 1 - May 30; the time-related semantic entity extracted from the search statement "the sky taken in the morning" is "morning". Querying the mapping table, the corresponding event range is obtained as: 7:00 - 12:00; the time-related semantic entity extracted from the search statement "photos of going out for fun during the National Day" is "National Day". Querying the mapping table, the corresponding time range is obtained as: October 1 - October 7; the time-related entity extracted from the search statement "the sky taken in Beijing this year" is "this year". Querying the mapping table, the corresponding time range is obtained as "January 1, 2023 to December 31, 2023".

[0350] The first parameter can be determined according to the duration corresponding to the time search range. Exemplarily, the start time of the time search range (which can be understood as the start timestamp) is T start and the end time (which can be understood as the end timestamp) is T end , and the duration of the time search range is: T end -T start , and the following formula can be used to calculate the first parameter:

[0351] ∈=α*(T end -T start ) (13)

[0352] Among them, α is a coefficient that can be adjusted manually, and its value can be set according to actual needs, which is not specifically limited in the present application.

[0353] 51022. Determine the second parameter involved in the density-based clustering algorithm according to the ratio of the number of multiple candidate visual media to the time search range.

[0354] The following formula can be used to calculate the second parameter:

[0355]

[0356] Among them, N is the total number of visual media in the candidate set whose acquisition time is within the time search range, N≥1; among them, n is the number of times the time search range repeats between T 1 and T 2 . T 1is the acquisition time of the earliest acquired visual media among multiple visual media stored via the mobile phone; T 2 is the acquisition time of the latest acquired visual media among multiple visual media stored via the mobile phone. Exemplarily, the search statement is "the sky photographed during the National Day", and its time search range is from October 1st to October 7th, with a duration of 7 days; the acquisition time of the earliest acquired visual media stored in the mobile phone is August 1st, 2020; the acquisition time of the latest acquired visual media stored in the mobile phone is October 20th, 2023; then, from August 1st, 2020 to October 20th, 2023, the time search range from October 1st to October 7th repeats 4 times (i.e., once a year).

[0357] where, (T end - T start ) * n can be understood as the total duration corresponding to the time search range.

[0358] where, β is a coefficient that can be adjusted manually, and its magnitude can be set according to actual needs. This application does not make specific limitations on this.

[0359] In this embodiment, according to the duration defined by the time search range included in the search statement, the first parameter ∈ and the second parameter MinPts involved in the density-based clustering algorithm are dynamically adjusted, which can ensure the rationality of clustering and thus improve the accuracy of the final search result.

[0360] When the search statement includes a time search range, determine the acquisition time range corresponding to each of the K (K≥1) clustering clusters; filter out the clustering clusters whose acquisition time range does not overlap with the time search range, and retain the clustering clusters whose acquisition time range overlaps with the time search range.

[0361] The overlap between the acquisition time range and the time search range can be partial overlap or full overlap. Whether it is partial overlap or full overlap, both belong to having an overlapping part.

[0362] In this way, G (G≥1) clustering clusters are selected from the K clustering clusters. The acquisition time of the earliest acquired visual media among these G clustering clusters is before the start time of the time search range, and / or, the acquisition time of the latest acquired visual media among these G clustering clusters is after the end time of the time search range. Note: There is no intersection among the G clustering clusters, and there is also no intersection among the acquisition time ranges of the G clustering clusters themselves.

[0363] The acquisition time range corresponding to a clustering cluster is from the acquisition time T min T1 of the earliest acquired visual media in this clustering cluster to the acquisition time T max of the latest acquired visual media in this clustering cluster, that is: [Tmin , T max .

[0364] Specifically, the time search range is [T start , T end , and the acquisition time range corresponding to the clustering cluster is [T min , T max . When [T min , T max and [T start , T end have an overlapping part, keep this clustering cluster; when [T min , T max and [T start , T end do not have an overlapping part, filter out this clustering cluster. For example: the time search range is from October 1, 2022 to October 7, 2022, and the acquisition time range of the clustering cluster is from September 30, 2022 to October 1, 2022. These two ranges have an overlapping part (i.e., October 1, 2022), and this clustering cluster is retained.

[0365] Continuing with the above example, the natural scenery pictures taken by the user on September 30, 2022 and the natural scenery pictures taken by the user on October 1, 2022 are relatively close in shooting time. Through the above clustering algorithm, they are grouped into the same clustering cluster. Assuming that this clustering cluster contains only these two pictures, then the acquisition time range of this clustering cluster is: from September 30, 2022 to October 1, 2022, and it has an overlapping part with the time search range: from October 1, 2022 to October 7, 2022. Therefore, this clustering cluster will be retained, that is to say, the natural scenery pictures taken by the user on September 30, 2022 will not be filtered out.

[0366] Similarly, the natural scenery pictures taken by the user on October 8, 2022 and the natural scenery pictures taken by the user on October 7, 2022 are relatively close in shooting time. Through the above clustering algorithm, they are grouped into the same clustering cluster. Assuming that this clustering cluster contains only these two pictures, then the acquisition time range of this clustering cluster is the pictures from October 7, 2022 to October 8, 2022, and it has an overlapping part with the time search range: from October 1, 2022 to October 7, 2022. Therefore, this clustering cluster will be retained. That is to say, the natural scenery pictures taken by the user on October 8, 2022 will not be filtered out.

[0367] In this way, for the search statement "scenery taken during last year's National Day holiday", among the multiple clustering clusters obtained by screening, the acquisition time of the earliest acquired visual media is September 30, 2022, and the acquisition time of the latest acquired visual media is October 8, 2022.

[0368] It can be seen that using the time filtering method provided by the embodiments of the present application can ensure that a series of photos with relatively close acquisition times are presented to the user, guarantee the coherence of the search results, and improve the user's search experience.

[0369] In addition, when the search statement also contains a semantic entity related to a location, location filtering can be further performed on the G clustered clusters filtered out. Specifically, visual media in the clustered clusters with acquisition locations that do not match the semantic entity related to the location in the search statement can be filtered out. Exemplarily, the acquisition location "Prince Kung's Mansion, Xicheng District, Beijing" and the semantic entity "Sujiatuo Town, Haidian District, Beijing" can be considered to match (belonging to municipal-level matching); the acquisition location "Tsinghua University, Haidian District, Beijing" and the semantic entity "Sujiatuo Town, Haidian District, Beijing" can be considered to match (belonging to district-level matching); the acquisition location "Shanghai" and the semantic entity "Beijing" can be considered not to match.

[0370] In the above embodiments, clustering is performed first, then time filtering, and finally location filtering. Of course, in practical applications, location filtering can also be performed first, then clustering, and finally time filtering. The specific execution order of these three steps can be set according to actual needs, and the embodiments of the present application do not make specific limitations on this.

[0371] 5103. Filtering of relationships between people.

[0372] When the search statement contains a semantic entity related to the relationship between people, visual media in each clustered cluster with a people relationship attribute that does not match the semantic entity are filtered out. Exemplarily, if the search statement contains "good friends" and the person name attribute of Picture B is "colleague", which do not match, then Picture B is filtered out.

[0373] 5104. Filtering of person names.

[0374] When the search statement includes a semantic entity related to a person name, visual media in each clustered cluster with a person name attribute that does not match the semantic entity are filtered out. Exemplarily, if the search statement contains "Zhang San" and the person name attribute of Picture A is "Li Si", which do not match, then Picture A is filtered out.

[0375] It should be added that the semantic entity related to the person name and the semantic entity related to the relationship between people in the search statement can also be identified through named entity recognition technology. The person name attribute and the people relationship attribute of the visual media are manually added by the user for the visual media in advance.

[0376] 511. Sorting.

[0377] For the candidate set obtained in step 508, the visual media in the candidate set may be sorted according to the collection time of the visual media in the candidate set to obtain the display order of the visual media. The collection time of the visual media at the front of the display order is earlier than the collection time of the visual media at the back of the display order. Subsequently, the mobile phone may display the candidate set according to the display order of the visual media in the candidate set.

[0378] The filtered candidate set obtained in step 510 (including the G clusters) may be sorted in one of the following ways:

[0379] Method 1: Sort the visual media in the G clusters from high to low according to the target matching degree of the visual media in the G clusters to obtain the display order of the visual media in the G clusters. The target matching degree can be the matching degree between the visual content of the visual media and the search sentence in the above text or the comprehensive matching degree in the above text (the specific calculation method can refer to the corresponding content in the above embodiment). Subsequently, the mobile phone can display the visual media in the G clusters according to the display order of the visual media in the G clusters.

[0380] Method 2: Sort the G clusters according to the start time of their collection time ranges to obtain the display order of the G clusters (i.e., inter-cluster sorting). Figure 6 In the interface 601 shown, the collection time range corresponding to cluster A is from September 30, 2022 to October 1, 2022; the collection time range corresponding to cluster B is from October 3, 2022 to October 5, 2022; then, the display order of cluster A precedes the display order of cluster B. For each cluster, the visual media in the cluster are sorted from high to low according to the target matching degree of the visual media in the cluster, and the display order between the visual media in the cluster (intra-cluster sorting) is obtained; or, for each cluster, the visual media in the cluster are sorted according to the collection time of the visual media in the cluster, and the display order of the visual media in the cluster is obtained. Subsequently, the mobile phone displays the visual media in the G clusters according to the inter-cluster sorting and intra-cluster sorting.

[0381] It should be noted that in order to better protect user privacy and security, meet the principle of minimizing user data, and avoid reporting user data to the cloud side as much as possible, the entire search process mentioned above is completed on the terminal side.

[0382] In addition, the present application provides an electronic device, including: a memory, a processor and a display, wherein the memory is used to store programs; the processor is coupled to the memory and the display, and is used to execute the program stored in the memory to implement the above-mentioned visual media search method.

[0383] The embodiments of the present application also provide a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a computer, one or more steps in any of the above visual media search methods can be implemented.

[0384] The computer-readable storage medium may be a non-transitory computer-readable storage medium. For example, the non-transitory computer-readable storage medium may be ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0385] Another embodiment of the present application also provides a computer program product containing instructions. When the computer program product is executed by a computer, one or more steps in any of the above methods can be implemented.

[0386] Among them, the electronic device, computer-readable storage medium, and computer program product provided in this embodiment are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above, and will not be elaborated here.

[0387] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces. The indirect coupling or communication connection of devices or units may be in electrical, mechanical or other forms.

[0388] The units described as separate components may or may not be physically separated. The components displayed as units may be one physical unit or multiple physical units, that is, they may be located in one place, or may be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0389] In addition, each functional unit in various embodiments of the present application may be integrated in a processing unit, or each unit may exist physically separately, or two or more units may be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0390] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods of the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0391] The above content is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A visual media search method applicable to an electronic device, characterized in that, it includes: displaying a first interface; the first interface includes a search box; receiving a search operation on a search statement input into the search box; if the semantic proportion related to visual content in the search statement is greater than or equal to a preset proportion threshold, determining a search result of the search statement according to multiple first visual media files; the multiple first visual media files are determined from the multiple visual media according to the matching situation between the search statement and the visual content of the multiple visual media; the visual content is data that needs to be obtained through a natural picture understanding model; if the semantic proportion related to visual content in the search statement is less than the preset proportion threshold, determining a search result of the search statement according to multiple second visual media; the multiple second visual media files are determined from the multiple visual media according to the matching situation between the semantic subject unrelated to visual content in the search statement and the attributes of the multiple visual media; displaying the search result.

2. The method according to claim 1, characterized in that, it further includes: removing the semantic subject unrelated to visual content in the search statement to obtain a rewritten search statement; determining multiple reference visual media from the multiple visual media according to the matching situation between the rewritten search statement and the visual content of the multiple visual media; determining a first representative visual semantic vector corresponding to the multiple first visual media files and a second representative visual semantic vector corresponding to the multiple reference visual media; determining the semantic proportion related to visual content in the search statement according to the vector similarity between a first difference vector and a second difference vector; the first difference vector represents the difference between the first sentence semantic vector and the first representative visual semantic vector; the second difference vector represents the difference between the second sentence semantic vector and the second representative visual semantic vector.

3. The method according to claim 2, characterized in that, removing the semantic subject unrelated to visual content in the search statement to obtain a rewritten search statement, including: removing the semantic subject related to time and / or the semantic subject related to location in the search statement to obtain a rewritten search statement.

4. The method according to any one of claims 1 to 3, characterized in that, it further includes: performing semantic understanding on the search statement to obtain a first sentence semantic vector; acquiring visual semantic vectors of the multiple visual media files; the visual semantic vector of each visual media file is obtained by performing semantic understanding on the image or image frame of the visual media file using a natural picture understanding model; determining M candidate visual media files with visual semantic vectors matching the first sentence semantic vector from the multiple visual media files; where M is an integer greater than 1; determining the multiple first visual media files according to the M candidate visual media files.

5. The method according to claim 4, characterized in that, determining the multiple first visual media files according to the M candidate visual media files, including: Determine N dimensions to be matched corresponding to the search statement according to the semantic entities related to visual content in the search statement; N is greater than 1 and is an integer; Obtain the weights of the N dimensions to be matched; wherein, the weight of the jth dimension to be matched is positively correlated with the degree of variation of the matching degree of the visual content of the M candidate visual media files in the jth dimension to be matched; j is an integer, and the value of j ranges from 1 to N in sequence; According to the weights of the N dimensions to be matched, perform weighted summation on the matching degrees of the visual content of the ith candidate visual media file in each of the N dimensions to be matched, to obtain the comprehensive matching degree of the ith candidate visual media file; i is an integer, and the value of i ranges from 1 to M in sequence; Determine the multiple first visual media files whose comprehensive matching degrees meet the preset requirements from the M candidate visual media files.

6. The method according to claim 5, wherein, further comprising: Match the search terms in the search statement with a plurality of preset tags, to determine whether the search terms belong to semantic entities related to the tags; The tags are used to describe visual content; Determine the semantic entities related to the tags in the search statement as the semantic entities related to visual content in the search statement.

7. The method according to any one of claims 1 to 3, wherein, further comprising: Obtain the attributes of the multiple visual media files; Determine the multiple second visual media whose attributes match the semantic entities unrelated to visual content in the search statement from the multiple visual media files.

8. The method according to any one of claims 1 to 3, wherein, Determining the search result of the search statement according to the multiple first visual media files includes: If the search statement includes a semantic entity related to time, filter the multiple first visual media files according to the semantic entity related to time and the acquisition time attribute of the multiple first visual media files; If the search statement includes a semantic entity related to location, filter the multiple first visual media files according to the semantic entity related to location and the acquisition location attribute of the multiple first visual media files; If the search statement includes a semantic entity related to the relationship between people, filter the multiple first visual media files according to the semantic entity related to the relationship between people and the relationship between people attribute of the multiple first visual media files; and / or If the search statement includes a semantic entity related to a person's name, filter the multiple first visual media files according to the semantic entity related to the person's name and the person's name attribute of the multiple first visual media files.

9. An electronic device, wherein, comprising: a memory, a processor and a display, wherein, The memory is used to store programs; The processor is coupled to the memory and the display, and is used to execute the programs stored in the memory to implement the method according to any one of claims 1 to 8.

10. A computer-readable storage medium storing a computer program, Characterized in that, When the computer program is executed by a computer, it can implement the method described in any one of claims 1 to 8.