Visual media searching method and device and storage medium

By obtaining the matching degree and ranking of visual media files on multiple dimensions to be matched, correcting the number of occurrences, and calculating the score, the problem of inaccurate visual media search in the prior art is solved, and the user search experience is improved.

CN120067371AActive Publication Date: 2025-05-30HONOR DEVICE CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202311571222.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-21
Publication Date
2025-05-30
Estimated Expiration
2043-11-21

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively support users to accurately search pictures and videos through complex search statements, especially in the multi-dimensional matching ranking of visual media, which lacks reasonable scoring statistics, resulting in poor user search experience.

Method used

Provide a visual media search method, by obtaining the matching degree and ranking of visual media files on multiple dimensions to be matched, correcting the number of occurrences of visual media files, calculating scores, and sorting or filtering search results based on the scores.

Benefits of technology

It improves the statistical rationality of scores of visual media, enhances the user's search experience, and can more accurately match users' complex search needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067371A_ABST
    Figure CN120067371A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a visual media searching method and device and a storage medium. The method is suitable for the electronic equipment and comprises the following steps: displaying a first interface; the first interface comprises a search box; receiving a search operation of a search statement input to the search box; determining W to-be-matched dimensions according to the search statement; obtaining the matching degree and the matching degree ranking of the Y visual media files on the wth dimension to be matched; according to the matching degree of the yth visual media file on each to-be-matched dimension in the W to-be-matched dimensions, correcting the total number of times that the yth visual media file appears as the uth visual media file; determining the score of the yth visual media according to the corrected total number of times that the yth visual media file appears in each rank; determining a search result of the search statement according to the scores of the Y visual media files; and displaying the search result. According to the technical scheme provided by the embodiment of the invention, the reasonability of score statistics of each visual media can be improved, so that the search experience of a user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of terminals, and in particular, to a visual media search method, device, and storage medium. Background Art

[0002] With the popularization of intelligent terminals, more and more users use intelligent terminals, such as mobile phones, to take pictures and videos, and store the taken pictures and videos in the picture gallery of the electronic device, so as to record every bit of life. In addition, users can also download pictures, take screenshots of the mobile phone interface, and store the downloaded pictures and screenshots in the picture gallery of the electronic device.

[0003] In order to facilitate users to manage and view pictures in the terminal, a picture management function and a picture search function are configured in the picture gallery application or other similar applications of the terminal device. For example: the picture gallery application in the terminal can classify the pictures in the terminal according to information such as the time and location of picture taking to generate corresponding albums, and users can view relevant pictures by searching for information such as time and location. Summary of the Invention

[0004] Multiple aspects of this application provide a visual media search method, device, and storage medium, which can improve the rationality of score statistics of each visual media, and further improve the user search experience.

[0005] In a first aspect, a visual media search method applicable to an electronic device is provided, including:

[0006] Display a first interface; the first interface includes a search box;

[0007] Receive a search operation on a search statement input into the search box;

[0008] Determine W dimensions to be matched according to the search statement; W is an integer greater than 1;

[0009] Obtain the matching degree and matching degree ranking of Y visual media files in the w-th dimension to be matched; Y is an integer greater than 1; w is an integer ranging from 1 to W;

[0010] According to the matching degree of the y-th visual media file in each dimension to be matched among the W dimensions to be matched, correct the total number of times the y-th visual media file appears in the u-th place, and obtain the corrected total number of times the y-th visual media file appears in the u-th place; where u is an integer ranging from 1 to Y; y is an integer ranging from 1 to Y;

[0011] Determine the score of the y-th visual media file according to the corrected total number of times the y-th visual media file appears in each ranking;

[0012] Determine the search results of the search statement according to the scores of the Y visual media files;

[0013] Display the search results.

[0014] Specifically, the Y visual media files can be sorted according to the scores of the Y visual media files to obtain the search results of the search statement; or, the Y visual media files can be filtered according to the scores of the Y visual media files to obtain the search results of the search statement.

[0015] That is to say, in the visual media search scenario, when counting the scores of each visual media file based on the matching degree ranking of the visual media file in each dimension to be matched, the matching degree size of the visual media file in each dimension to be matched is additionally introduced to correct the total number of times the visual media appears in each ranking, which helps to improve the rationality of score statistics and thus improve the user search experience.

[0016] It can be understood that this solution is a comprehensive ranking of multi-dimensional fusion of visual media files based on an improved borda counting method.

[0017] In a possible implementation manner, according to the matching degree of the y-th visual media file in each dimension to be matched among the W dimensions to be matched, correct the total number of times the y-th visual media file appears in the u-th place to obtain the corrected total number of times the y-th visual media file appears in the u-th place, including:

[0018] Determine the first weight of the w-th dimension to be matched for the y-th visual media file according to the matching degree of the y-th visual media file in the w-th dimension to be matched;

[0019] Perform a weighted sum of the number of times the y-th visual media file appears in the u-th place in each dimension to be matched among the W dimensions to be matched according to the first weights of the W dimensions to be matched for the y-th visual media file respectively, so as to obtain the corrected total number of times the y-th visual media file appears in the u-th place.

[0020] In a possible implementation manner, determining the first weight of the w-th dimension to be matched for the y-th visual media file according to the matching degree of the y-th visual media file in the w-th dimension to be matched includes:

[0021] Obtain the preset weight of the w-th dimension to be matched;

[0022] According to the preset weight of the w-th dimension to be matched, correct the matching degree of the y-th visual media file in the w-th dimension to be matched, so as to obtain the corrected matching degree of the y-th visual media file in the w-th dimension to be matched;

[0023] Determine the corrected matching degree of the y-th visual media file in the w-th dimension to be matched as the first weight of the w-th dimension to be matched for the y-th visual media file.

[0024] That is to say, weight configuration channels are reserved for each dimension in the comprehensive sorting, and the preset weights of each dimension can be set according to prior knowledge.

[0025] In a possible implementation manner, the determination method of the dimension to be matched includes multiple of the following methods:

[0026] Determine the search statement as the dimension to be matched;

[0027] Determine the semantic entity irrelevant to the visual content in the search statement as the dimension to be matched;

[0028] Determine the number of user clicks as the dimension to be matched.

[0029] Among them, the number of user clicks can represent the personalized preferences of users. When performing multi-dimensional comprehensive sorting on visual media, considering information such as the personalized preferences of users can improve the rationality of score statistics, and thus improve the user search experience.

[0030] In a possible implementation manner, the W dimensions to be matched include: the number of user clicks;

[0031] According to the number of times the y-th visual media file has been clicked by the user historically, determine the matching degree of the y-th visual media file in terms of the number of user clicks.

[0032] In a possible implementation manner, the W dimensions to be matched include: the search statement;

[0033] According to the vector similarity between the visual semantic vector of the y-th visual media file and the first sentence semantic vector of the search statement, determine the matching degree of the y-th visual media file in the search statement.

[0034] Among them, the matching degree of the y-th visual media file in the search statement represents the matching degree between the visual content of the visual media file and the entire search statement. That is to say, when performing comprehensive sorting, the matching degree between the visual content of the visual media file and the entire search statement will be considered.

[0035] In a possible implementation, among the W dimensions to be matched, it includes: the semantic entity in the search statement that is irrelevant to the visual content;

[0036] According to the matching situation of the semantic entity irrelevant to the visual content of the y-th visual media file's attributes, determine the matching degree of the y-th visual media file on the semantic entity irrelevant to the visual content in the search statement.

[0037] That is to say, when comprehensively sorting, the matching degrees of attributes such as the time and location of the visual media file with the search statement will be considered.

[0038] In a possible implementation, before correcting the total number of times the y-th visual media file appears in the u-th place according to the matching degrees of the y-th visual media file on each dimension to be matched among the W dimensions to be matched, the method further includes:

[0039] Perform percentage processing on the matching degrees of the Y visual media files on the w-th dimension to be matched.

[0040] In this solution, the matching degree on the dimension of the number of user clicks is the actual number of user clicks. To ensure the effectiveness of subsequent calculations, it is necessary to perform percentage processing on the matching degree on the dimension of the number of user clicks. Specifically, the optimal matching degree on the dimension of the number of user clicks can be set, and then the matching degree on the dimension of the number of user clicks is percentage-processed based on the optimal matching degree.

[0041] In a possible implementation, the method further includes:

[0042] Perform semantic understanding on the search statement to obtain the first sentence semantic vector;

[0043] Obtain the visual semantic vectors of multiple visual media files; the visual semantic vector of each visual media file is obtained by performing semantic understanding on the image or image frame of the visual media file using a natural picture understanding model;

[0044] From the multiple visual media files, determine multiple first visual media files whose visual semantic vectors match the first sentence semantic vector;

[0045] Determine the Y visual media files according to the multiple first visual media files.

[0046] In this solution, the above-mentioned Y visual media files are recalled by matching the visual semantic vectors of the visual media files with the sentence semantic vectors of the search statements. That is to say, the visual content of the above-mentioned Y visual media is semantically matched with the entire search statement, fitting the user's visual focus. Subsequently, only the visual media files that fit the user's visual focus need to be comprehensively sorted. In this way, not only can the search accuracy be guaranteed, but also the computational amount involved in the comprehensive sorting can be reduced, thereby shortening the search latency and reducing the power consumption.

[0047] In a second aspect, the present application provides an electronic device, including: a memory, a processor, and a display, where,

[0048] The memory is used to store programs;

[0049] The display is used to display a search page;

[0050] The processor is coupled to the memory and the display, and is used to execute the programs stored in the memory to implement any of the above-mentioned methods.

[0051] In a third aspect, the present application provides a computer-readable storage medium storing a computer program, where the computer program, when executed by a computer, can implement any of the above-mentioned methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation to the present application. In the drawings:

[0053] Figure 1A A set of interface diagrams of the search interface for a mobile phone to enter the gallery application provided by an embodiment of the present application;

[0054] Figure 1B A schematic diagram of the search interface after clearing the search history provided by an embodiment of the present application;

[0055] Figure 1C A set of interface diagrams involved in the search in the gallery application provided by an embodiment of the present application;

[0056] Figure 1D A schematic diagram of the search failure interface provided by an embodiment of the present application;

[0057] Figure 2A A schematic diagram of the structure of an electronic device provided by another embodiment of the present application;

[0058] Figure 2B A software structure block diagram of an electronic device provided by another embodiment of the present application;

[0059] Figure 3A Another set of interface diagrams related to search in the gallery application provided by an embodiment of the present application;

[0060] Figure 3B A set of interface diagrams related to search in the negative first screen provided by an embodiment of the present application;

[0061] Figure 3C Schematic diagram one of the search result interface provided by an embodiment of the present application;

[0062] Figure 3D Schematic diagram two of the search result interface provided by an embodiment of the present application;

[0063] Figure 3E Schematic diagram three of the search result interface provided by an embodiment of the present application;

[0064] Figure 3F Schematic diagram of the search result interface provided by an embodiment of the present application Figure Four ;

[0065] Figure 4 Interaction diagram of the visual media search method provided by an embodiment of the present application;

[0066] Figure 5 Flow schematic diagram of the visual media search method provided by an embodiment of the present application;

[0067] Figure 6 Schematic diagram of the search result interface provided by an embodiment of the present application Figure Five ;

[0068] Figure 7 Flow schematic diagram of the visual media search method provided by an embodiment of the present application. Detailed implementation manners

[0069] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application. Among them, in the description of the embodiments of the present application, unless otherwise specified, " / " means "or", for example, A / B may mean A or B; herein, "and / or" is only a description of the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B may mean: A exists alone, A and B exist simultaneously, and B exists alone these three situations. In addition, in the description of the embodiments of the present application, "a plurality of" means two or more than two.

[0070] Hereinafter, the terms "first", "second", and "third" are only used for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first", "second", and "third" may explicitly or implicitly include one or more of such features.

[0071] First, the vocabulary involved in the embodiments of this application will be described. It can be understood that this description is for a clearer understanding of the embodiments of this application and does not necessarily constitute a limitation on the embodiments of this application.

[0072] Visual media: refers to pictures or videos.

[0073] Semantic entity: The named entity recognition technology can identify text and recognize entities with specific meanings in the text, such as person names (PER), place names (LOC), etc. In this solution, the entities with specific meanings identified are called semantic entities.

[0074] Visual content-related and visual content-unrelated: Visual content refers to the objects presented by visual media and their interrelationships, etc. Computer vision can endow a computer with capabilities similar to human vision, including perceiving, understanding, analyzing, and interpreting visual content. Currently, the Generative Pre-trained Transformer 4 (GPT-4) can support inputting an image into the model and then outputting human natural language describing the important information in the picture.

[0075] In the context of image search in this solution, the data that a visual media file needs to go through the natural picture understanding of the model to obtain is called "visual content-related". In the following text, visual content is the data that needs to be obtained through the natural picture understanding model. This solution refers to the data that is related to the visual media file and can be obtained without going through the picture understanding ability of the model as "visual content-unrelated", such as the shooting location, shooting time, name, file attributes, etc. that can be obtained and saved when the terminal device collects the visual media file.

[0076] For example, in the "photo taken in Beijing this year", "this year" (shooting time), "Beijing" (shooting location), and "photo" (file attribute) are all data that can be obtained and saved when the terminal device collects the visual media file. Therefore, "this year", "Beijing", and "photo" are visual content-unrelated; in the "sky taken in Beijing this year", "sky" needs to be obtained through the picture understanding ability of the model to understand the image or image frame of the visual media file. Therefore, "sky" is visual content-related.

[0077] Text semantic vector: It can be obtained by sending text into a text encoder, which is a vector capable of representing the semantic features of the entire sentence. The text encoder can adopt models such as Transformer commonly used in Natural Language Processing (NLP). This solution places no restrictions here. In this solution, the text semantic vector obtained for a sentence is called a sentence semantic vector, the text semantic vector obtained for the semantic subject in a sentence is called a subject semantic vector, and the text semantic vector obtained for a label is called a label semantic vector.

[0078] Visual semantic vector: It can be obtained by sending the image or image frame of a visual media file into an image encoder. Commonly used CNN (Convolutional Neural Network) models or VIT (Vision Transformer) models can be adopted. This solution places no restrictions here.

[0079] Density-based clustering algorithm: It is based on a set of neighborhoods to describe the tightness of a sample set. (The first parameter ∈, the second parameter MinPts) is used to describe the tightness of the sample distribution in the neighborhood. The first parameter ∈ is used to describe the neighborhood radius of a data point; the second parameter MinPts is used to describe the minimum number of data points in the neighborhood of a data point. Its representative algorithms include: DBSCAN (Density-Based Spatial Clustering of Application with Noise); the DBSCAN algorithm is a relatively representative density-based clustering algorithm that can divide regions with sufficient high density into clusters and can discover clusters of any shape in a spatial database with noise;

[0080] Vector similarity: It is used to describe the similarity between two vectors (for example: between a sentence semantic vector and a visual semantic vector). In the embodiments of this application, the visual media that matches the search statement can be determined by comparing the similarity between the sentence semantic vector and the visual semantic vector. Generally, the vector similarity can be calculated through the cosine similarity calculation formula. Of course, it can also be calculated through other methods.

[0081] In the prior art, a mobile phone manages visual media files such as pictures and videos of users through a gallery application (hereinafter referred to as: visual media). Taking the example of a mobile phone taking a photo, after the mobile phone takes a photo, the gallery application can obtain and save attributes unrelated to the visual content such as the shooting location, shooting time, and photo name of the photo.

[0082] In practical applications, when the mobile phone is in the state of charging and the screen is off, photos can also be input into the natural image understanding model to generate and save tags for the photos by the natural image understanding model. These tags can be regarded as attributes related to the visual content of the photos, and the tags can be: "sky", "cat", "dog", etc. The gallery application can establish an index for the photos based on their attributes. After the index is established, the gallery application can provide corresponding search services to users. Specifically, users can search for pictures or videos by entering keywords in the gallery application. Exemplarily, users can enter keywords such as "Beijing", "sky", "National Day" in the search box provided by the gallery application. The gallery application matches the keywords entered by the user with the indexes of visual media such as pictures and videos in the gallery application to obtain search results.

[0083] The following describes the interfaces involved in the search process of the gallery application in the prior art with reference to the accompanying drawings:

[0084] As Figure 1A shown in (a) of [reference], the mobile phone can display the main interface 101, which can also be called the desktop. The main interface 101 can include the icon 102 of the gallery application. The mobile phone can receive the operation of the user clicking on the icon 102. In response to this operation, the mobile phone can start the gallery application and display the interface 103 as shown in Figure 1A (b) of [reference]. Among them, the interface 103 can be an album interface. It should be noted that in response to the operation of the user clicking on the icon 102, the mobile phone can start the gallery application and display the photo interface of the gallery. The photo interface includes thumbnails of the photos (i.e., pictures) in the gallery or the large picture of a certain photo. In the photo interface, in response to the operation of the user on the "Album" control, the above-mentioned album interface 103 is displayed.

[0085] As Figure 1A shown in (b) of [reference], the interface 103 includes multiple albums. Among them, the "All Photos" album includes 2,023 photos, the "Camera" album includes 1,502 photos and videos, the "Screenshots & Screen Recordings" album includes 102 photos and videos, the "My Favorites" album includes 48 photos and videos, the "One-Click Multi-Capture" album has 34 photos and videos, the "Video Editing" album has 65 videos, the "Self-created Album" has 57 photos and videos, and the "Shared Album" has 100 photos and videos.

[0086] As Figure 1A shown in (b) of [reference], the interface 103 can include a search box 104. The mobile phone can receive the operation of the user clicking on the search box 104. In response to this operation, the mobile phone can display as Figure 1AThe interface 105 shown in (c) therein can be referred to as a search interface. Among them, the interface 105 can display classification information of photos to the user. For example, in the interface 105, the mobile phone classifies the photos of the local machine according to time, portraits, and things, etc. For example, in the dimension of time, the mobile phone classifies the photos of the local machine according to three time periods: "this month", "last month", and "this year". Among them, the "this month" album includes the photos or videos taken by the mobile phone this month, the "last month" album includes the photos or videos taken by the mobile phone last month, and the "this year" album includes the photos or videos taken by the mobile phone this year. In the dimension of portraits, the mobile phone classifies the photos of the local machine according to different people, such as the four different people in the interface 105. In the dimension of things, the mobile phone classifies and displays the photos of the local machine according to "scenery", "animals", "documents", and "buildings". It should be noted that the above classification dimensions can also be others, and no specific restrictions are made here. In the interface 105, the user can see this classification information without entering keywords.

[0087] Optionally, the interface 105 can also include options for search history 107 and "clear" 108. The search history includes the keywords that the user has entered, such as "flowers", "coffee", "cat", etc. The mobile phone can receive the operation of the user clicking "clear" 108, and in response to this operation, the mobile phone can clear the search history. After the mobile phone clears the search history, the keywords that the user has entered are no longer displayed on the search interface 105. For example, in response to the operation of the user clicking "clear" 108, as Figure 1B shown, the search history 107 and the option of "clear" 108 are no longer displayed on the search interface 105, and the content displayed below moves up.

[0088] In response to the operation of the user entering the keyword "sky" in the interface 105, the mobile phone displays the interface 109 shown in (a) of Figure 1C it. As Figure 1C shown in (a) of it, 100 photos related to "sky" and 32 photos related to the photos containing the word "sky". Among them, the 100 photos related to "sky" can be recalled because the tags of these 100 photos match "sky"; the 32 photos related to the photos containing the word "sky" can be recalled because through the OCR (Optical Character Recognition) technology, it is recognized that these 32 photos contain words such as "sky". In practical applications, the mobile phone can also associate the keywords entered by the user to obtain associated words and perform searches based on the associated words.

[0089] The interface 109 also displays some search results related to the keyword "sky" and a "More" option 110 corresponding to the search results of the keyword "sky". The mobile phone receives a click operation from the user on the "More" option 110 and displays an interface 111 as shown in Figure 1C (b) of the figure. Among them, the interface 111 is used to display photos and videos in the search results of the keyword "sky". Optionally, the photos and videos can be classified and displayed according to time. In addition, the interface 111 also includes a return key 112 and a title 113. In response to the user's operation on the return key 112, the mobile phone can redisplay the interface 109. The title 113 may include the keyword "sky".

[0090] That is to say, in the existing gallery applications, when a user enters a simple keyword in the search box, such as: sky, Beijing, National Day, etc., corresponding search results can be obtained. However, because the number of photo attributes is simple and limited, and the mobile phone's ability to understand and associate with search statements is also limited. When the user enters a more complex search statement in the search box, if the keywords in the search statement cannot match the attributes of the picture or the text in the picture, no photos can be found. That is to say, the existing gallery applications do not support the search function based on complex search statements. As Figure 1D shown, when the user enters a more complex search statement "warming oneself by the fire and brewing tea" in the search box of the interface 114, the mobile phone cannot understand the associated words of "warming oneself by the fire and brewing tea", and since the photos do not have attributes that can match "warming oneself by the fire and brewing tea" or its associated words, the search result shows "no pictures".

[0091] However, in practical applications, users have a strong demand for the function of searching for pictures based on complex search statements. This is because users can describe the pictures or videos they want more comprehensively through complex search statements, thereby achieving precise search. To meet this demand of users, an embodiment of the present application provides a visual media search method. This method can be applied to an electronic device, and the electronic device can be a mobile phone, a tablet computer, a wearable device, a vehicle-mounted device, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), and other terminal devices.

[0092] Exemplarily, Figure 2AThe structural schematic diagram of the electronic device 200 is shown. The electronic device 100 may include a processor 210, an external memory interface 220, an internal memory 221, a universal serial bus (USB) interface 230, a charging management module 240, a power management module 241, a battery 242, an antenna 1, an antenna 2, a mobile communication module 250, a wireless communication module 260, an audio module 270, a speaker 270A, a receiver 270B, a microphone 270C, a headphone interface 270D, a sensor module 280, a key 290, a motor 291, an indicator 292, a camera 293, a display screen 294, and a subscriber identification module (SIM) card interface 295, etc.

[0093] Among them, the sensor module 280 may include a pressure sensor 280A, a gyroscope sensor 280B, a barometric pressure sensor 280C, a magnetic sensor 280D, an acceleration sensor 280E, a distance sensor 280F, a proximity light sensor 280G, a fingerprint sensor 280H, a temperature sensor 280J, a touch sensor 280K, an ambient light sensor 280L, a bone conduction sensor 280M, etc.

[0094] It can be understood that the structure schematically shown in the embodiments of the present application does not constitute a specific limitation on the electronic device 200. In other embodiments of the present application, the electronic device 200 may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0095] The processor 210 may include one or more processing units. For example, the processor 210 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.

[0096] The controller may generate an operation control signal according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.

[0097] A memory can also be set in the processor 210 for storing instructions and data. In some embodiments, the memory in the processor 210 is a cache memory. This memory can save the instructions or data that the processor 210 has just used or recycled. If the processor 210 needs to use the instruction or data again, it can be directly called from the memory. This avoids repeated accesses, reduces the waiting time of the processor 210, and thus improves the efficiency of the system.

[0098] In some embodiments, the processor 210 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0099] It can be understood that the interface connection relationships between the modules illustrated in the embodiments of the present application are only illustrative descriptions and do not constitute a structural limitation on the electronic device 200. In other embodiments of the present application, the electronic device 200 may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.

[0100] The electronic device 200 realizes the display function through a GPU (Graphics Processing Unit), a display screen 294, and an application processor, etc. The GPU is connected to the display screen 294 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 210 may include one or more GPUs, which execute program instructions to generate or change display information.

[0101] The display screen 294 is used to display images, videos, etc. The display screen 294 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include one or N display screens 294, where N is a positive integer greater than 1.

[0102] The electronic device 200 can implement the shooting function through an ISP, a camera 293, a video codec, a GPU, a display screen 294, an application processor, etc.

[0103] The camera 293 is used to capture still images or videos. An object generates an optical image through a lens and projects it onto a photosensitive element. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the ISP to be converted into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard format such as RGB, YUV, etc. In some embodiments, the electronic device 100 may include one or N cameras 293, where N is a positive integer greater than 1.

[0104] The video codec is used to compress or decompress digital videos. The electronic device 200 can support one or more video codecs. In this way, the electronic device 200 can play or record videos in multiple coding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0105] The NPU is a neural-network (NN) computing processor. By drawing on the structure of the biological neural network, such as the transmission pattern between human brain neurons, it can quickly process the input information and can also continuously learn on its own. Through the NPU, applications such as intelligent cognition of the electronic device 200 can be realized, such as image recognition, face recognition, voice recognition, text understanding, etc.

[0106] The external memory interface 220 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 200. The external memory card communicates with the processor 210 through the external memory interface 220 to implement the data storage function. For example, files such as music and videos are saved in the external memory card.

[0107] The internal memory 221 can be used to store computer-executable program code, and the executable program code includes instructions. The internal memory 221 can include a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, image playback function, etc.). The data storage area can store the data created during the use of the electronic device 200 (such as audio data, phone book, etc.). In addition, the internal memory 221 can include high-speed random access memory and can also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. The processor 210 executes various functional applications and data processing of the electronic device 200 by running the instructions stored in the internal memory 221 and / or the instructions stored in the memory provided in the processor.

[0108] Figure 2B It is a software structure block diagram of the electronic device 200 in the embodiment of the present application. The software system of the electronic device 200 can adopt a layered architecture. The layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. Taking the Android system as an example, in some embodiments, the Android system can be divided into four layers, from top to bottom are the application layer, application framework layer, Android runtime and system libraries, and kernel layer.

[0109] As Figure 2B shown, the application layer can include application programs such as a gallery service module, a search module, a multimodal understanding module, a natural language understanding module, and a camera application.

[0110] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions.

[0111] The system library can include multiple functional modules. For example: surface manager, Media Libraries, 3D graphics processing library (such as: OpenGL ES), 2D graphics engine (such as: SGL), etc. Among them, the media library supports the playback and recording of multiple common audio and video formats, as well as static image files, etc. The media library can support multiple audio and video coding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.

[0112] The kernel layer is the layer between hardware and software.

[0113] Next, in combination with the capture and photo-taking scenario, the working processes of the software and hardware of the electronic device 200 will be exemplarily described.

[0114] When the touch sensor 280K receives a touch operation, the corresponding hardware interrupt is sent to the kernel layer. The kernel layer processes the touch operation into a raw input event (including information such as touch coordinates and the timestamp of the touch operation). The raw input event is stored in the kernel layer. The application framework layer obtains the raw input event from the kernel layer and identifies the control corresponding to the input event. Taking the touch operation as a touch click operation and the control corresponding to the click operation being the control of the camera application icon as an example, the camera application calls the interface of the application framework layer to start the camera application, and then starts the camera driver by calling the kernel layer to capture a static image or video through the camera 293.

[0115] Taking the electronic device as a mobile phone as an example, the visual media search method provided by the embodiments of the present application will be introduced. The visual media search method provided by the embodiments of the present application can be applied in application programs such as the gallery application and the file management application.

[0116] Next, the interfaces and search logics involved in the visual media search method provided by the embodiments of the present application will be described in conjunction with the accompanying drawings.

[0117] As Figure 3A shown in (a) of [reference], the interface 301 displays search history 303 and a "Clear" option 304. The search history includes search statements that the user has entered, such as: "Watching the sunrise on the mountain top", "The Great Wall photographed in Beijing this year". Other contents that can be displayed on the interface 301 can refer to the relevant display contents of the above interface 305 and will not be elaborated here. Among them, the interface 301 can be called a search interface. The mobile phone can respond to the user's operation onFigure 1A By clicking the search box 104 in the interface 103 shown in (b), the interface 301 is displayed.

[0118] like Figure 3A In the interface 305 (i.e., the search result interface) shown in (b) of FIG. 305, the user enters the search sentence "cooking tea around the fire" in the search box 306 of the interface 305, and the mobile phone searches for 239 pictures. The mobile phone displays part of the search results for the search sentence "cooking tea around the fire" (e.g., thumbnails of 8 pictures) and the "more" option 307 corresponding to the search results for the search sentence "cooking tea around the fire" on the interface 305. In response to the user's operation on the "more" option 307, the mobile phone displays the following Figure 3A Interface 308 shown in (c) of FIG. Interface 308 can display the images in the search results of the search sentence "cooking tea around the fire" in descending order according to the matching degree between the visual content of the images and the search sentence "cooking tea around the fire" (specifically, the similarity between the visual semantic vector of the image and the sentence semantic vector of the search sentence). The user can also perform an upward sliding operation on interface 308 to view the images that are not displayed.

[0119] At present, mobile phones are also equipped with a negative screen, a drop-down search interface, etc. It can be understood that the negative screen can be the leftmost split screen of the electronic device, which is used to provide users with functions such as search and quick services. Among them, the negative screen can also be used to display notification messages that need to be pushed to users, such as application messages subscribed by users, real-time hot search messages, segment selection, itinerary information, etc. The drop-down search interface is an interface displayed in response to the user's drop-down operation performed on the main interface. This interface is used to provide users with functions such as search and application suggestions. This interface is similar to Figure 3B The interface 315 in can be the same interface.

[0120] The following will take the negative one screen as an example for introduction. When the user needs to view the negative one screen of the mobile phone, the electronic device can display the negative one screen by sliding the screen of the mobile phone.

[0121] For example, refer to Figure 3B As shown in (a) of FIG. 1 , the mobile phone can receive a first operation performed by the user on the interface 309 (which can be called the desktop) of the mobile phone. For example, the first operation can be as follows: Figure 3B In response to the first operation, the mobile phone may display the following Figure 3B The negative first screen 310 shown in (b) of FIG. The negative first screen 310 may include: a search box 311, a quick service 312, a default card 313, a recommended card 314, etc. The quick service 312 may be a quick entry to a page or function of an application, such as: scan, payment code, ride code, etc.; the default card may be: a gallery card, a remaining battery card, etc.; the recommended card may be a recommended application card.

[0122] The mobile phone receives a click operation by the user on the search box 311 on the negative first screen 310, and displays the interface 315 shown in (c) below. The interface 315 may include: a search box 316, application suggestions, and a search history 217. The application suggestions include: icons of various applications recommended for use. The interface 215 may also include: the search history 217 and its corresponding "Clear" option 318. In response to the user's trigger operation on the "Clear" option 318, the search history 317 and the "Clear" option 318 are no longer displayed on the interface 315. In addition, a hot news title may be displayed in the search box 316, such as: "Tianjin Marathon". Figure 3B As shown in (d) below, in the interface 319, the search box of the interface 319 displays the search statement "The Great Wall photographed in Beijing during the National Day" input by the user; the interface 319 also displays a preview area 322 of the search results of the image gallery for the search statement "The Great Wall photographed in Beijing during the National Day" and the "Search in App" option 323 corresponding to the image gallery. In response to the user's trigger operation on the preview area 322, the mobile phone enters the photo details interface provided by the image gallery application for the user to flip through the search results of the search statement "The Great Wall photographed in Beijing during the National Day". In response to the user's trigger operation on the "Search in App" option 323, the mobile phone displays the interface 324 provided by the image gallery application shown in (e) below. The interface 324 displays some search results of the search statement "The Great Wall photographed in Beijing during the National Day" and the "More" option corresponding to the search results of the search statement "The Great Wall photographed in Beijing during the National Day". In response to the user's trigger operation on the "More" option, the mobile phone may display the search result details interface, and the search result details interface displays the pictures in the search results of the search statement "The Great Wall photographed in Beijing during the National Day". The interface 319 may also display an online search option 321. In response to the user's trigger operation on the online search option 321, the mobile phone displays a search web page and displays the online search results on the search web page.

[0123] As Figure 3B shown in (d) below Figure 3B As shown in (e) below, the interface 324 provided by the image gallery application shown in (e) below. The interface 324 displays some search results of the search statement "The Great Wall photographed in Beijing during the National Day" and the "More" option corresponding to the search results of the search statement "The Great Wall photographed in Beijing during the National Day". In response to the user's trigger operation on the "More" option, the mobile phone may display the search result details interface, and the search result details interface displays the pictures in the search results of the search statement "The Great Wall photographed in Beijing during the National Day". The interface 319 may also display an online search option 321. In response to the user's trigger operation on the online search option 321, the mobile phone displays a search web page and displays the online search results on the search web page.

[0124] As Figure 3CThe interface 325 shown. The user enters the search statement "The Great Wall photographed during last year's National Day" in the search box of the interface 325, and the mobile phone searches for 419 pictures. The mobile phone displays the thumbnails of each picture or video in the search results of the search statement "The Great Wall photographed during last year's National Day" on the interface 325. Among them, the shooting time of the picture or video corresponding to thumbnail A is 23:22 on September 30, 2022; the shooting time of the picture or video corresponding to thumbnail B is 00:19 on October 1, 2022. Among them, the time search range corresponding to National Day is from 00:00 on October 1, 2022 (i.e., the first time point) to 00:00 on October 8, 2022 (i.e., the second time point). The shooting time of the picture or video corresponding to thumbnail A is before 00:00 on October 1, 2022. The shooting time of the picture or video corresponding to thumbnail B is 00:19 on October 1, 2022, which is between 00:00 on October 1, 2022 and 00:00 on October 8, 2022. Among them, the shooting times of the pictures or videos corresponding to thumbnail A and thumbnail B are relatively close.

[0125] As Figure 3D For the interface 326 shown, for the search statement "The sky photographed during last year's National Day", the shooting time of the picture or video corresponding to the displayed thumbnail C is 22:19 on October 7, 2022; the shooting time of the picture or video corresponding to the displayed picture D is 01:24 on October 8, 2022. Among them, the time search range corresponding to National Day is from 00:00 on October 1, 2022 (i.e., the first time point) to 00:00 on October 8, 2022 (i.e., the second time point). The shooting time of the picture or video corresponding to thumbnail D is after 00:00 on October 8, 2022. The shooting time of the picture or video corresponding to thumbnail C is 00:19 on October 1, 2022, which is between 00:00 on October 1, 2022 and 00:00 on October 8, 2022. Among them, the shooting times of the pictures or videos corresponding to thumbnail C and thumbnail D are relatively close.

[0126] In practical applications, the search statement may include, in addition to the time search range (such as National Day), a first keyword (such as the Great Wall or the sky). Then, for the search statement, the visual media file corresponding to its search result matches the first keyword.

[0127] When the first keyword belongs to a semantic entity unrelated to visual content (such as: location), the first keyword can be matched with the attributes of the visual media to determine the visual media file that matches the first keyword.

[0128] When the first keyword belongs to a semantic entity related to visual content (such as the Great Wall or the sky), match the first keyword with the tags of the visual media file to determine the visual media file whose visual content matches the first keyword; or, match the text semantic vector of the first keyword with the visual semantic vector of the visual media file to determine the visual media file whose visual content matches the first keyword (including the above-mentioned first, second, and third visual media files).

[0129] The above thumbnail is a scaled thumbnail of the corresponding visual media file or a scaled thumbnail of any image frame.

[0130] As Figure 3E Shown in the interface 327, in which, for the search statement "the sky photographed during last year's National Day", the two thumbnails E and F shown are sorted and displayed in descending order according to the matching degree between the visual content of their respective visual media files and the first keyword "sky". Among them, the matching degree between the visual content of the visual media file corresponding to the thumbnail E and "sky" is 0.87; the matching degree between the visual content of the visual media file corresponding to the thumbnail F and "sky" is 0.76.

[0131] As Figure 3F Shown in the interface 328, in which, for the search statement "photos taken in Nankai District, Tianjin during last year's National Day", the two thumbnails J and H shown are sorted and displayed in descending order according to the matching degree between the location attributes of their respective visual media files and the first keyword "Nankai District, Tianjin". Among them, the matching degree between the location attribute of the visual media file corresponding to the thumbnail J and "Nankai District, Tianjin" is 0.87; the matching degree between the location attribute of the visual media file corresponding to the thumbnail H and "Nankai District, Tianjin" is 0.8.

[0132] Figure 4 This is the interaction diagram of the visual media search method provided by the embodiments of the present application. As Figure 4 Shown, on the mobile phone, there are provided: a gallery service module (i.e., a gallery application) 41, a search module 42, a multimodal understanding module 43, and a natural language understanding module 44.

[0133] As Figure 4 Shown, the visual media search method provided by the embodiments of the present application can be divided into two stages: an index construction stage and a search stage.

[0134] In the index construction stage, the following steps are included:

[0135] S401, Add and / or modify visual media and its attributes.

[0136] In the above S401, for the newly added visual media, the attributes that the gallery application can automatically generate and are unrelated to the visual content may include, but are not limited to: the collection location, the collection time, and the name of the visual media. Taking the captured video or picture as an example, the collection location refers to the shooting location, and the collection time refers to the shooting time; taking the screenshot as an example, the collection location refers to the screenshot location, and the collection time refers to the screenshot time; taking the downloaded video or picture as an example, the collection location refers to the download location, and the collection time refers to the download time.

[0137] Users can add new visual media by means such as shooting, downloading, and taking screenshots. In addition, users can also modify the existing visual media. Such modifications include, but are not limited to: beautification, custom naming, adding watermarks, and other operations.

[0138] S402. Store the visual media and its attributes.

[0139] The gallery service module 41 can, in response to the above-mentioned addition or modification operations, store the visual media and its attributes locally on the mobile phone. In actual applications, with the user's authorization, the mobile phone can store the locally stored visual media and its attributes in the cloud to relieve the storage pressure on the local mobile phone.

[0140] S403. Request visual semantic understanding of the visual media.

[0141] Since visual semantic understanding requires a large amount of computing resources, in order not to affect the user's use, the above step S403 can be executed when the mobile phone is in the state of being charged and the screen is off.

[0142] The gallery service module 41 can request the multimodal understanding module 43 to perform visual semantic understanding on the newly added or modified visual media to obtain the visual semantic vector of the visual media.

[0143] Among them, the multimodal understanding module 43 can perform visual semantic understanding on the visual media based on the multimodal model to obtain the visual semantic vector of the visual media.

[0144] Among them, the multimodal model is not only used for: performing visual semantic understanding on the visual media to obtain the visual semantic vector of the visual media; but also for: performing semantic understanding on the search statement to obtain the sentence semantic vector of the search statement; performing semantic understanding on the rewritten search statement in the following text to obtain the sentence semantic vector of the rewritten search statement; performing semantic understanding on the semantic subject in the search statement to obtain the subject semantic vector of the semantic subject; performing semantic understanding on the label of the visual media to obtain the label semantic vector of the label. The multimodal model can be trained based on training samples.

[0145] Exemplarily, the multimodal model may specifically be CLIP (Contrastive Language-Image Pre-training). The CLIP model can map visual media and text (i.e., search statements, rewritten search statements, semantic entities, tags) into a unified vector space to understand the relationships between different modal resources visually and textually, and then be used for image retrieval. That is, in the embodiments of the present application, the visual media file and text can be specifically matched through the CLIP model.

[0146] The multimodal model can map visual media and text into vectors of the same dimension. That is to say, the dimension of the visual semantic vector of visual media is the same as that of the semantic vector of text (e.g., the sentence semantic vector of the search statement). The multimodal model includes: the above-mentioned image encoder and text encoder.

[0147] S404. Return the visual semantic vector of the visual media.

[0148] The multimodal understanding module 44 returns the visual semantic vector of the visual media to the gallery service module 41.

[0149] S405. Store the visual semantic vector of the visual media.

[0150] The gallery service module 41 can locally store the visual semantic vector of the visual media.

[0151] S406. Send the attribute information of the visual media and its visual semantic vector.

[0152] Exemplarily, the gallery service module 41 can store the visual semantic vector of the visual media returned by the multimodal understanding module 44, and then batch send the attributes of the visual media and its visual semantic vector to the search module 42 for the search module 42 to construct an index of the visual media.

[0153] S407. Construct an index

[0154] The index of the visual media constructed by the search module 42 may include: the attributes of the visual media, the visual semantic vector of the visual media.

[0155] In the search stage, the following steps are included:

[0156] S408. Input a search statement.

[0157] The user can input a search statement through the search interface provided by the gallery service module 41, for example: Figure 3A the interface 301 shown in (a) of Figure 3AFor the interface 305 shown in (b) in [context], enter "warming the tea around the stove" in the search box 306.

[0158] S409: Send the search statement.

[0159] After the gallery service module 41 receives the search statement input by the user, it sends the search statement to the search module 42 for searching.

[0160] S410: Request semantic entity recognition for the search statement.

[0161] The search module 42 requests the natural language understanding module 44 to perform semantic entity recognition on the search statement to obtain the semantic entities contained in the search statement. Among them, the natural language understanding module 44 performs semantic entity recognition based on a natural language understanding model. Specifically, named entity recognition technology (NER) can be used to perform semantic entity recognition on the search statement to obtain the semantic entities contained in the search statement. In the embodiments of the present application, semantic entities can also be referred to as entities.

[0162] Using named entity recognition technology, semantic entities related to time, location, and tags in the search statement can be recognized. Among them, semantic entities related to time and location are semantic entities unrelated to visual content; the tags of visual media files are data that can only be obtained through the natural picture understanding of the model. Therefore, semantic entities related to tags are semantic entities related to visual content. In practical applications, based on practical experience, multiple tags that users are more concerned about can be statistically obtained, such as: "sky", "cat", "dog", "birthday", "child", etc. These tags are used to describe visual content. Subsequently, named entity recognition technology can match the keywords (or search terms) in the search statement with multiple pre-set tags to determine whether the keyword belongs to a semantic entity related to the tag.

[0163] Exemplarily, using named entity recognition technology to perform semantic entity recognition on the search statement "the sky photographed in Beijing during the National Day", it is determined that "National Day" belongs to a semantic entity related to time, "Beijing" belongs to a semantic entity related to location, and "sky" belongs to a semantic entity related to the tag.

[0164] S411: Return the semantic entity.

[0165] The natural language understanding module 44 returns the recognized semantic entity to the search module 42.

[0166] S412: Request semantic understanding of the search statement.

[0167] The search module 42 can send the search statement to the multimodal understanding module 43, and the multimodal understanding module 43 performs semantic understanding on the search statement to obtain the sentence semantic vector of the search statement (i.e., the first sentence semantic vector). For the specific semantic understanding process, reference can be made to the corresponding content in the above embodiments, which will not be elaborated here.

[0168] It should be additionally supplemented that when the search statement includes a semantic entity related to visual content (i.e., a semantic entity related to a label), the search module 42 can also send the semantic entity related to visual content to the multimodal understanding module 43, so that the multimodal understanding module 43 performs semantic understanding on the semantic entity related to visual content to obtain the entity semantic vector of this semantic entity. Continuing with the above example, "sky" belongs to the semantic entity related to visual content, and the multimodal understanding module 43 can perform semantic understanding on "sky" to obtain the entity semantic vector corresponding to "sky".

[0169] S413. Return the text vector.

[0170] The multimodal understanding module 43 can return the sentence semantic vector of the search statement to the search module 42.

[0171] S414. Perform recall respectively based on the attributes of the visual media and the visual semantic vector of the visual media.

[0172] The search module 42 includes different branches of search methods:

[0173] Exemplarily, the first branch is: performing recall based on the visual semantic vector of the visual media. Specifically, obtain the visual semantic vectors of each visual media among the multiple visual media stored in the mobile phone; calculate the vector similarity between the sentence semantic vector of the search statement and the visual semantic vectors of each visual media; according to the vector similarity, determine M candidate visual media (i.e., M candidate visual media files) whose visual semantic vectors match the sentence semantic vector from the multiple visual media stored in the mobile phone; where M is an integer greater than or equal to 1; these M visual media can be used as the visual media set recalled by the first branch.

[0174] Exemplarily, use the visual media with a vector similarity greater than the preset similarity threshold as the visual media whose visual semantic vector matches the sentence semantic vector.

[0175] Exemplarily, sort the multiple visual media according to the vector similarity from high to low, and use the top F (F≥1) visual media as the visual media whose visual semantic vector matches the sentence semantic vector.

[0176] Exemplarily, Q (Q≥1) visual media with vector similarity greater than a preset similarity threshold are determined from multiple visual media stored via the mobile phone; if Q is greater than or equal to a preset quantity threshold D, these Q visual media can be sorted in descending order of vector similarity; the top D visual media in the sorting are used as the visual media whose visual semantic vectors match the semantic vector of this sentence; if Q is less than the preset quantity threshold D, these Q visual media can be directly used as the visual media whose visual semantic vectors match the semantic vector of this sentence.

[0177] It should be noted that the multiple visual media stored via the mobile phone may include: visual media stored locally by the mobile phone and / or visual media stored in the cloud by the mobile phone (for example: the cloud storage space applied for by the mobile phone). For the purpose of protecting user privacy, all the multiple visual media stored via the mobile phone are stored in the mobile phone.

[0178] It should be noted that in the recall process of the first branch, the attributes of the visual media are not understood, and only the overall visual semantic information of the visual media can be understood.

[0179] Exemplarily, the second branch is: recall based on the attributes of the visual media.

[0180] Specifically, obtain the attributes of each visual media among the multiple visual media stored via the mobile phone; match the semantic subject in the search statement with the attributes of each visual media to determine the visual media (i.e., the second visual media) that matches this semantic subject; use the visual media that matches this semantic subject as the set of visual media recalled by the second branch. When the number of semantic subjects in the search statement is one, the set of visual media recalled by the second branch includes the visual media that matches this one semantic subject; when the number of semantic subjects in the search statement is multiple, the set of visual media recalled by the second branch includes the visual media that each of these multiple semantic subjects matches.

[0181] Exemplarily, for a semantic subject related to time, the corresponding time range (i.e., the time search range) of this semantic subject can be determined; match this time range with the acquisition time attribute of each visual media to determine the visual media whose acquisition time is within this time range; use the visual media whose acquisition time is within this time range as the visual media that matches this semantic subject. For example: the semantic subject is "National Day", its corresponding time range is "from October 1st to October 7th", the acquisition time of Picture 1 is "October 2nd", and the acquisition time of Picture 2 is "October 8th", then, according to the above matching method, Picture 1 matches the semantic subject "National Day", and Picture 2 does not match the semantic subject "National Day".

[0182] Exemplarily, for a semantic entity related to a location, the geographical range corresponding to the semantic entity (i.e., the geographical search range) can be determined; the geographical range is matched with the attribute of the collection location of each visual medium to determine the visual media whose collection location is within the geographical range; the visual media whose collection location is within the geographical range is used as the visual media matched by the semantic entity. For example: the semantic entity is "Beijing", its corresponding geographical range is the whole Beijing City, the collection location of Picture 3 is "Xicheng District, Beijing City", and the collection location of Picture 4 is "Nankai District, Tianjin City". Then, according to the above matching method, Picture 3 is matched with the semantic entity "Beijing", and Picture 4 is not matched with the semantic entity "Beijing".

[0183] Exemplarily, for a semantic entity related to a tag, the main semantic vector of the semantic entity and the tag semantic vectors of the tags of each visual medium can be obtained; the vector similarity between the main semantic vector of the semantic entity and the tag semantic vectors of the tags of each visual medium is calculated; according to the vector similarity, the visual media matched by the semantic entity is determined. For example: the semantic entity is "human cub", and Picture 5 has a tag of "child". Through calculation, it is found that the main semantic vector of "human cub" is similar to the tag semantic vector of "child", that is, Picture 5 is matched with the semantic entity "human cub". In practical applications, after obtaining the set of visual media recalled by the first branch and the set of visual media recalled by the second branch, the candidate set of visual media can be determined according to the set of visual media recalled by the first branch and the set of visual media recalled by the second branch. In an optional implementation manner, the union or intersection of the set of visual media recalled by the first branch and the set of visual media recalled by the second branch can be used as the candidate set of visual media.

[0184] In practical applications, when users search for pictures and the like on their mobile phones, sometimes they focus on the visual semantic information of the pictures, sometimes they focus on the attribute information such as the shooting location and shooting time of the pictures, and sometimes they focus on both. Exemplarily, when the user searches for "pictures taken today", the user focuses on the shooting time of the pictures; when the user searches for "the sky taken today", the user not only focuses on the shooting time of the pictures, but also focuses on the visual semantics of the pictures, that is, whether the picture content is the sky; when the user searches for "pictures taken while walking in Beijing", the user focuses on the shooting location of the pictures; when the user searches for "pictures of walking taken in Beijing", the user not only focuses on the shooting location attribute of the pictures, but also focuses on the visual semantics of the pictures, that is, whether the picture content is a walking picture.

[0185] Taking the two search statements of "photos taken in Beijing this year" and "the sky taken in Beijing this year" as examples, referring to the foregoing introduction, the semantic proportion of visual content in the search statement of "the sky taken in Beijing this year" is greater than that in the search statement of "photos taken in Beijing this year". Obviously, for the search statement of "photos taken in Beijing this year", it is more appropriate to use the visual media set recalled by the second branch as the candidate visual media set. Subsequently, the intersection of the visual media matching "this year" and the visual media matching "Beijing" in the candidate visual media set can be used as the final search result. The collection locations of all visual media in the final search result are Beijing, and the collection time is this year, which meets the user's search requirements. If the visual media set recalled by the first branch is used as the candidate visual media set, the following situation may occur: There is a picture of "a girl holding a camera to take pictures" in the visual media set recalled by the first branch, but the shooting location of this photo is Shanghai and the shooting time is last year. Since the word "take" exists in the search statement of "photos taken in Beijing this year", and the action of "taking" exists in the picture of "a girl holding a camera to take pictures", there is a certain similarity between the sentence semantic vector of the search statement of "photos taken in Beijing this year" and the visual semantic vector of the picture of "a girl holding a camera to take pictures". That is to say, the picture of "a girl holding a camera to take pictures" may be recalled. That is, when the semantic proportion of relevant visual content is relatively low, it is inappropriate to use the visual media set recalled by the first branch as the candidate visual media set.

[0186] In order to solve the above problems, in an optional implementation manner, the semantic proportion related to visual content in the search statement can be determined; according to the semantic proportion related to visual content in the search statement, the first extraction ratio corresponding to the visual media set recalled by the first branch and the second extraction ratio corresponding to the visual media set recalled by the second branch are determined; when the semantic proportion is greater than or equal to the preset proportion threshold, the first extraction ratio is greater than the second extraction ratio. Exemplarily, if the semantic proportion is S, the first extraction ratio is: β*S, and the second extraction ratio is: 1-β*S, where the value of β can be set according to actual needs, and the embodiments of the present application do not make specific limitations on this.

[0187] In another alternative embodiment, the semantic proportion related to visual content in the search statement can be determined; when the semantic proportion related to visual content in the search statement is greater than or equal to a preset proportion threshold, the visual media set recalled by the first branch is used as the candidate visual media set; when the semantic proportion related to visual content in the search statement is less than the preset proportion threshold, the visual media set recalled by the second branch is used as the candidate visual media set. That the semantic proportion related to visual content in the search statement is greater than or equal to the preset proportion threshold indicates that the search statement belongs to a high-semantic search statement; that the semantic proportion related to visual content in the search statement is less than the preset proportion threshold indicates that the search statement belongs to a low-semantic search statement. The specific process will be described in detail in the embodiments below in conjunction with Figure 5 will be introduced in detail.

[0188] S415. Visual media filtering.

[0189] The search module 42 can perform semantic subject filtering, spatio-temporal filtering, and / or person-relationship filtering on the candidate visual media set. The specific filtering method will be described in detail in the embodiments below in conjunction with Figure 5 will be introduced in detail.

[0190] S416. Visual media ranking.

[0191] The search module 42 ranks the candidate visual media set.

[0192] The specific ranking method will also be described in detail in the following embodiments.

[0193] S417. Return search results.

[0194] The search module 42 sends the search results obtained after ranking to the gallery service module 41.

[0195] Specifically, the search module 42 can select the top P (P≥1) visual media and their ranking information as the final search results.

[0196] S418. Display search results.

[0197] The gallery service module 41 can display the search results to the user. For example, through Figure 3A interface 305 in Figure 3A interface 308 in

[0198] As Figure 3A in interface 305, the search results include 239 pictures, and interface 305 only displays the thumbnails of 8 pictures in the search results; if the user wants to view these 239 pictures, the user can click the "More" option in interface 305, and in response to this operation, the mobile phone displays as Figure 3AThe interface 308 in it. In one example, the display order of pictures in the search results is related to the matching degree between the pictures and the search results. For example, the matching degree between the pictures displayed earlier and the search statement is greater than or equal to the matching degree between the pictures displayed later and the search statement. The calculation of the matching degree and the sorting method will be introduced in detail in the following embodiments.

[0199] Next, the search process executed by the search module 42 of the present application will be introduced in detail in conjunction with Figure 5 :

[0200] 501. Receive a search statement.

[0201] 502. Execute the search process corresponding to the search statement based on the visual semantic vector of the visual media.

[0202] The visual media set recalled by the first branch obtained by executing step 502 can be specifically referred to the search process corresponding to the first branch above.

[0203] Exemplarily, assume that the multimodal model can encode the visual media and the search statement into a k-dimensional vector space. The search statement is Q, and the corresponding sentence semantic vector V Q ={α z}, z = 1, 2,..., Z, and a total of M picture sets R are recalled, and their vector representations are:

[0204]

[0205] Among them, the i-th row corresponds to the visual semantic vector of the i-th picture, and the value range of i is [1, M].

[0206] 503. Semantic entity recognition.

[0207] For the specific process of semantic entity recognition of the search statement, reference can be made to the corresponding content in the above embodiments, and details will not be repeated here.

[0208] 504. Perform a search based on the attributes of the visual media.

[0209] Specifically, match the semantic entity irrelevant to the time content in the search statement with the attributes of the visual media to obtain the visual media set recalled by the second branch. For details, reference can be made to the search process corresponding to the second branch above.

[0210] 505. Rewrite the search statement.

[0211] Specifically, delete the semantic entity irrelevant to the visual content in the search statement to obtain the rewritten search statement. The semantic entity irrelevant to the visual content specifically refers to: the semantic entity related to time and the semantic entity related to location.

[0212] Exemplary: For the search statement "sky photographed this year", where "today" is the semantic entity related to time, the rewritten search statement is "photographed sky".

[0213] In practical applications, after deleting the semantic entities unrelated to visual content, there may be some redundant stop words. For example, for the search statement "sky photographed in Beijing this year", where "this year" is the semantic entity related to time and "Beijing" is the semantic entity related to location, after deleting "this year" and "Beijing", the stop word "in" becomes a redundant word and thus also needs to be deleted. Specifically, delete the semantic entities unrelated to visual content in the search statement and their related stop words to obtain the rewritten search statement. Exemplarily, the rewritten search statement corresponding to the search statement "sky photographed in Beijing this year" is "photographed sky".

[0214] It should be noted that there is no order restriction in the execution of the above steps 502, 503, and 505. In an optional example, to improve efficiency, these three steps can be executed simultaneously.

[0215] 506. Perform the search process corresponding to the rewritten search statement based on the visual semantic vector of the visual media.

[0216] Performing step 506 obtains the set of visual media recalled by the third branch.

[0217] Specifically, obtain the visual semantic vectors of each visual media among the multiple visual media stored in the mobile phone; calculate the vector similarity between the sentence semantic vector of the rewritten search statement (i.e., the second sentence semantic vector) and the visual semantic vectors of each visual media; determine, according to the vector similarity, multiple visual media (i.e., reference visual media) whose visual semantic vectors match the sentence semantic vector from the multiple visual media stored in the mobile phone; these multiple visual media can be used as the set of visual media recalled by the third branch.

[0218] Among them, the specific implementation process of the step of "determining, according to the vector similarity, multiple visual media whose visual semantic vectors match the sentence semantic vector from the multiple visual media stored in the mobile phone" can refer to the corresponding content in the above embodiments and will not be elaborated here.

[0219] Exemplarily, the rewritten search statement is Q′, and the corresponding vector V Q′ ={α′ z}, z = 1, 2, …, Z, and a total of M′ picture sets R′ are recalled, and their vector representation is:

[0220]

[0221] Among them, the i-th row corresponds to the visual semantic vector of the i-th picture, and the value range of i is [1, M′].

[0222] 507. Calculate the semantic proportion related to visual content in the search statement.

[0223] The following will introduce a method for determining the semantic proportion related to visual content:

[0224] 5071. Determine the representative visual semantic vector V I (i.e., the first representative visual semantic vector) corresponding to the visual media set recalled by the first branch and the representative visual semantic vector V I′ (i.e., the second representative visual semantic vector) corresponding to the visual media set recalled by the third branch.

[0225] 5072. Determine the sentence semantic vector V Q and the difference from the representative visual semantic vector V I to obtain a difference vector (V Q -V I )(i.e., the first difference vector).

[0226] 5073. Determine the sentence semantic vector V Q′ and the difference from the representative visual semantic vector V I′ to obtain a difference vector (V Q′ -V I′ )(i.e., the second difference vector).

[0227] 5074. Determine the semantic proportion related to visual content in the search statement according to the vector similarity between the difference vector (V Q -V I ) and the difference vector (V Q′ -V I′ ).

[0228] Among them, the semantic proportion related to visual content is positively correlated with this vector similarity.

[0229] In the above 5071, in one example, the vector similarity between the visual semantic vector of each visual media in the visual media set recalled by the first branch and the sentence semantic vector of the search statement can be obtained; the visual media in the visual media set recalled by the first branch are sorted according to the vector similarity from high to low; the average vector of the visual semantic vectors of the top T (T≥1) visual media in the sorting is used as the representative visual semantic vector V I .

[0230] Continuing with the above example, the representative visual semantic vector V I is:

[0231]

[0232] The vector similarity between the visual semantic vector of each visual medium recalled by the third branch and the sentence semantic vector of the rewritten search statement can be obtained; the visual media in the visual media set recalled by the third branch are sorted according to the vector similarity from high to low; the average vector of the visual semantic vectors of the top H (H≥1) visual media in the sorting is used as the representative visual semantic vector V I′ 。

[0233] Continuing with the above example: the representative visual semantic vector V I′ is:

[0234]

[0235] The values of H and T above can be the same or different, and the embodiments of the present application do not make specific limitations on this. In another example, according to the vector similarity between the visual semantic vector of each visual medium in the visual media set recalled by the first branch and the sentence semantic vector of the search statement, the visual media in the visual media set recalled by the first branch can be clustered using a clustering algorithm to obtain the clustering center point; the visual semantic vector of the clustering center point is used as the first representative visual semantic vector.

[0236] According to the vector similarity between the visual semantic vector of each visual medium in the visual media set recalled by the third branch and the second sentence semantic vector of the rewritten search statement, the visual media in the visual media set recalled by the third branch can be clustered using a clustering algorithm to obtain the clustering center point; the visual semantic vector of the clustering center point is used as the second representative visual semantic vector.

[0237] In the embodiments of the present application, the representative visual semantic vector is used to represent the visual semantics of the entire set. The specific clustering algorithm can be selected according to actual needs, and the embodiments of the present application do not make any limitations on this.

[0238] In the above 5074, in an optional embodiment, the vector similarity between the difference vector (V Q -V I ) and the difference vector (V Q′ -VI′) can be directly used as the semantic proportion related to the visual content in the search statement.

[0239] Among them, the vector similarity between the difference vector (V Q -V I ) and the difference vector (V Q′ -V I′ ), the larger it is, the greater the semantic proportion of the visual content in the search statement; conversely, it indicates that the semantic proportion of the visual content in the search statement is smaller.

[0240] It should be noted that in the embodiments of the present application, the calculation method of the semantic proportion related to visual content in the search statement draws on the word analogy feature of the distribution representation vector, that is, the additivity of word meanings is directly reflected in the additivity of the distribution representation vector.

[0241] When the semantic proportion related to visual content in the search statement is less than the preset proportion threshold, step 508 is executed subsequently.

[0242] When the semantic proportion related to visual content in the search statement is greater than or equal to the preset proportion threshold, steps 509 and 510 are executed subsequently.

[0243] 508. Determine the visual media recalled by the second branch as the candidate set.

[0244] 509. Determine the visual media recalled by the first branch as the candidate set.

[0245] 510. Visual media filtering.

[0246] To improve the search accuracy, one or more of the following processing operations can also be performed on the visual media recalled by the first branch: semantic entity filtering, spatio-temporal filtering, person name filtering, and person relationship filtering. Among them, semantic entity filtering refers to filtering the visual media recalled by the first branch by using the matching degree between the visual media recalled by the first branch and the semantic entity related to visual content in the search statement. Spatio-temporal filtering includes: time filtering and space (i.e., location) filtering. Person name filtering refers to filtering the visual media recalled by the first branch by using the person name attribute of the visual media recalled by the first branch and the matching degree with the semantic entity related to the person name in the search statement. Person relationship filtering refers to filtering the visual media recalled by the first branch by using the person relationship attribute of the visual media recalled by the first branch and the matching degree with the semantic entity related to the person relationship in the search statement. In practical applications, in addition to the above-mentioned attributes such as the collection location, collection time, and visual media name stored in the mobile phone, the visual media may also include the person name attribute and person relationship attribute manually input by the user. Therefore, in practical applications, the named entity recognition technology can also be used to match the keywords in the search statement with a variety of preset person relationships to determine whether the keyword belongs to the semantic entity related to the person relationship.

[0247] When multiple processing operations such as semantic entity filtering, spatio-temporal filtering, person name filtering, and person relationship filtering need to be performed on the candidate set recalled by the first branch, the execution order between these multiple processing operations can be set according to actual needs, and the embodiments of the present application do not make specific limitations on this.

[0248] Exemplarily, as Figure 5 shown, the visual media filtering includes:

[0249] 5101. Semantic entity filtering.

[0250] 5102. Spatiotemporal filtering.

[0251] 5103. Character relationship filtering.

[0252] 5104. Person name filtering.

[0253] In the above 5101, generally, when a user searches, the images they hope to retrieve contain some visual content described in their search query, such as: sky, puppy, child, etc. That is to say, the semantic entities related to visual content in the search query represent the visual focuses that the user pays attention to during the search.

[0254] However, when the first branch retrieves, it considers the matching degree between the visual content of visual media such as images and videos and the user's entire search query, without fully considering the role of certain specific semantic entities (i.e., the semantic entities related to visual content) in the search query in the mobile phone gallery search scenario, resulting in the retrieval of some inaccurate photos.

[0255] For example: When the user searches for "photos of the Great Wall taken on National Day", the semantic entity related to visual content among them includes "the Great Wall". When the first branch makes the match, it considers the matching degree between the entire search query and the visual content of the image, which may cause the first branch to retrieve images that contain the "taking" behavior but do not contain the "Great Wall". This is because the image contains the "taking" behavior and the search query contains the word "taking", that is, there is a certain similarity between the visual semantic vector of the image and the sentence semantic vector of the search query. Then, this image has the possibility of being retrieved.

[0256] Another example: When the user searches for "last year when the child had a birthday holding a cake", the semantic entities related to visual content among them include "child", "birthday", and "cake". When the first branch makes the match, it considers the matching degree between the entire search query and the image, which may cause the model to retrieve images of an adult having a birthday holding a cake. This is because the similarity between the visual semantic vector of the image of an adult having a birthday holding a cake and the sentence semantic vector of "last year when the child had a birthday holding a cake" is very high.

[0257] Therefore, in order to further improve the accuracy of searching for visual media on the mobile phone, on the basis of the first branch retrieving visual media, the "semantic entity" can be strengthened, that is, the results retrieved by the first branch are fine-tuned by means of the matching degree between the visual media and the "semantic entity".

[0258] Specifically, the "semantic entity filtering" in the above 5101 may include the following steps:

[0259] 5101a. Determine the dimension to be matched according to the semantic entity related to the visual content in the search statement.

[0260] 5101b. Filter the M candidate visual media according to the matching degree of the visual content of the M candidate visual media in the dimension to be matched.

[0261] In the above 5101a, in one example, the semantic entity related to the visual content is used as the dimension to be matched. When there are multiple semantic entities related to the visual content in the search statement, these multiple semantic entities are respectively used as different dimensions to be matched to obtain multiple dimensions to be matched. Among them, the number of multiple dimensions to be matched is the same as the number of semantic entities related to the visual content.

[0262] Exemplarily, for the search statement "The child held a cake on his / her birthday last year", the semantic entities related to the visual content are "child", "birthday", and "cake". Then, the corresponding three dimensions to be matched are: "child", "birthday", and "cake".

[0263] In practical applications, in addition to using the semantic entity as the dimension to be matched, the search statement itself can also be used as the dimension to be matched. In this way, when filtering the semantic entity, the matching situation between the visual content of the candidate visual media and the entire search statement can also be considered to improve the rationality of filtering. Specifically, the semantic entity related to the visual content and the search statement can be used as different dimensions to be matched to obtain multiple dimensions to be matched. When there are multiple semantic entities related to the visual content in the search statement, these multiple semantic entities and the search statement are respectively used as different dimensions to be matched. Among them, the number of multiple dimensions to be matched is one more than the number of semantic entities related to the visual content.

[0264] Exemplarily, for the search statement "The child held a cake on his / her birthday last year", the semantic entities related to the visual content are "child", "birthday", and "cake". Then, the corresponding four dimensions to be matched are: "child", "birthday", "cake", and "The child held a cake on his / her birthday last year".

[0265] In the above 5101b, the matching degree of the visual content of the candidate visual media in the dimension to be matched refers to the matching degree between the visual content of the candidate visual media and the dimension to be matched. The vector similarity between the visual semantic vector of the candidate visual media and the semantic vector to be matched in the dimension to be matched can be determined; among them, when the dimension to be matched is the semantic subject related to the visual content, the semantic vector to be matched in this dimension to be matched is the subject semantic vector of this semantic subject; when the dimension to be matched is the search statement, the semantic vector to be matched in this dimension to be matched is the sentence semantic vector of this search statement; according to this vector similarity, the matching degree of the visual content of the candidate visual media in the dimension to be matched is determined. Among them, the matching degree is positively correlated with the vector similarity.

[0266] When the number of dimensions to be matched is one, the candidate visual media with a matching degree less than or equal to the preset matching degree threshold can be filtered out according to the matching degrees of the visual contents of multiple candidate visual media in the dimension to be matched.

[0267] For example: the search statement is "photos of the sky taken on National Day", and the only semantic subject related to the visual content is "sky"; then, the recalled pictures that contain the "taking" behavior but do not contain the "sky" have a relatively low matching degree with the "sky" and will be filtered out.

[0268] When the number of dimensions to be matched is multiple, for each candidate visual media, the comprehensive matching degree of this candidate visual media is determined according to the matching degrees of the visual content of this candidate visual media in multiple dimensions to be matched; according to the comprehensive matching degrees of each candidate visual media, multiple candidate visual media are filtered. Specifically, from M candidate visual media, multiple first visual media with a comprehensive matching degree greater than or equal to the preset matching degree threshold (that is, meeting the preset requirements) are determined, which is equivalent to filtering out the candidate visual media with a comprehensive matching degree less than the preset matching degree threshold; or, the M candidate visual media are sorted in descending order of the comprehensive matching degree, and the top Z′ (Z′≥1) candidate visual media (that is, meeting the preset requirements) are used as multiple first visual media, which is equivalent to filtering out the (M - Z′) candidate visual media ranked behind.

[0269] In an optional implementation manner, any one of the following three methods can be used to determine the comprehensive matching degree of the candidate visual media:

[0270] Method 1: Sum the matching degrees of the visual content of the candidate visual media in multiple dimensions to be matched to obtain the comprehensive matching degree of the candidate visual media.

[0271] Method 2: Perform a weighted sum of the matching degrees of the visual content of the candidate visual media in multiple dimensions to be matched to obtain the comprehensive matching degree of the candidate visual media.

[0272] Among them, the weights of the respective multiple dimensions to be matched can be pre-configured by the user.

[0273] Method 3: Using a machine learning model, determine the comprehensive matching degree of the candidate visual media according to the matching degrees of the visual content of the candidate visual media on multiple dimensions to be matched.

[0274] Among them, the machine learning model needs to be trained based on a data set, and the purpose of its training is essentially to learn the weights of each dimension to be matched.

[0275] In the above Method 1, the contribution degrees of the matching degrees on different dimensions to the comprehensive matching degree are not distinguished, which may lead to the numerical values of the comprehensive matching degrees of multiple candidate visual media calculated finally being relatively close or equal, and further lead to the inability to screen multiple candidate visual media.

[0276] Exemplarily, assume that there are Picture 1, Picture 2, and Picture 3 among multiple candidate visual media, and there are Dimension A, Dimension B, and Dimension C among multiple dimensions to be matched. Calculate the matching degrees of the visual content of each picture on each dimension respectively, and the results are shown in Table 1.

[0277] Table 1:

[0278] Dimension A Dimension B Dimension C Picture 1 0.30 0.35 0.55 Picture 2 0.40 0.38 0.42 Picture 3 0.38 0.37 0.45

[0279] If calculated according to Method 1, the comprehensive matching scores of Picture 1, Picture 2, and Picture 3 are all 1.2, which will lead to the inability to screen the three pictures.

[0280] In the above Method 2, when there are too many semantic entities related to visual content to be concerned about (that is, the number of preset multiple tags is too large), it is difficult to accurately configure the weights of different dimensions.

[0281] In the above Method 3, the construction of the data set is inseparable from user data. However, user data belongs to user privacy content, and users do not want their data to be reported to the cloud side.

[0282] In an optional implementation manner, to solve the above problems, the following steps can be adopted to determine the comprehensive matching degree:

[0283] S51: Determine the weights of the respective N dimensions to be matched.

[0284] Among them, the weight of the j-th dimension to be matched is positively correlated with the degree of variation of the matching degree of the visual content of the M candidate visual media files on the j-th dimension to be matched; j is an integer, and the value of j ranges from 1 to N in sequence;

[0285] S52. According to the weights of the N dimensions to be matched, perform a weighted sum of the degrees of match of the visual content of the i-th candidate visual media file on each dimension to be matched among the N dimensions to be matched, so as to obtain the comprehensive degree of match of the i-th candidate visual media file.

[0286] Among them, i is an integer, and the values of i sequentially range from 1 to M.

[0287] Taking the search statement "Last year, the child held a cake on his birthday" as an example, the semantic entities related to the visual content include: child, birthday, and cake. If multiple recalled photos all contain a cake, then the degree of variation in the degrees of match of the visual content of the multiple recalled photos on the dimension of "cake" will be relatively small, and the weight corresponding to the dimension of "cake" will be relatively small; if some of the multiple recalled photos contain a child and some do not, then the degree of variation in the degrees of match of the visual content of the multiple recalled photos on the dimension of "child" will be relatively large, and the weight corresponding to the dimension of "child" will be relatively large.

[0288] In one example, for each dimension to be matched, determine the degree of variation in the degrees of match of the visual content of multiple candidate visual media on this dimension to be matched according to the information entropy of the degrees of match of the visual content of multiple candidate visual media on this dimension to be matched; among them, the degree of variation is inversely proportional to the information entropy. The calculation method of the information entropy will be introduced in detail in the following embodiments.

[0289] In this embodiment, the information entropy is used to measure the degree of variation in the degrees of match under each dimension to be matched, and based on this, the weight corresponding to this dimension to be matched is determined.

[0290] In order to ensure that the degrees of match of the visual content of candidate visual media on different dimensions to be matched have a unified dimension, a normalization processing step can be performed. Specifically, for each candidate visual media, determine the initial degree of match of the candidate visual media on the dimension to be matched according to the vector similarity between the visual semantic vector of the candidate visual media and the semantic vector to be matched of the dimension to be matched. Exemplarily, the vector similarity between the visual semantic vector of the candidate visual media and the semantic vector to be matched of the dimension to be matched can be used as the initial degree of match of the candidate visual media on the dimension to be matched. Perform normalization processing on the initial degrees of match of the visual content of multiple candidate visual media on the dimension to be matched to obtain the degrees of match of the visual content of multiple candidate visual media on this dimension to be matched.

[0291] The normalization process and the calculation processes of the weights and the comprehensive degree of match will be introduced in detail below:

[0292] Suppose there are m candidate visual media and n dimensions to be matched. The initial matching degrees of the visual content of the m candidate visual media on each dimension to be matched among the n dimensions to be matched can be regarded as a data matrix:

[0293] X = (x ij ) m*n , i = 1, 2, …, m; j = 1, 2, …, n (5)

[0294] where x ij is the initial matching degree of the visual content of the i-th candidate visual media on the j-th dimension to be matched.

[0295] Continuing with the above example, there are multiple candidate visual media including Picture 1, Picture 2, and Picture 3, and multiple dimensions to be matched including Dimension A, Dimension B, and Dimension C. That is, m is 3 and n is 3.

[0296] Step 1: Perform normalization processing on the above data matrix.

[0297] The normalized matrix is:

[0298] R = (r ij ) m*n , i = 1, 2, …, m; j = 1, 2, …, n (6)

[0299] where,

[0300]

[0301] where, max(x j ) refers to the maximum value among the initial matching degrees of the visual content of multiple candidate visual media on the j-th dimension to be matched; min(x j ) refers to the minimum value among the initial matching degrees of the visual content of multiple candidate visual media on the j-th dimension to be matched.

[0302] In practical applications, other normalization methods can also be used, and the embodiments of this application do not make specific limitations in this regard.

[0303] Exemplarily, the results obtained by normalizing the example data in Table 1 above are shown in Table 2.

[0304] Table 2:

[0305] Dimension A Dimension B Dimension C Picture 1 0.00 0.00 1.00 Picture 2 1.00 1.00 0.00 Picture 3 0.80 0.67 0.23

[0306] Step 2: Calculate the information entropy corresponding to each dimension to be matched.

[0307] Formulas (8) and (9) can be used for calculation:

[0308]

[0309] Among them,

[0310]

[0311] Among them, e j refers to the information entropy corresponding to the j-th dimension to be matched.

[0312] Among them, since the domain of the ln(x) function is x > 0. In actual calculations, to avoid the case where p in ln(p ij ) takes 0, ln(p ij ) in formula (8) can be replaced by ln(p ij ) + α), where α << 0.001. ij + α), where α << 0.001.

[0313] The smaller the information entropy corresponding to the dimension to be matched, the greater the degree of variation in the matching degree of the visual content of multiple candidate visual media on this dimension to be matched, and the greater the amount of information provided. It can be considered that the role played by this dimension to be matched in the comprehensive evaluation is also greater.

[0314] Exemplarily, for the example data in Table 2 above, the information entropy calculated according to the above formula (8) and formula (9) is shown in Table 3:

[0315] Table 3:

[0316] Dimension A Dimension B Dimension C Information Entropy 0.69 0.67 0.48

[0317] Step 3: Calculate the weights corresponding to each dimension to be matched.

[0318] The formula (10) can be used to calculate the weights corresponding to each dimension to be matched:

[0319]

[0320] Among them, d j refers to the weight corresponding to the j-th dimension to be matched.

[0321] It can be seen that the above formula (10) is a monotonically decreasing function of the information entropy. In this embodiment, the smaller the information entropy, the greater the degree of variation; the greater the degree of variation, the greater the weight. That is to say, the weight is negatively correlated with the information entropy.

[0322] To ensure that the sum of the weights corresponding to multiple dimensions to be matched is 1, the formula (11) can be used for the following calculation as the final weight w j :

[0323]

[0324] Among them, w jRefers to the final weight corresponding to the j-th dimension to be matched.

[0325] In an alternative embodiment, other monotonically decreasing functions may also be used to calculate the above weights, and the embodiments of the present application do not make specific limitations thereto.

[0326] Exemplarily, for the example data in Table 3 above, the weights of different dimensions to be matched can be calculated according to the above formulas (10) and (11), as shown in Table 4:

[0327] Table 4:

[0328] Dimension A Dimension B Dimension C Weight 0.28 0.29 0.42

[0329] It should be noted that since rounding is introduced in the process of calculating the information entropy, the sum of the three dimensions in Table 4 above is not 1.

[0330] Step 4: Weighted summation.

[0331] Perform a weighted summation on the normalized matching degrees of any candidate visual media to obtain the comprehensive matching degree of the candidate visual media, which can be specifically calculated using formula (12):

[0332]

[0333] where s i refers to the comprehensive matching degree of the i-th visual media.

[0334] Exemplarily, for the example data in Tables 2 and 4 above, the comprehensive matching degrees of different pictures are calculated using the above formula (12), as shown in Table 5:

[0335] Table 5:

[0336] Picture 1 Picture 2 Picture 3 Comprehensive Matching Degree 0.42 0.58 0.52

[0337] In the above 5102, when the first branch recalls visual media, it only considers the visual semantic information of the visual media and does not consider the attribute information such as the time and location of the visual media. Therefore, it is necessary to perform time filtering, location filtering, etc. on the visual media recalled by the first branch. Specifically, when the user performs semantic search, if the search statement contains time information, it is necessary to filter out the pictures that do not meet the time limit.

[0338] However, the user's description form of time information is rich and diverse and is fuzzy. For example, when the user searches for "the sky photographed in the afternoon", the "afternoon" in the user's search statement has no clear and standardized definition. When the user uses a fuzzy time expression in the search statement, the time window used for time filtering will affect the user experience.

[0339] Exemplary: The user starts traveling to other places on September 30, 2022 and arrives home on October 9, 2022. During this period, the user takes a lot of photos. One day in 2023, the user wants to view the beautiful scenery taken during this trip. Then the user is very likely to enter the search statement "Scenery taken during last year's National Day holiday". If directly according to "last year's National Day" in the search statement, the time window used for time filtering is set to "from October 1, 2022 to October 7, 2022", then the scenic photos taken by the user on September 30, 2022, October 8, 2022, and October 9, 2022 will be filtered out, which obviously does not meet the user's expectations. To improve the rationality of time filtering, the embodiments of the present application provide a new time filtering method. Specifically, using a clustering algorithm, cluster multiple candidate visual media according to the acquisition time of each of the multiple candidate visual media to obtain K (K≥1) clustering clusters.

[0340] In this way, pictures of the same series with relatively close acquisition times can be grouped into the same clustering cluster.

[0341] Continuing with the above example, the natural scenery picture taken by the user on September 30, 2022 and the natural scenery picture taken by the user on October 1, 2022 have relatively close shooting times. Through the above clustering algorithm, they are grouped into the same clustering cluster.

[0342] The above clustering algorithm may include but is not limited to: K-Means clustering algorithm, Mean shift clustering algorithm, and density-based clustering algorithm.

[0343] Taking the density-based clustering algorithm as an example, in the scenario of semantic search including time information, it is unreasonable to set fixed first parameter ∈ and second parameter MinPts. For example, when the user searches for "Photos taken during the outing in 2022", the time range is 1 whole year; while when the user searches for "Photos taken during the morning outing", the time range is several hours. In these two cases, the same first parameter ∈ and second parameter MinPts should not be set. To improve the rationality of clustering, the following steps can be used to determine the first parameter ∈ and the second parameter MinPts:

[0344] 51021. Determine the first parameter involved in the density-based clustering algorithm according to the time search range included in the search statement.

[0345] Among them, the first parameter is positively correlated with the duration corresponding to the time search range.

[0346] The time search range can be determined according to the time-related semantic entities in the search statement. Specifically, the time range corresponding to the time-related semantic entity can be returned through a mapping table (i.e., the time search range). Among them, the mapping table can be constructed in advance as needed, and the specific form is not specifically limited in the embodiments of the present application.

[0347] Exemplarily, the time-related semantic entity extracted from the search statement "photos taken in spring" is "spring". Querying the mapping table, the corresponding time range is obtained as: February 1 - May 30; the time-related semantic entity extracted from the search statement "the sky taken in the morning" is "morning". Querying the mapping table, the corresponding event range is obtained as: 7:00 - 12:00; the time-related semantic entity extracted from the search statement "photos of going out for fun during the National Day" is "National Day". Querying the mapping table, the corresponding time range is obtained as: October 1 - October 7; the time-related entity extracted from the search statement "the sky taken in Beijing this year" is "this year". Querying the mapping table, the corresponding time range is obtained as "January 1, 2023 to December 31, 2023".

[0348] The first parameter can be determined according to the duration corresponding to the time search range. Exemplarily, the start time of the time search range (which can be understood as the start timestamp) is T start and the end time (which can be understood as the end timestamp) is T end , and the duration of the time search range is: T end -T start , and the following formula can be used to calculate the first parameter:

[0349] ∈=α*(T end -T start ) (13)

[0350] Among them, α is a coefficient that can be adjusted manually, and its value can be set according to actual needs, which is not specifically limited in the present application.

[0351] 51022. Determine the second parameter involved in the density-based clustering algorithm according to the ratio of the number of multiple candidate visual media to the time search range.

[0352] The following formula can be used to calculate the second parameter:

[0353]

[0354] Among them, N is the total number of visual media whose acquisition time is within the time search range in the candidate set, N≥1; among them, n is the number of times the time search range repeats between T 1 and T 2 . T 1is the acquisition time of the earliest acquired visual media among multiple visual media stored via the mobile phone; T 2 is the acquisition time of the latest acquired visual media among multiple visual media stored via the mobile phone. Exemplarily, the search statement is "the sky photographed during the National Day", and its time search range is from October 1st to October 7th, with a duration of 7 days; the acquisition time of the earliest acquired visual media stored in the mobile phone is August 1st, 2020; the acquisition time of the latest acquired visual media stored in the mobile phone is October 20th, 2023; then, from August 1st, 2020 to October 20th, 2023, the time search range from October 1st to October 7th repeats 4 times (i.e., once a year).

[0355] where, (T end - T start ) * n can be understood as the total duration corresponding to the time search range.

[0356] where, β is a coefficient that can be adjusted manually, and its magnitude can be set according to actual needs. This application does not make specific limitations on this.

[0357] In this embodiment, according to the duration defined by the time search range included in the search statement, the first parameter ∈ and the second parameter MinPts involved in the density-based clustering algorithm are dynamically adjusted, which can ensure the rationality of clustering and thus improve the accuracy of the final search result.

[0358] When the search statement includes a time search range, determine the acquisition time range corresponding to each of the K (K≥1) clustering clusters; filter out the clustering clusters whose acquisition time range does not overlap with the time search range, and retain the clustering clusters whose acquisition time range overlaps with the time search range.

[0359] The overlap between the acquisition time range and the time search range can be partial overlap or full overlap. Whether it is partial overlap or full overlap between the two, both belong to having an overlapping part.

[0360] In this way, from the K clustering clusters, G (G≥1) clustering clusters are selected. The acquisition time of the earliest acquired visual media among these G clustering clusters is before the start time of the time search range, and / or, the acquisition time of the latest acquired visual media among these G clustering clusters is after the end time of the time search range. Note: There is no intersection among the G clustering clusters, and there is also no intersection among the acquisition time ranges of the G clustering clusters themselves.

[0361] The acquisition time range corresponding to the clustering cluster is from the acquisition time T min T1 of the earliest acquired visual media in this clustering cluster to the acquisition time T max of the latest acquired visual media in this clustering cluster, that is: [Tmin , T max .

[0362] Specifically, the time search range is [T start , T end , and the acquisition time range corresponding to the clustering cluster is [T min , T max . When [T min , T max and [T start , T end have an overlapping part, keep this clustering cluster; when [T min , T max and [T start , T end have no overlapping part, filter this clustering cluster. For example: the time search range is from October 1, 2022 to October 7, 2022, and the acquisition time range of the clustering cluster is from September 30, 2022 to October 1, 2022. These two ranges have an overlapping part (i.e., October 1, 2022), and this clustering cluster is retained.

[0363] Continuing with the above example, the natural scenery pictures taken by the user on September 30, 2022 and the natural scenery pictures taken by the user on October 1, 2022 are close in shooting time. Through the above clustering algorithm, they are grouped into the same clustering cluster. Assuming this clustering cluster contains only these two pictures, then the acquisition time range of this clustering cluster is: from September 30, 2022 to October 1, 2022, and it has an overlapping part with the time search range: from October 1, 2022 to October 7, 2022. Therefore, this clustering cluster will be retained, that is to say, the natural scenery pictures taken by the user on September 30, 2022 will not be filtered out.

[0364] Similarly, the natural scenery pictures taken by the user on October 8, 2022 and the natural scenery pictures taken by the user on October 7, 2022 are close in shooting time. Through the above clustering algorithm, they are grouped into the same clustering cluster. Assuming this clustering cluster contains only these two pictures, then the acquisition time range of this clustering cluster is the pictures from October 7, 2022 to October 8, 2022, and it has an overlapping part with the time search range: from October 1, 2022 to October 7, 2022. Therefore, this clustering cluster will be retained. That is to say, the natural scenery pictures taken by the user on October 8, 2022 will not be filtered out.

[0365] In this way, for the search statement "scenery pictures taken during last year's National Day holiday", the acquisition time of the earliest acquired visual media among the multiple clustering clusters obtained by screening is September 30, 2022, and the acquisition time of the latest acquired visual media is October 8, 2022.

[0366] It can be seen that the time filtering method provided by the embodiments of the present application can ensure that a series of photos with relatively close acquisition times are displayed to the user, guarantee the coherence of the search results, and improve the user's search experience.

[0367] In addition, when the search statement further includes a semantic entity related to a location, location filtering can be continued for the G clustering clusters filtered out. Specifically, visual media in the clustering clusters whose acquisition locations do not match the semantic entity related to the location in the search statement can be filtered out. Exemplarily, the acquisition location "Prince Kung's Mansion, Xicheng District, Beijing" and the semantic entity "Sujiatuo Town, Haidian District, Beijing" can be considered as matching (belonging to municipal-level matching); the acquisition location "Tsinghua University, Haidian District, Beijing" and the semantic entity "Sujiatuo Town, Haidian District, Beijing" can be considered as matching (belonging to district-level matching); the acquisition location "Shanghai" and the semantic entity "Beijing" can be considered as not matching.

[0368] In the above embodiments, clustering is performed first, then time filtering, and finally location filtering. Of course, in practical applications, location filtering can also be performed first, then clustering, and finally time filtering. The specific execution order of these three steps can be set according to actual needs, and the embodiments of the present application do not make specific limitations in this regard.

[0369] 5103. Filtering of relationships between people.

[0370] When the search statement includes a semantic entity related to the relationship between people, visual media in each clustering cluster whose relationship attribute between people does not match the semantic entity are filtered out. Exemplarily, if the search statement includes "good friends" and the person name attribute of Picture B is "colleague", and these two do not match, then Picture B is filtered out.

[0371] 5104. Filtering of person names.

[0372] When the search statement includes a semantic entity related to a person name, visual media in each clustering cluster whose person name attribute does not match the semantic entity are filtered out. Exemplarily, if the search statement includes "Zhang San" and the person name attribute of Picture A is "Li Si", and these two do not match, then Picture A is filtered out.

[0373] It should be added that the semantic entity related to the person name and the semantic entity related to the relationship between people in the search statement can also be recognized through named entity recognition technology. The person name attribute and the relationship attribute between people of the visual media are manually added by the user for the visual media in advance.

[0374] 511. Sorting.

[0375] For the candidate set obtained in the above step 508, the visual media in the candidate set can be sorted according to the acquisition time of the visual media in the candidate set, so as to obtain the display order of the visual media. The acquisition time of the visual media with a higher display order is earlier than that of the visual media with a lower display order. Subsequently, the mobile phone can display the candidate set according to the display order of the visual media in the candidate set.

[0376] For the filtered candidate set obtained in the above step 510, the following method can be used for sorting:

[0377] The above filtered candidate set includes: the above G clustering clusters; the G clustering clusters are sorted according to the start time of the acquisition time range of the G clustering clusters, so as to obtain the display order of the G clustering clusters (that is, the inter-cluster sorting). Exemplarily, as Figure 6 shown in the interface 601, the acquisition time range corresponding to the clustering cluster A is from September 30, 2022 to October 1, 2022; the acquisition time range corresponding to the clustering cluster B is from October 3, 2022 to October 5, 2022; then, the display order of the clustering cluster A is prior to the display order of the clustering cluster B. For each clustering cluster, the visual media within the cluster are sorted according to the target matching degree of the visual media within the cluster from high to low, so as to obtain the display order (intra-cluster sorting) among the visual media within the cluster; or, for each clustering cluster, the visual media within the cluster are sorted according to the acquisition time of the visual media within the cluster, so as to obtain the display order of the visual media within the cluster. Subsequently, the mobile phone displays the visual media in the G clustering clusters according to the inter-cluster sorting and the intra-cluster sorting.

[0378] In addition, in the terminal device, taking pictures as an example, when the user searches for pictures, not only the visual information of the pictures is concerned, but also the attribute information of the pictures (for example: the shooting location, shooting time, etc.) of the pictures is concerned. Therefore, how to perform fusion sorting on the recalled pictures is particularly important. Assume that the candidate set obtained in the above step 508 or the filtered candidate set obtained in the above step 510 includes: Y visual media; Y is an integer greater than 1. The following steps can be used to sort the Y visual media, so as to obtain the final search result. As Figure 7 shown, it includes:

[0379] 5111. Determine W dimensions to be matched according to the search statement.

[0380] Among them, W is an integer greater than 1.

[0381] 5112. Obtain the matching degree and the matching degree ranking of the Y visual media files on the w-th dimension to be matched.

[0382] Among them, Y is an integer greater than 1; w is an integer and ranges from 1 to W.

[0383] 5113. Correct the total number of times the y-th visual media file appears in the u-th place according to the matching degrees of the y-th visual media file on each of the W dimensions to be matched, to obtain the corrected total number of times the y-th visual media file appears in the u-th place.

[0384] Where u is an integer ranging from 1 to Y; y is an integer ranging from 1 to Y.

[0385] 5114. Determine the score of the y-th visual media file according to the corrected total number of times the y-th visual media file appears in each place.

[0386] 5115. Sort the Y visual media files according to the scores of the Y visual media files.

[0387] In the above 5111, the determination methods of the dimensions to be matched include multiple methods as follows:

[0388] Determine the search statement as the dimension to be matched;

[0389] Determine the semantic entity irrelevant to the visual content in the search statement as the dimension to be matched;

[0390] Determine the number of user clicks as the dimension to be matched.

[0391] In an example, the W dimensions to be matched may include: the search statement itself, the semantic entity related to time in the search statement, the semantic entity related to location in the search statement, and the number of user clicks. Among them, the number of user clicks reflects the personalized preference of the user. The more the number of user clicks on a certain visual media, the more the user prefers the visual media.

[0392] In the above 5112, the W dimensions to be matched include: the number of user clicks; determine the matching degree of the y-th visual media file in terms of the number of user clicks according to the number of times the y-th visual media file has been clicked by the user historically.

[0393] The W dimensions to be matched include: the search statement; the matching degree of the y-th visual media file on the search statement refers to the matching degree between the visual semantic vector of the y-th visual media file and the first sentence semantic vector of the search statement.

[0394] Among the W dimensions to be matched, there are semantic entities in the search statement that are irrelevant to visual content. According to the matching situation of the semantic entities irrelevant to visual content based on the attributes of the y-th visual media file, the matching degree of the y-th visual media file on the semantic entities irrelevant to visual content in the search statement is determined. Among them, the semantic entities irrelevant to visual content include: semantic entities related to time and / or semantic entities related to location. Specifically, according to the matching situation of the attributes of the y-th visual media file with the semantic entities related to time, the matching degree of the y-th visual media file on the semantic entities related to time in the search statement is determined. According to the matching situation of the attributes of the y-th visual media file with the semantic entities related to location, the matching degree of the y-th visual media file on the semantic entities related to location in the search statement is determined.

[0395] The matching degree ranking of the Y visual media files in the w-th dimension to be matched is obtained by sorting the matching degrees of the Y visual media files in the w-th dimension to be matched from large to small.

[0396] In the above 5113, the total number of times the y-th visual media file appears in the u-th place is corrected by using the matching degrees of the y-th visual media file in each dimension to be matched among the W dimensions to be matched. That is to say, the corrected total number of times not only considers the true total number of times, but also considers the size of the matching degree, which helps to optimize the subsequent sorting effect.

[0397] Exemplarily, the matching degree of the first visual media file in the first dimension to be matched is 1.0 (the optimal matching degree is 1.0), and it is ranked first; the matching degree of the first visual media file in the second dimension to be matched is 0.5, and it is ranked first. The matching degree of the first visual media file in the first dimension to be matched is very large. Then, the value or contribution of the first place obtained by the first visual media file in the first dimension to be matched should be relatively large; the matching degree of the first visual media file in the second dimension to be matched is very small. Then, the value or contribution of the first place obtained by the first visual media file in the second dimension to be matched should be relatively small. Therefore, by correcting the total number of times according to the size of the matching degree, the subsequent sorting effect can be optimized.

[0398] In an implementable solution, in the above 5113, "correct the total number of times the y-th visual media file appears in the u-th place according to the matching degrees of the y-th visual media file in each dimension to be matched among the W dimensions to be matched, and obtain the corrected total number of times the y-th visual media file appears in the u-th place" may include:

[0399] S61. Determine the first weight of the w-th dimension to be matched for the y-th visual media file according to the matching degree of the y-th visual media file in the w-th dimension to be matched.

[0400] S62. According to the first weight of each of the W dimensions to be matched for the y-th visual media file, perform a weighted sum of the number of times the y-th visual media file appears in the u-th place in each dimension to be matched among the W dimensions to be matched, so as to obtain the corrected total number of times the y-th visual media file appears in the u-th place.

[0401] In the above S61, in one example, the matching degree of the y-th visual media file in the w-th dimension to be matched can be determined as the first weight of the w-th dimension to be matched for the y-th visual media file.

[0402] In another example, obtain the preset weight of the w-th dimension to be matched; according to the preset weight of the w-th dimension to be matched, correct the matching degree of the y-th visual media file in the w-th dimension to be matched to obtain the corrected matching degree of the y-th visual media file in the w-th dimension to be matched; determine the corrected matching degree of the y-th visual media file in the w-th dimension to be matched as the first weight of the w-th dimension to be matched for the y-th visual media file.

[0403] The preset weights of the W dimensions to be matched can be set based on prior knowledge, and the embodiments of the present application do not make specific limitations on this.

[0404] The product of the matching degree of the y-th visual media file in the w-th dimension to be matched and the preset weight of the w-th dimension to be matched can be determined as the corrected matching degree of the y-th visual media file in the w-th dimension to be matched.

[0405] It should be noted that the first weights of the w-th dimension to be matched for different visual media files may be different.

[0406] In the above S62, the number of times the y-th visual media file appears in the u-th place in the w-th dimension to be matched is 1 or 0.

[0407] In addition, before performing the above step 5113, the matching degrees of the Y visual media files in the w-th dimension to be matched can also be normalized.

[0408] In the above 5114, according to the corrected total number of times the y-th visual media file appears in each ranking and the ranking scores of each ranking, determine the score of the y-th visual media.

[0409] In the above 5115, sort the Y visual media files in descending order according to the scores of the Y visual media files to obtain the display order of the Y visual media files. Subsequently, display the visual media files according to the display order.

[0410] The following will introduce the calculation process of the score of visual media files through examples:

[0411] Suppose there are m dimensions and n pictures in total. In the following text, is the matching score of the i-th (1 ≤ i ≤ n) picture in the j-th (1 ≤ j ≤ m) dimension.

[0412] Example: Suppose the user enters a search statement "The sky photographed in Sujiatuo Town, Haidian District, Beijing this year", and a total of 4 pictures are recalled; and the search statement "The sky photographed in Sujiatuo Town, Haidian District, Beijing this year" corresponds to four dimensions to be matched: search statement, this year, Sujiatuo Town, Haidian District, Beijing, and click count. The matching situation of each picture in each dimension is counted respectively, and the results are shown in Table 6.

[0413] Table 6:

[0414] Search Statement This Year Sujiatuo Town, Haidian District, Beijing Click Count Picture 1 0.30 Match Information Missing 50 Picture 2 0.40 No Match No Match 0 Picture 3 0.38 No Match Province / City Match 35 Picture 4 0.50 Match Street Match 50

[0415] The data in Table 6 can be numericalized to obtain the matching degrees of each picture in each dimension, as shown in Table 7.

[0416] Table 7:

[0417] Search Statement This Year Sujiatuo Town, Haidian District, Beijing Click Count Picture 1 0.30 1 0.2 50 Picture 2 0.40 0 0 0 Picture 3 0.38 0 0.8 35 Picture 4 0.50 1 1 50

[0418] 1. Ranking score conversion

[0419] Calculate the "ranking" score vector R = (r 1 , r 2 , …, r h ) T for the scoring of h (1 ≤ h ≤ n), where r h = n - h + 1. Here, h represents the ranking.

[0420] Example: For the case of 4 pictures, the ranking can be converted as follows:

[0421] Table 8:

[0422] First Place Second Place Third Place Fourth Place Scoring 4 3 2 1

[0423] 2. Percentage processing of matching degree

[0424] Calculate the matching degree after percentage processing of the i-th picture in the j-th dimension

[0425]

[0426] where is the matching degree of the i-th picture in the j-th dimension; represents the optimal matching score in the j-th dimension,

[0427] Exemplarily, if the optimal matching scores in the three dimensions of search statement, time, and location are all 1.0; the optimal matching score of click-through rate: 100. The percentage-based picture information calculated according to the above formula is as follows:

[0428] Table 9:

[0429] Search Statement This Year Sujiatuo Town, Haidian District, Beijing Click Count Picture 1 0.30 1 0.2 0.50 Picture 2 0.40 0 0 0 Picture 3 0.38 0 0.8 0.35 Picture 4 0.50 1 1 0.50

[0430] 3. Preset weight configuration for each dimension

[0431] Based on prior knowledge, etc., configure the preset weight for each dimension: W = (w (1) , w (2) , …, w (m) ) T .

[0432] Example, based on expert experience, it is considered that the importance of time matching degree is twice that of other dimensions, then W = (1, 2, 1, 1) T .

[0433] 4. Fusion scoring:

[0434] Calculate the fusion score for picture i:

[0435] s i = R T f i = R T δ i U i W (16)

[0436] Among them, R T represents the transpose operation on R; represents the ranking matrix of the picture in different dimensions. Specifically, if the ranking of the i-th picture in the j-th dimension is the h-th, then

[0437] For example, for picture 1, the positions in the respective rankings of the 4 dimensions are (4, 1, 3, 1), then there is:

[0438]

[0439] Optionally, the scoring of the picture can be controlled within a certain range, for example, [r mim , r max , r max is the value of the largest element in R, r min is the value of the smallest element in R. f i = (f i1 , fi2 ,…, f in ) T After normalization, f' i =(f' i1 , f' i2 ,…, f' in ) T Then calculate the fusion score, that is

[0440] s i = R T f' i (17)

[0441]

[0442] where f ij is the corrected total number of times the i-th picture appears in the l-th place.

[0443] That is to say, this solution improves the Borda Count method to incorporate the matching degree and the preset weights corresponding to each dimension

[0444] When calculating the score of each visual medium in this solution

[0445] It should be noted that, in order to better protect user privacy and security and meet the principle of minimizing user data, the entire search process is completed on the device side as much as possible to avoid reporting user data to the cloud side.

[0446] In addition, this application provides an electronic device, including: a memory, a processor, and a display. Among them, the memory is used to store programs; the processor is coupled to the memory and the display and is used to execute the programs stored in the memory to implement the above-mentioned visual medium search method.

[0447] This application embodiment also provides a computer-readable storage medium storing a computer program, where the computer program, when executed by a computer, can implement one or more steps in any of the above-mentioned visual medium search methods.

[0448] The computer-readable storage medium can be a non-temporary computer-readable storage medium. For example, the non-temporary computer-readable storage medium can be ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage devices, etc.

[0449] Another embodiment of this application also provides a computer program product containing instructions. When the computer program product is executed by a computer, it can implement one or more steps in any of the above-mentioned methods.

[0450] Among them, the electronic device, computer-readable storage medium, and computer program product provided in this embodiment are all used to execute the corresponding methods provided above. Therefore, the beneficial effects they can achieve can refer to the beneficial effects in the corresponding methods provided above, and will not be elaborated here.

[0451] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.

[0452] The units described as separate components may or may not be physically separated. The components displayed as units may be one physical unit or multiple physical units, that is, they can be located in one place, or they can be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0453] In addition, in each embodiment of this application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0454] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiments of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods in the embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0455] The above content is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A visual media search method applicable to an electronic device, characterized in that, it includes: displaying a first interface; the first interface includes a search box; receiving a search operation on the search statement input into the search box; determining W dimensions to be matched according to the search statement; W is an integer greater than 1; obtaining the matching degree and the matching degree ranking of Y visual media files in the wth dimension to be matched; Y is an integer greater than 1; w is an integer and ranges from 1 to W; correcting the total number of times the yth visual media file appears in the uth place according to the matching degrees of the yth visual media file in each dimension to be matched among the W dimensions to be matched, to obtain the corrected total number of times the yth visual media file appears in the uth place; wherein, u is an integer and ranges from 1 to Y; y is an integer and ranges from 1 to Y; determining the score of the yth visual media file according to the corrected total number of times the yth visual media file appears in each ranking; determining the search result of the search statement according to the scores of the Y visual media files; displaying the search result.

2. The method according to claim 1, characterized in that, correcting the total number of times the yth visual media file appears in the uth place according to the matching degrees of the yth visual media file in each dimension to be matched among the W dimensions to be matched, to obtain the corrected total number of times the yth visual media file appears in the uth place, includes: determining a first weight of the wth dimension to be matched for the yth visual media file according to the matching degree of the yth visual media file in the wth dimension to be matched; weighted summing the number of times the yth visual media file appears in the uth place in each dimension to be matched among the W dimensions to be matched according to the first weights of the W dimensions to be matched for the yth visual media file respectively, to obtain the corrected total number of times the yth visual media file appears in the uth place.

3. The method according to claim 2, characterized in that, determining a first weight of the wth dimension to be matched for the yth visual media file according to the matching degree of the yth visual media file in the wth dimension to be matched, includes: obtaining a preset weight of the wth dimension to be matched; correcting the matching degree of the yth visual media file in the wth dimension to be matched according to the preset weight of the wth dimension to be matched, to obtain the corrected matching degree of the yth visual media file in the wth dimension to be matched; determining the corrected matching degree of the yth visual media file in the wth dimension to be matched as the first weight of the wth dimension to be matched for the yth visual media file.

4. The method according to any one of claims 1 to 3, characterized in that, the determination methods of the dimensions to be matched include multiple of the following methods: determining the search statement as the dimension to be matched; determining the semantic entity irrelevant to the visual content in the search statement as the dimension to be matched; determining the number of user clicks as the dimension to be matched.

5. The method according to any one of claims 1 to 3, It is characterized in that among the W dimensions to be matched, it includes: the number of user clicks; According to the number of times the y-th visual media file has been clicked by the user historically, determine the matching degree of the y-th visual media file in terms of the number of user clicks.

6. The method according to any one of claims 1 to 3, It is characterized in that among the W dimensions to be matched, it includes: the search statement; According to the vector similarity between the visual semantic vector of the y-th visual media file and the first sentence semantic vector of the search statement, determine the matching degree of the y-th visual media file in the search statement.

7. The method according to any one of claims 1 to 3, It is characterized in that among the W dimensions to be matched, it includes: the semantic entity irrelevant to visual content in the search statement; According to the matching situation between the attributes of the y-th visual media file and the semantic entity irrelevant to visual content, determine the matching degree of the y-th visual media file in the semantic entity irrelevant to visual content in the search statement.

8. The method according to any one of claims 1 to 3, It is characterized in that before correcting the total number of times the y-th visual media file appears in the u-th place according to the matching degrees of the y-th visual media file in each dimension to be matched among the W dimensions to be matched, the method further includes: Perform percentage processing on the matching degrees of the Y visual media files in the w-th dimension to be matched.

9. The method according to any one of claims 1 to 3, It is characterized in that further includes: Perform semantic understanding on the search statement to obtain a first sentence semantic vector; Obtain the visual semantic vectors of multiple visual media files; The visual semantic vector of each visual media file is obtained by performing semantic understanding on the image or image frame of the visual media file using a natural image understanding model; Determine multiple first visual media files whose visual semantic vectors match the first sentence semantic vector from the multiple visual media files; Determine the Y visual media files according to the multiple first visual media files.

10. An electronic device, It is characterized in that includes: a memory, a processor, and a display, wherein the memory is used for storing programs; the display is used for displaying a search page; the processor is coupled to the memory and the display, and is used for executing the program stored in the memory to implement the method according to any one of claims 1 to 9.

11. A computer-readable storage medium storing a computer program, It is characterized in that when the computer program is executed by a computer, it can implement the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Method and device for displaying video retrieval results

    CN104462573A

  • Visual media personalized search method and device

    CN113641857A

  • Image retrieval method and related equipment

    CN116431855A