Visual media searching method and device and storage medium
By receiving search statements containing time search range and keywords in the gallery application of electronic devices, recalling photos and videos whose acquisition time is outside the time search range but matches the keywords, the problem of difficult to recall media files that match keywords but do not match the time in the prior art is solved, and the user's search experience is improved.
Patent Information
- Application Number
- CN202311562033.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-21
- Publication Date
- 2025-05-30
AI Technical Summary
It is difficult for the prior art to recall photos and videos whose acquisition time is outside the time search range but match the keyword in search statements that contain both the time search range and the keyword.
By displaying the first interface, a search statement containing the time search range and keywords is received, a visual media file matching the keywords is displayed, and photos and videos that are collected from outside the time search range but match the keywords are recalled.
Ensure that the same series of photos or videos taken by the user before and after the specified time range can be recalled, improving the user's search experience.
Smart Images

Figure CN120067368A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of terminals, and in particular, to a visual media search method, device, and storage medium. Background Art
[0002] With the popularization of intelligent terminals, more and more users use intelligent terminals, such as mobile phones, to take photos and videos, and store the taken photos and videos in the photo gallery of the electronic device, thereby recording every bit of life. In addition, users can also download pictures, take screenshots of the mobile phone interface, and store the downloaded pictures and screenshots in the photo gallery of the electronic device.
[0003] To facilitate users to manage and view pictures in the terminal, a picture management function and a picture search function are configured in the photo gallery application or other similar applications of the terminal device. For example: The photo gallery application in the terminal can classify the pictures in the terminal according to information such as the shooting time and location of the pictures to generate corresponding photo albums, and users can view relevant pictures by searching for information such as time and location. Summary of the Invention
[0004] Multiple aspects of this application provide a visual media search method, device, and storage medium, which enable users to search for photos and videos using a search statement that includes both a time search range and a keyword, and can search for photos and videos that match the keyword but have a capture time outside the time search range, so as to ensure that photos or videos of the same series taken before and after the time search range by the user can be recalled, improving the user search experience.
[0005] In a first aspect, a visual media search method is provided, which is applicable to an electronic device and includes:
[0006] Display a first interface; the first interface includes a search box;
[0007] Receive a search operation on the search statement input into the search box; wherein, the search statement includes a time search range and a first keyword, the start time of the time search range is a first time point, and the end time is a second time point;
[0008] Display a first search result, the first search result corresponding to a first visual media file;
[0009] Wherein, the first visual media file matches the first keyword;
[0010] The capture time of the first visual media file is before the first time point or after the second time point.
[0011] It can be seen that in the technical solution provided by the present application, when a user uses a search statement that includes both a time search range and a keyword to search for pictures and videos, not all the photos and videos taken by the user before the start time point or after the end time point of the time search range will be filtered out. Instead, photos and videos that match the first keyword and were taken by the user before the start time point or after the end time point of the time search range will be recalled, improving the user's search experience.
[0012] Optionally, when the acquisition time of the first visual media file is before the first time point, the time difference between the acquisition time of the first visual media file and the first time point is less than a second threshold.
[0013] Optionally, when the acquisition time of the first visual media file is after the second time point, the time difference between the acquisition time of the first visual media file and the second time point is less than a third threshold.
[0014] That is to say, it is possible to recall photos and videos that match the first keyword and were taken by the user within a limited time period before the start time point or within a limited time period after the end time of the time search range.
[0015] The magnitudes of the above-mentioned second threshold and the third threshold may be equal or unequal, and the specific values can be set according to actual needs. The present application does not make specific limitations in this regard.
[0016] In a possible implementation manner, after receiving a search operation on the search statement input into the search box, it includes: displaying a second search result, where the second search result corresponds to a second visual media file;
[0017] The visual content of the second visual media file matches the first keyword;
[0018] The acquisition time of the second visual media file is between the first time point and the second time point, and the time difference between the acquisition time of the first visual media file and the acquisition time of the second visual media file is less than a first threshold.
[0019] That is to say, this solution can not only recall photos and videos that match the first keyword and were taken by the user between the start time point and the end time point of the time search range, but also recall photos and videos that match the first keyword and were taken by the user within a limited time period before the start time point or within a limited time period after the end time of the time search range.
[0020] In a possible implementation manner, displaying a first interface includes:
[0021] In response to an operation of opening the gallery application, the first interface is displayed;
[0022] Alternatively, in response to an operation of opening the negative first screen triggered on the home screen of the electronic device, display the first interface;
[0023] Alternatively, in response to a pull-down search operation triggered on the home screen of the electronic device, display the first interface.
[0024] That is, the user can perform searches for media such as photos and videos through the gallery application interface, the negative first screen interface, or the pull-down search interface.
[0025] In a possible implementation manner, the displaying the first search result includes:
[0026] Display a first picture, where the first picture is a proportional thumbnail of the first visual media file or a proportional thumbnail of any image frame.
[0027] When the first visual media file is a picture, the first picture is a proportional thumbnail of the picture; when the first visual media file is a video, the first picture is a proportional thumbnail of any image frame (such as the first frame) of the video.
[0028] It can be understood that the thumbnails of the search results are displayed through the grid page.
[0029] In a possible implementation manner, the method further includes: displaying a third picture; the third picture is a proportional thumbnail of a third visual media file or a proportional thumbnail of any image frame; the third visual media file matches the first keyword;
[0030] The first picture and the third picture are displayed on the second interface; the first picture is displayed in front of the third picture; the matching degree of the visual media file corresponding to the first picture with the first keyword is the first matching degree; the matching degree of the visual media file corresponding to the second picture with the first keyword is the second matching degree; the first matching degree is greater than the second matching degree.
[0031] It can be understood that the thumbnails of multiple visual media files obtained by the search are sorted and displayed on the grid page in descending order of the matching degree of their respective visual media files with the first keyword, facilitating the user to quickly find the visual media they want.
[0032] In a possible implementation manner, the visual content of the first visual media file matches the first keyword; the visual content is data that needs to be obtained through a natural picture understanding model.
[0033] In a possible implementation manner, the first keyword and the first visual media file are matched through a contrastive text-image pre-trained CLIP model.
[0034] Through model training, the CLIP model can map visual media and text to a unified vector space to fully understand the relationship between different modal data visually and textually, which can improve the matching accuracy and thus the recall accuracy.
[0035] In a second aspect, the present application provides a visual media search method applicable to an electronic device, including:
[0036] Displaying a first interface; the first interface includes a search box;
[0037] Receiving a search operation on a search statement input into the search box, the search statement including a time search range;
[0038] Using a clustering algorithm to cluster a plurality of first visual media files according to the acquisition time of the plurality of first visual media files to obtain K clustering clusters, the acquisition time ranges of the K clustering clusters are different; where K≥1 and is an integer;
[0039] Determining a target clustering cluster according to whether the acquisition time range overlaps partially with the time search range;
[0040] Determining a search result of the search statement according to the target clustering cluster;
[0041] Displaying the search result.
[0042] In this solution, visual media files with relatively close acquisition times are clustered into one cluster by a clustering algorithm, that is to say, the acquisition times of visual media files in each clustering cluster are relatively close.
[0043] The clustering cluster whose acquisition time range overlaps partially with the time search range means that: the visual media whose acquisition time in this clustering cluster is before the start time point of the time search range or after the end time point of the time search range is included. And, the acquisition time of the visual media whose acquisition time in this clustering cluster is before the start time point of the time search range or after the end time point of the time search range is relatively close to the acquisition times of other visual media in the cluster. It can be seen that through this solution, the same series of photos or videos continuously taken by the user before or after the start time point or the end time point of the time search range can be recalled to improve the user search experience.
[0044] In a possible implementation manner, the start time of the acquisition time range of each clustering cluster is the acquisition time of the earliest acquired visual media file in the clustering cluster, and the end time is the acquisition time of the latest acquired visual media file in the clustering cluster; there is no intersection between the acquisition time ranges of the K clustering clusters.
[0045] In a possible implementation, the start time of the time search range is the first time point, and the end time is the second time point;
[0046] Determine the target clustering clusters according to whether the acquisition time range overlaps with the time search range partially, including:
[0047] When the start time of the acquisition time range of the k-th clustering cluster among the K clustering clusters is between the first time point and the second time point and the end time is not between the first time point and the second time point, determine that the acquisition time range of the k-th clustering cluster overlaps with the time search range partially;
[0048] When the end time of the acquisition time range of the k-th clustering cluster is between the first time point and the second time point and the start time is not between the first time point and the second time point, determine that the acquisition time range of the k-th clustering cluster overlaps with the time search range partially;
[0049] When the start time of the acquisition time range of the k-th clustering cluster is not between the first time point and the second time point and the end time is not between the first time point and the second time point, determine that the acquisition time range of the k-th clustering cluster overlaps with the time search range partially; k is an integer, 1 ≤ k ≤ K;
[0050] Determine the clustering clusters among the K clustering clusters whose acquisition time ranges overlap with the time search range partially as the target clustering clusters.
[0051] Among them, the start time being between the first time point and the second time point means that the start time is greater than or equal to the first time point and less than or equal to the second time point; the start time not being between the first time point and the second time point means that the start time is less than the first time point or greater than the second time point; the end time being between the first time point and the second time point means that the end time is greater than or equal to the first time point and less than or equal to the second time point; the end time not being between the first time point and the second time point means that the end time is less than the first time point or greater than the second time point.
[0052] In a possible implementation, the search result includes multiple target clustering clusters; the method further includes:
[0053] Determine the display order of the multiple target clustering clusters according to the chronological order of the start times of the acquisition time ranges of the multiple target clustering clusters;
[0054] Display the search result, including:
[0055] Display the multiple target clustering clusters in accordance with the display order of the multiple target clustering clusters.
[0056] It is understandable that on the grid page, multiple target clustering clusters are displayed in sequence according to the start time of the acquisition time range of each target clustering cluster. The start time of the acquisition time range of the target clustering cluster displayed earlier on the grid page is earlier than that of the target clustering cluster displayed later.
[0057] In a possible implementation, the method further includes:
[0058] Performing semantic understanding on the search statement to obtain a first sentence semantic vector;
[0059] Obtaining visual semantic vectors of multiple visual media files; the visual semantic vector of each visual media file is obtained by performing semantic understanding on the image or image frame of the visual media file using a natural image understanding model;
[0060] Determining M candidate visual media files from the multiple visual media files, whose visual semantic vectors match the first sentence semantic vector; where M is an integer greater than 1;
[0061] Determining the multiple first visual media files according to the M candidate visual media files.
[0062] That is to say, the above-mentioned multiple first visual media files are recalled by matching the visual semantic vectors of the visual media files with the sentence semantic vectors of the search statement (hereinafter referred to as the first branch recall). In this way, it can be ensured that the visual content of the visual media in the finally obtained search results is semantically matched with the search statement.
[0063] Specifically, the CLIP model can be used to match visual media files with search statements.
[0064] In a possible implementation, using a clustering algorithm to cluster the multiple first visual media files according to the acquisition time of the multiple first visual media files to obtain K clustering clusters, including:
[0065] If the semantic proportion related to visual content in the search statement is greater than or equal to a preset proportion threshold, using a clustering algorithm to cluster the multiple first visual media files according to the acquisition time of the multiple first visual media files to obtain K clustering clusters; the visual content is data that needs to be obtained through a natural image understanding model;
[0066] The method further includes:
[0067] Obtaining the attributes of the multiple visual media files;
[0068] From the multiple visual media files, determine multiple second visual media whose attributes match the semantic entities in the search statement (hereinafter referred to as the second branch recall);
[0069] If the semantic proportion related to visual content in the search statement is less than the preset proportion threshold, determine the search result of the search statement according to the multiple second visual media.
[0070] That is to say, when the semantic proportion related to visual content in the search statement is greater than or equal to the preset proportion threshold (i.e., a high visual semantic search statement), use the visual media recalled by the first branch as the candidate set.
[0071] When the semantic proportion related to visual content in the search statement is less than the preset proportion threshold (i.e., a low visual semantic search statement), use the visual media recalled by the second branch as the candidate set.
[0072] That is to say, for high visual semantic search statements, recall through the visual semantics of visual media to obtain a candidate set; for low visual semantic search statements, recall through the attributes of visual media (e.g., time, location) to obtain a candidate set. In this way, the search accuracy can be improved.
[0073] In this solution, for high visual semantic search statements and low visual semantic search statements, recall the candidate set using the respective adapted recall methods, thereby improving the integrity and accuracy of the search results, and avoiding the incompleteness or inaccuracy of the search results caused by using the same recall method.
[0074] For example: When the search statement belongs to a low visual semantic search statement (e.g., photos taken in Beijing this year), if recalled according to the matching situation between the search statement and the visual content of multiple visual media, then many photos taken by users in Beijing this year will be filtered out, resulting in the incompleteness of the search results. When the search statement belongs to a high visual semantic search statement (e.g., the sky taken in Beijing this year), if recalled according to the matching situation between the time or location in the search statement and the attributes of multiple visual media, then many photos irrelevant to "the sky" will be recalled, resulting in the inaccuracy of the search results.
[0075] In a possible implementation manner, the method further includes:
[0076] Remove the semantic entities irrelevant to visual content in the search statement to obtain a rewritten search statement;
[0077] Perform semantic understanding on the rewritten search statement to obtain a second sentence semantic vector;
[0078] From the multiple visual media files, determine multiple reference visual media whose visual semantic vectors match the second sentence semantic vector;
[0079] Determine a first representative visual semantic vector corresponding to the multiple first visual media files and a second representative visual semantic vector corresponding to the multiple reference visual media;
[0080] Determine the semantic proportion related to visual content in the search statement according to the vector similarity between the first difference vector and the second difference vector;
[0081] The first difference vector represents the difference between the first sentence semantic vector and the first representative visual semantic vector; the second difference vector represents the difference between the second sentence semantic vector and the second representative visual semantic vector.
[0082] The semantic entities unrelated to visual content mentioned above mainly refer to semantic entities related to time (e.g., this year) and semantic entities related to location (e.g., Shanghai). These two types of semantic entities can be identified through Named Entity Recognition (NER). Therefore, removing the semantic entities unrelated to visual content in the search statement to obtain the rewritten search statement may specifically include: removing the semantic entities related to time and / or the semantic entities related to location in the search statement to obtain the rewritten search statement.
[0083] The greater the vector similarity between the first difference vector and the second difference vector, the greater the semantic proportion related to visual content in the search statement; the smaller the vector similarity between the first difference vector and the second difference vector, the smaller the semantic proportion related to visual content in the search statement. That is: the semantic proportion related to visual content in the search statement is positively correlated with the vector similarity between the first difference vector and the second difference vector.
[0084] In a possible implementation manner, determining the multiple first visual media files according to the M candidate visual media files includes:
[0085] According to the semantic entities related to visual content in the search statement, determine N matching dimensions corresponding to the search statement; N is greater than 1 and is an integer;
[0086] Obtain the weights of the N matching dimensions; wherein, the weight of the jth matching dimension is positively correlated with the variation degree of the matching degree of the visual content of the M candidate visual media files on the jth matching dimension; j is an integer, and the values of j range from 1 to N in sequence;
[0087] According to the weights of the N dimensions to be matched, the matching degrees of the visual content of the i-th candidate visual media file on each dimension to be matched among the N dimensions to be matched are weighted and summed to obtain the comprehensive matching degree of the i-th candidate visual media file; i is an integer, and the values of i range from 1 to M in sequence;
[0088] From the M candidate visual media files, determine the multiple first visual media files whose comprehensive matching degrees meet the preset requirements.
[0089] The above M candidate visual media files are obtained by matching the visual semantic vectors of the visual media files with the sentence semantic vectors of the search statements. That is to say, when recalling the above M candidate visual media files, the matching degree between the visual semantics of the visual media files and the entire search statement of the user is considered, rather than the matching degree between the visual semantics of the visual media files and the semantic entities that the user visually focuses on in the search statement, resulting in the recall of some inaccurate photos and videos. Therefore, in this solution, the results recalled by the above first branch are fine-tuned by means of the matching degree between the visual semantics of the visual media and the "semantic entities related to the visual content", and inaccurate visual media are filtered out.
[0090] Specifically, each semantic entity related to the visual content in the search statement and / or the search statement itself can be used as different dimensions to be matched. The roles played by different dimensions to be matched in the fine-tuning are also different, that is, the weights of different dimensions to be matched are different. The weight of each dimension to be matched is determined according to the degree of variation of the matching degree of the visual content of the M candidate visual media files on this dimension to be matched. The greater the degree of variation, the greater the role played by the dimension to be matched in the fine-tuning. Therefore, its weight is greater. In this way, some inaccurate visual media such as photos and videos can be effectively excluded.
[0091] In a possible implementation manner, the method further includes:
[0092] Match the search terms in the search statement with a preset plurality of tags to determine whether the search terms belong to semantic entities related to the tags; the tags are used to describe visual content;
[0093] Determine the semantic entities related to the tags in the search statement as the semantic entities related to the visual content in the search statement.
[0094] That is, the search terms that match a certain tag among the preset plurality of tags belong to the semantic entities related to the tag. Since the tag is related to the visual content, the semantic entities related to the tag also belong to the semantic entities related to the visual content. In this solution, by presetting a plurality of tags, the semantic entities that the user visually focuses on can be relatively simply determined from the search statement.
[0095] In a possible implementation, the clustering algorithm includes: a density-based clustering algorithm;
[0096] The method further includes:
[0097] According to the time search range, determine a first parameter involved in the density-based clustering algorithm; the first parameter is used to describe the neighborhood radius of data points; the first parameter is positively correlated with the duration corresponding to the time search range.
[0098] The method may further include:
[0099] Determine the number of visual media files among the multiple first visual media files whose acquisition times are within the time search range;
[0100] According to the ratio of the number of visual media files to the total duration corresponding to the time search range, determine a second parameter involved in the density-based clustering algorithm; the second parameter is used to describe the minimum number of data points in the neighborhood of data points.
[0101] In this embodiment, the first parameter and the second parameter involved in the density-based clustering algorithm are dynamically adjusted according to the time search range in the user's search statement to adapt to different search scenarios, thereby improving the user's search experience.
[0102] In a possible implementation, determining the search result of the search statement according to the target clustering cluster includes:
[0103] If the search statement includes a semantic entity related to a location, filter the visual media files in the target clustering cluster according to the semantic entity related to the location and the acquisition location attribute of the visual media files in the target clustering cluster;
[0104] If the search statement includes a semantic entity related to a person relationship, filter the visual media files in the target clustering cluster according to the semantic entity related to the person relationship and the person relationship attribute of the visual media files in the target clustering cluster; and / or
[0105] If the search statement includes a semantic entity related to a person name, filter the visual media files in the target clustering cluster according to the semantic entity related to the person name and the person name attribute of the visual media files in the target clustering cluster.
[0106] In this solution, location filtering, person relationship filtering, and person name filtering are performed on the recalled visual media.
[0107] In a third aspect, the present application provides an electronic device, including: a memory, a processor, and a display, where
[0108] the memory is configured to store programs;
[0109] the display is configured to display a search page;
[0110] the processor is coupled to the memory and the display, and is configured to execute the programs stored in the memory to implement any one of the methods.
[0111] In a fourth aspect, the present application provides a computer-readable storage medium storing a computer program, where the computer program, when executed by a computer, can implement any one of the methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0112] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation to the present application. In the drawings:
[0113] Figure 1A FIG. is a set of interface diagrams of a search interface for a mobile phone to enter a gallery application provided by an embodiment of the present application;
[0114] Figure 1B FIG. is a schematic diagram of a search interface after clearing search history provided by an embodiment of the present application;
[0115] Figure 1C FIG. is a set of interface diagrams related to searching in a gallery application provided by an embodiment of the present application;
[0116] Figure 1D FIG. is a schematic diagram of a search failure interface provided by an embodiment of the present application;
[0117] Figure 2A FIG. is a schematic structural diagram of an electronic device provided by another embodiment of the present application;
[0118] Figure 2B FIG. is a software structure block diagram of an electronic device provided by another embodiment of the present application;
[0119] Figure 3A FIG. is another set of interface diagrams related to searching in a gallery application provided by an embodiment of the present application;
[0120] Figure 3B FIG. is a set of interface diagrams related to searching in the negative first screen provided by an embodiment of the present application;
[0121] Figure 3C FIG. is a first schematic diagram of a search result interface provided by an embodiment of the present application;
[0122] Figure 3D Schematic diagram II of the search result interface provided by an embodiment of the present application;
[0123] Figure 3E Schematic diagram III of the search result interface provided by an embodiment of the present application;
[0124] Figure 3F Schematic diagram of the search result interface provided by an embodiment of the present application Figure Four ;
[0125] Figure 4 Interaction diagram of the visual media search method provided by an embodiment of the present application;
[0126] Figure 5 Flow schematic diagram of the visual media search method provided by an embodiment of the present application;
[0127] Figure 6 Schematic diagram of the search result interface provided by an embodiment of the present application Figure Five 。 Detailed implementation manners
[0128] Next, the technical solutions in the embodiments of the present application will be described in conjunction with the accompanying drawings in the embodiments of the present application. Among them, in the description of the embodiments of the present application, unless otherwise specified, " / " means "or". For example, A / B may represent A or B; herein, "and / or" is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of the present application, "a plurality of" means two or more than two.
[0129] Hereinafter, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first", "second", and "third" may explicitly or implicitly include one or more of such features.
[0130] First, the terms involved in the embodiments of the present application will be described. It can be understood that this description is for a clearer understanding of the embodiments of the present application and does not necessarily constitute a limitation to the embodiments of the present application.
[0131] Visual media: refers to pictures or videos.
[0132] Semantic subject: The named entity recognition technology can identify text and recognize entities with specific meanings in the text, such as person names (PER), place names (LOC), etc. In this solution, the entities identified with specific meanings are called semantic subjects.
[0133] Visual content related and visual content unrelated: Visual content refers to the objects presented by visual media and their interrelationships, etc. Computer vision enables a computer to have capabilities similar to human vision, including perceiving, understanding, analyzing, and interpreting visual content. Currently, the Generative Pre-trained Transformer 4 (GPT-4) can support inputting an image into the model and then outputting human natural language describing the important information in the image.
[0134] In the context of image search in this solution, the data that can only be obtained after the natural image understanding of the model by the visual media file is called "visual content related". That is to say, visual content is the data that needs to be obtained through the natural image understanding model. In this solution, the data that is related to the visual media file and can be obtained without the image understanding ability of the model is called "visual content unrelated", such as the shooting location, shooting time, name, file attributes, etc. that can be obtained and saved when the terminal device collects the visual media file.
[0135] For example, in the "photo taken in Beijing this year", "this year" (shooting time), "Beijing" (shooting location), and "photo" (file attribute) are all data that can be obtained and saved when the terminal device collects the visual media file. Therefore, "this year", "Beijing", and "photo" are unrelated to visual content; in the "sky in the photo taken in Beijing this year", "sky" can only be obtained by the model's image understanding ability to understand the image or image frame of the visual media file. Therefore, "sky" is visual content related.
[0136] Text semantic vector: It can be obtained by sending the text into a text encoder. It is a vector that can represent the semantic features of the entire sentence. The text encoder can use models such as Transformer commonly used in Natural Language Processing (NLP). This solution does not limit this here. In this solution, the text semantic vector obtained for a sentence is called a sentence semantic vector, the text semantic vector obtained for the semantic subject in a sentence is called a subject semantic vector, and the text semantic vector obtained for a label is called a label semantic vector.
[0137] Visual semantic vector: It can be obtained by sending the image or image frame of the visual media file into an image encoder. Commonly used CNN (Convolutional Neural Network) models or VIT (Vision Transformer) models can be used. This solution does not limit this here.
[0138] Density-based clustering algorithm: It describes the tightness of a sample set based on a group of neighborhoods. (The first parameter ∈, the second parameter MinPts) is used to describe the tightness of the sample distribution in the neighborhood. The first parameter ∈ is used to describe the neighborhood radius of a data point; the second parameter MinPts is used to describe the minimum number of data points in the neighborhood of a data point. Its representative algorithms include: DBSCAN (Density-Based Spatial Clustering of Application with Noise, density-based spatial clustering of applications with noise); the DBSCAN algorithm is a relatively representative density-based clustering algorithm that can divide regions with sufficient high density into clusters and can discover clusters of any shape in a spatial database with noise;
[0139] Vector similarity: It is used to describe the similarity between two vectors (for example: between a sentence semantic vector and a visual semantic vector). In the embodiments of the present application, the visual media that matches the search statement can be determined by comparing the similarity between the sentence semantic vector and the visual semantic vector. Generally, the vector similarity can be calculated through the cosine similarity calculation formula. Of course, it can also be calculated in other ways.
[0140] In the prior art, a mobile phone manages visual media files such as pictures and videos of users through a gallery application (hereinafter referred to as: visual media). Taking the example of a mobile phone taking a photo, after the mobile phone takes a photo, the gallery application can obtain and save attributes unrelated to the visual content such as the shooting location, shooting time, and photo name of the photo.
[0141] In practical applications, the photo can also be input into the natural picture understanding model in the state of the mobile phone being charged and the screen being turned off, so that the natural picture understanding model generates and saves the label of the photo. This label can be regarded as an attribute related to the visual content of the photo. The label can be: "sky", "cat", "dog", etc. The gallery application can establish an index for the photo based on the attributes of the photo. After the index is established, the gallery application can provide corresponding search services to users. Specifically, users can search for pictures or videos by entering keywords in the gallery application. Exemplarily, users can enter keywords such as "Beijing", "sky", "National Day" in the search box provided by the gallery application, and the gallery application matches the keywords entered by the user with the indexes of visual media such as pictures and videos in the gallery application, and then obtains the search results.
[0142] The following will describe the interface involved in the search process of the gallery application in the prior art with reference to the accompanying drawings:
[0143] As Figure 1AAs shown in (a) of [Figure ID], the mobile phone can display the main interface 101, which can also be called the desktop. The main interface 101 may include the icon 102 of the gallery application. The mobile phone can receive the operation of the user clicking on the icon 102. In response to this operation, the mobile phone can launch the gallery application and display the interface 103 as shown in Figure 1A (b) of [Figure ID]. Among them, the interface 103 can be an album interface. It should be noted that in response to the operation of the user clicking on the icon 102, the mobile phone can launch the gallery application and display the photo interface of the gallery. The photo interface includes thumbnails of the photos (i.e., pictures) in the gallery or the large picture of a certain photo. In the photo interface, in response to the operation of the user on the "Album" control, the above-mentioned album interface 103 is displayed.
[0144] As Figure 1A (b) of [Figure ID] shows, the interface 103 includes multiple albums. Among them, the "All Photos" album includes 2,023 photos, the "Camera" album includes 1,502 photos and videos, the "Screenshots & Screen Recordings" album includes 102 photos and videos, the "My Favorites" album includes 48 photos and videos, the "One-Take, Many Gains" album has 34 photos and videos, the "Video Editing" album has 65 videos, the "Self-created Album" has 57 photos and videos, and the "Shared Album" has 100 photos and videos.
[0145] As Figure 1A (b) of [Figure ID] shows, the interface 103 may include a search box 104. The mobile phone can receive the operation of the user clicking on the search box 104. In response to this operation, the mobile phone can display the interface 105 as shown in Figure 1A (c) of [Figure ID]. This interface 105 can be called a search interface. Among them, the interface 105 can display the classification information of the photos to the user. For example, in the interface 105, the mobile phone classifies the photos of the local machine according to time, people, and things, etc. For example, in the dimension of time, the mobile phone classifies the photos of the local machine according to the three time periods of "This Month", "Last Month", and "This Year" respectively. Among them, the "This Month" album includes the photos or videos taken by the mobile phone this month, the "Last Month" album includes the photos or videos taken by the mobile phone last month, and the "This Year" album includes the photos or videos taken by the mobile phone this year. In the dimension of people, the mobile phone classifies the photos of the local machine according to different people, such as the four different people in the interface 105. In the dimension of things, the mobile phone classifies and displays the photos of the local machine according to "Scenery", "Animals", "Documents", and "Buildings". It should be noted that the above classification dimensions can also be others, which are not specifically limited here. In the interface 105, the user can see this classification information without entering keywords.
[0146] Optionally, the interface 105 may further include a search history 107 and an option “Clear” 108. The search history includes keywords that the user has entered, such as “flowers”, “coffee”, “cats”, etc. The mobile phone can receive an operation where the user clicks “Clear” 108, and in response to this operation, the mobile phone can clear the search history. After the mobile phone clears the search history, the keywords that the user has entered are no longer displayed on the search interface 105. For example, in response to the user clicking the “Clear” 108 operation, as Figure 1B shown, the search history 107 and the option “Clear” 108 are no longer displayed on the search interface 105, and the content displayed below moves up.
[0147] In response to the user's operation of entering the keyword “sky” on the interface 105, the mobile phone displays the interface 109 as shown in Figure 1C (a). As shown in Figure 1C (a), there are 100 photos related to “sky” and 32 photos related to photos containing the word “sky”. Among them, the 100 photos related to “sky” can be recalled because the tags of these 100 photos match “sky”; the 32 photos related to photos containing the word “sky” can be recalled because through OCR (Optical Character Recognition) technology, it is recognized that these 32 photos contain the word “sky”. In practical applications, the mobile phone can also associate with the keywords entered by the user to obtain associated words and perform a search based on the associated words.
[0148] The interface 109 also shows some search results of the keyword “sky” and a “More” option 110 corresponding to the search results of the keyword “sky”. The mobile phone receives a click operation by the user on the “More” option 110 and displays the interface 111 as shown in Figure 1C (b). Among them, the interface 111 is used to display photos and videos in the search results of the keyword “sky”. Optionally, the photos and videos can be classified and displayed according to time. In addition, the interface 111 also includes a return key 112 and a title 113. In response to the user's operation on the return key 112, the mobile phone can redisplay the interface 109. The title 113 may include the keyword “sky”.
[0149] That is to say, in the existing gallery applications, when a user enters simple keywords in the search box, such as: sky, Beijing, National Day, etc., corresponding search results can be obtained. However, since the number of photo attributes is simple and limited, and the mobile phone's ability to understand and associate with search statements is also limited. When the user enters a relatively complex search statement in the search box, if the keywords in the search statement cannot match the attributes of the picture or the text in the picture, no photos can be found. That is to say, the existing gallery applications do not support the search function based on complex search statements. As Figure 1D shown, when the user enters a relatively complex search statement "warming oneself by the fire and brewing tea" in the search box of interface 114, the mobile phone cannot understand the associated words of "warming oneself by the fire and brewing tea". Since the photo has no attributes that can match "warming oneself by the fire and brewing tea" or its associated words, the search result shows "no pictures".
[0150] However, in practical applications, users have a strong demand for the function of searching for pictures based on complex search statements. This is because users can describe the pictures or videos they want more comprehensively through complex search statements, thereby achieving precise search. To meet this demand of users, an embodiment of this application provides a visual media search method. This method can be applied to an electronic device, and the electronic device can be a mobile phone, a tablet computer, a wearable device, a vehicle-mounted device, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), and other terminal devices.
[0151] Exemplarily, Figure 2A FIG. shows a schematic structural diagram of an electronic device 200. The electronic device 100 may include a processor 210, an external memory interface 220, an internal memory 221, a universal serial bus (USB) interface 230, a charging management module 240, a power management module 241, a battery 242, an antenna 1, an antenna 2, a mobile communication module 250, a wireless communication module 260, an audio module 270, a speaker 270A, a receiver 270B, a microphone 270C, a headphone interface 270D, a sensor module 280, a button 290, a motor 291, an indicator 292, a camera 293, a display screen 294, and a subscriber identification module (SIM) card interface 295, etc.
[0152] Among them, the sensor module 280 may include a pressure sensor 280A, a gyroscope sensor 280B, a barometric pressure sensor 280C, a magnetic sensor 280D, an acceleration sensor 280E, a distance sensor 280F, a proximity light sensor 280G, a fingerprint sensor 280H, a temperature sensor 280J, a touch sensor 280K, an ambient light sensor 280L, a bone conduction sensor 280M, etc.
[0153] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device 200. In some other embodiments of the present application, the electronic device 200 may include more or fewer components than those shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0154] The processor 210 may include one or more processing units. For example, the processor 210 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0155] The controller may generate an operation control signal according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.
[0156] A memory may also be provided in the processor 210 for storing instructions and data. In some embodiments, the memory in the processor 210 is a cache memory. This memory may store the instructions or data that the processor 210 has just used or recycled. If the processor 210 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 210, and thus improves the efficiency of the system.
[0157] In some embodiments, the processor 210 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0158] It can be understood that the interface connection relationships among the modules illustrated in the embodiments of the present application are only illustrative descriptions and do not constitute a structural limitation on the electronic device 200. In other embodiments of the present application, the electronic device 200 may also adopt different interface connection manners in the above embodiments, or a combination of multiple interface connection manners.
[0159] The electronic device 200 realizes the display function through a GPU (Graphics Processing Unit), a display screen 294, and an application processor, etc. The GPU is connected to the display screen 294 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 210 may include one or more GPUs, which execute program instructions to generate or change display information.
[0160] The display screen 294 is used to display images, videos, etc. The display screen 294 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include one or N display screens 294, where N is a positive integer greater than 1.
[0161] The electronic device 200 can implement the shooting function through the ISP, the camera 293, the video codec, the GPU, the display screen 294, and the application processor, etc.
[0162] The camera 293 is used to capture still images or videos. An object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the ISP to convert it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in standard RGB, YUV, etc. formats. In some embodiments, the electronic device 100 may include one or N cameras 293, where N is a positive integer greater than 1.
[0163] The video codec is used to compress or decompress digital videos. The electronic device 200 can support one or more video codecs. In this way, the electronic device 200 can play or record videos in multiple coding formats, such as: Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0164] The NPU is a neural-network (NN) computing processor. By drawing on the structure of biological neural networks, such as the transmission pattern between human brain neurons, it can quickly process input information and can also continuously learn on its own. Through the NPU, applications such as the intelligent cognition of the electronic device 200 can be realized, such as: image recognition, face recognition, speech recognition, text understanding, etc.
[0165] The external memory interface 220 can be used to connect to an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 200. The external memory card communicates with the processor 210 through the external memory interface 220 to implement the data storage function. For example, files such as music and videos are saved in the external memory card.
[0166] The internal memory 221 can be used to store computer-executable program code, and the executable program code includes instructions. The internal memory 221 can include a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, image playback function, etc.). The data storage area can store the data created during the use of the electronic device 200 (such as audio data, phone book, etc.). In addition, the internal memory 221 can include high-speed random access memory and can also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. The processor 210 executes various functional applications and data processing of the electronic device 200 by running the instructions stored in the internal memory 221 and / or the instructions stored in the memory provided in the processor.
[0167] Figure 2B It is a software structure block diagram of the electronic device 200 in the embodiment of the present application. The software system of the electronic device 200 can adopt a layered architecture. The layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. Taking the Android system as an example, in some embodiments, the Android system can be divided into four layers, from top to bottom are the application layer, application framework layer, Android runtime and system libraries, and kernel layer.
[0168] As Figure 2B shown, the application layer can include application programs such as a gallery service module, a search module, a multimodal understanding module, a natural language understanding module, and a camera application.
[0169] The application framework layer provides application programming interfaces (APIs) and programming frameworks for applications in the application layer. The application framework layer includes some predefined functions.
[0170] The system library can include multiple functional modules. For example: surface manager, Media Libraries, 3D graphics processing library (e.g., OpenGL ES), 2D graphics engine (e.g., SGL), etc. Among them, the media library supports the playback and recording of multiple common audio and video formats, as well as static image files, etc. The media library can support multiple audio and video coding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.
[0171] The kernel layer is the layer between hardware and software.
[0172] Next, in combination with the capture and photo-taking scenario, the working processes of the software and hardware of the electronic device 200 will be exemplarily described.
[0173] When the touch sensor 280K receives a touch operation, the corresponding hardware interrupt is sent to the kernel layer. The kernel layer processes the touch operation into a raw input event (including information such as touch coordinates and the timestamp of the touch operation). The raw input event is stored in the kernel layer. The application framework layer obtains the raw input event from the kernel layer and identifies the control corresponding to the input event. Taking the touch operation as a touch click operation and the control corresponding to the click operation being the control of the camera application icon as an example, the camera application calls the interface of the application framework layer to start the camera application, and then starts the camera driver by calling the kernel layer to capture a static image or video through the camera 293.
[0174] Taking the electronic device as a mobile phone as an example, the visual media search method provided by the embodiments of the present application will be introduced. The visual media search method provided by the embodiments of the present application can be applied to application programs such as the gallery application and the file management application.
[0175] Next, the interfaces and search logics related to the visual media search method provided by the embodiments of the present application will be described in conjunction with the accompanying drawings.
[0176] As Figure 3A shown in (a) of [reference], the interface 301 displays a search history 303 and a "Clear" option 304. The search history includes search statements that the user has entered, such as: "Watching the sunrise on the mountain top", "The Great Wall photographed in Beijing this year". Other content that can be displayed on the interface 301 can refer to the relevant display content of the above interface 305 and will not be elaborated here. Among them, the interface 301 can be called a search interface. The mobile phone can respond to the user's operation onFigure 1A Performing a click operation on the search box 104 in the interface 103 shown in (b) in
[0177] As Figure 3A shown in (b) in Figure 3A the interface 305 (i.e., the search result interface), the user enters the search statement "warming the tea around the stove" in the search box 306 of the interface 305, and the mobile phone searches for 239 pictures. The mobile phone displays some search results of the search statement "warming the tea around the stove" (e.g., thumbnails of 8 pictures) and the "More" option 307 corresponding to the search results of the search statement "warming the tea around the stove" on the interface 305. In response to the user's operation on the "More" option 307, the mobile phone displays the interface 308 shown in (c) in
[0178] Currently, the mobile phone also has a negative first screen, a pull-down search interface, etc. It can be understood that the negative first screen can be the leftmost split screen of the electronic device, which is used to provide functions such as search and quick services for users. Among them, the negative first screen can also be used to display notification messages to be pushed to users, such as application messages subscribed by users, real-time hot search messages, segment selections, itinerary information, etc. The pull-down search interface is an interface displayed in response to the user's pull-down operation on the main interface, and this interface is used to provide functions such as search and application suggestions for users. This interface and Figure 3B the interface 315 in
[0179] can be the same interface.
[0180] Exemplarily, referring to Figure 3B shown in (a) in Figure 3B the mobile phone can receive a first operation performed by the user on the interface 309 (which can be called the desktop) of the mobile phone. Exemplarily, this first operation can be a rightward sliding operation as shown in (a) in Figure 3B . In response to this first operation, the mobile phone can display the negative first screen 310 shown in (b) in
[0181] The mobile phone receives a click operation by the user on the search box 311 on the negative first screen 310, and displays the interface 315 as shown in Figure 3B (c) below. The interface 315 may include: a search box 316, application suggestions, and a search history 217. The application suggestions include: icons of various applications recommended for use. The interface 215 may also include: the search history 217 and its corresponding "Clear" option 318. In response to the user's trigger operation on the "Clear" option 318, the search history 317 and the "Clear" option 318 are no longer displayed on the interface 315. In addition, a hot news title, such as: "Tianjin Marathon", may be displayed in the search box 316.
[0182] As shown in Figure 3B (d) below, the search box of the interface 319 displays the search statement "Great Wall photographed in Beijing during the National Day" entered by the user; the interface 319 also displays a preview area 322 of the search results of the gallery for the search statement "Great Wall photographed in Beijing during the National Day" and the "Search in App" option 323 corresponding to the gallery. In response to the user's trigger operation on the preview area 322, the mobile phone enters the photo details interface provided by the gallery application for the user to flip through the search results of the search statement "Great Wall photographed in Beijing during the National Day". In response to the user's trigger operation on the "Search in App" option 323, the mobile phone displays the interface 324 provided by the gallery application as shown in Figure 3B (e) below. The interface 324 displays some search results of the search statement "Great Wall photographed in Beijing during the National Day" and the "More" option corresponding to the search results of the search statement "Great Wall photographed in Beijing during the National Day". In response to the user's trigger operation on the "More" option, the mobile phone may display a search result details interface that shows the pictures in the search results of the search statement "Great Wall photographed in Beijing during the National Day". The interface 319 may also display an online search option 321. In response to the user's trigger operation on the online search option 321, the mobile phone displays a search web page and shows the online search results in the search web page.
[0183] As shown in Figure 3CAs shown in the interface 325, the user enters the search statement "The Great Wall taken during last year's National Day" in the search box of the interface 325, and the mobile phone searches for 419 pictures. The mobile phone displays the thumbnails of each picture or video in the search results of the search statement "The Great Wall taken during last year's National Day" on the interface 325. Among them, the shooting time of the picture or video corresponding to thumbnail A is 23:22 on September 30, 2022; the shooting time of the picture or video corresponding to thumbnail B is 00:19 on October 1, 2022. Among them, the time search range corresponding to National Day is from 00:00 on October 1, 2022 (i.e., the first time point) to 00:00 on October 8, 2022 (i.e., the second time point). The shooting time of the picture or video corresponding to thumbnail A is before 00:00 on October 1, 2022. The shooting time of the picture or video corresponding to thumbnail B is 00:19 on October 1, 2022, which is between 00:00 on October 1, 2022 and 00:00 on October 8, 2022. Among them, the shooting times of the pictures or videos corresponding to thumbnail A and thumbnail B are relatively close.
[0184] As Figure 3D shown in the interface 326, for the search statement "The sky taken during last year's National Day", the shooting time of the picture or video corresponding to the displayed thumbnail C is 22:19 on October 7, 2022; the shooting time of the displayed picture D corresponding to the picture or video is 01:24 on October 8, 2022. Among them, the time search range corresponding to National Day is from 00:00 on October 1, 2022 (i.e., the first time point) to 00:00 on October 8, 2022 (i.e., the second time point). The shooting time of the picture or video corresponding to thumbnail D is after 00:00 on October 8, 2022. The shooting time of the picture or video corresponding to thumbnail C is 00:19 on October 1, 2022, which is between 00:00 on October 1, 2022 and 00:00 on October 8, 2022. Among them, the shooting times of the pictures or videos corresponding to thumbnail C and thumbnail D are relatively close.
[0185] In practical applications, the search statement may include, in addition to the time search range (such as National Day), a first keyword (such as The Great Wall or the sky). Then, for the search statement, the visual media file corresponding to its search result matches the first keyword.
[0186] When the first keyword belongs to a semantic entity unrelated to visual content (such as: location), the first keyword can be matched with the attributes of the visual media to determine the visual media file that matches the first keyword.
[0187] When the first keyword belongs to a semantic entity related to visual content (such as the Great Wall or the sky), match the first keyword with the tags of the visual media file to determine the visual media file whose visual content matches the first keyword; or, match the text semantic vector of the first keyword with the visual semantic vector of the visual media file to determine the visual media file whose visual content matches the first keyword (including the above-mentioned first, second, and third visual media files).
[0188] The above thumbnail is a proportional thumbnail of the corresponding visual media file or a proportional thumbnail of any image frame.
[0189] As Figure 3E Shown in the interface 327, in which, for the search statement "the sky photographed during last year's National Day", the two thumbnails E and F shown are sorted and displayed in descending order according to the matching degree between the visual content of their respective visual media files and the first keyword "sky". Among them, the matching degree between the visual content of the visual media file corresponding to the thumbnail E and "sky" is 0.87; the matching degree between the visual content of the visual media file corresponding to the thumbnail F and "sky" is 0.76.
[0190] As Figure 3F Shown in the interface 328, in which, for the search statement "photos taken in Nankai District, Tianjin during last year's National Day", the two thumbnails J and H shown are sorted and displayed in descending order according to the matching degree between the location attributes of their respective visual media files and the first keyword "Nankai District, Tianjin". Among them, the matching degree between the location attribute of the visual media file corresponding to the thumbnail J and "Nankai District, Tianjin" is 0.87; the matching degree between the location attribute of the visual media file corresponding to the thumbnail H and "Nankai District, Tianjin" is 0.8.
[0191] Figure 4 This is an interaction diagram of the visual media search method provided by the embodiments of the present application. As Figure 4 Shown in the figure, a mobile phone is provided with: a gallery service module (i.e., a gallery application) 41, a search module 42, a multimodal understanding module 43, and a natural language understanding module 44.
[0192] As Figure 4 Shown in the figure, the visual media search method provided by the embodiments of the present application can be divided into two stages: an index construction stage and a search stage.
[0193] In the index construction stage, the following steps are included:
[0194] S401, Add and / or modify visual media and its attributes.
[0195] In the above S401, for the newly added visual media, the attributes that the gallery application can automatically generate and are unrelated to the visual content may include but are not limited to: collection location, collection time, and visual media name. Taking the captured video or picture as an example, the collection location refers to the shooting location, and the collection time refers to the shooting time; taking the screenshot as an example, the collection location refers to the screenshot location, and the collection time refers to the screenshot time; taking the downloaded video or picture as an example, the collection location refers to the download location, and the collection time refers to the download time.
[0196] Users can add new visual media by means such as shooting, downloading, and taking screenshots. In addition, users can also modify the existing visual media. Such modifications include but are not limited to: beautification, custom naming, adding watermarks, etc.
[0197] S402. Store the visual media and its attributes.
[0198] The gallery service module 41 can, in response to the above-mentioned addition or modification operations, store the visual media and its attributes locally on the mobile phone. In actual applications, with the user's authorization, the mobile phone can store the locally stored visual media and its attributes in the cloud to relieve the storage pressure on the local mobile phone.
[0199] S403. Request visual semantic understanding of the visual media.
[0200] Since visual semantic understanding requires a large amount of computing resources, in order not to affect the user's use, the above step S403 can be executed when the mobile phone is in the state of being charged and the screen is off.
[0201] The gallery service module 41 can request the multimodal understanding module 43 to perform visual semantic understanding on the newly added or modified visual media to obtain the visual semantic vector of the visual media.
[0202] Among them, the multimodal understanding module 43 can perform visual semantic understanding on the visual media based on the multimodal model to obtain the visual semantic vector of the visual media.
[0203] Among them, the multimodal model is not only used for: performing visual semantic understanding on the visual media to obtain the visual semantic vector of the visual media; but also for: performing semantic understanding on the search statement to obtain the sentence semantic vector of the search statement; performing semantic understanding on the rewritten search statement in the following text to obtain the sentence semantic vector of the rewritten search statement; performing semantic understanding on the semantic subject in the search statement to obtain the subject semantic vector of the semantic subject; performing semantic understanding on the label of the visual media to obtain the label semantic vector of the label. The multimodal model can be trained according to the training samples.
[0204] Exemplarily, the multimodal model can specifically be CLIP (Contrastive Language-Image Pre-training). The CLIP model can map visual media and text (i.e., search statements, rewritten search statements, semantic entities, labels) into a unified vector space to understand the relationships between different modal resources visually and textually, and then be used for image retrieval. That is, in the embodiments of the present application, the visual media file and text can be specifically matched through the CLIP model.
[0205] The multimodal model can map visual media and text into vectors of the same dimension. That is to say, the dimension of the visual semantic vector of visual media is the same as that of the semantic vector of text (e.g., the sentence semantic vector of the search statement). The multimodal model includes: the above-mentioned image encoder and text encoder.
[0206] S404. Return the visual semantic vector of the visual media.
[0207] The multimodal understanding module 44 returns the visual semantic vector of the visual media to the gallery service module 41.
[0208] S405. Store the visual semantic vector of the visual media.
[0209] The gallery service module 41 can locally store the visual semantic vector of the visual media.
[0210] S406. Send the attribute information of the visual media and its visual semantic vector.
[0211] Exemplarily, the gallery service module 41 can store the visual semantic vector of the visual media returned by the multimodal understanding module 44, and then batch send the attributes of the visual media and its visual semantic vector to the search module 42 for the search module 42 to construct an index of the visual media.
[0212] S407. Construct an index
[0213] The index of the visual media constructed by the search module 42 can include: the attributes of the visual media, the visual semantic vector of the visual media.
[0214] In the search stage, the following steps are included:
[0215] S408. Input a search statement.
[0216] The user can input a search statement through the search interface provided by the gallery service module 41, such as: Figure 3A the interface 301 shown in (a) of Figure 3AFor the interface 305 shown in (b) in [reference], enter "warming the tea around the stove" in the search box 306.
[0217] S409: Send the search statement.
[0218] After the gallery service module 41 receives the search statement input by the user, it sends the search statement to the search module 42 for searching.
[0219] S410: Request semantic entity recognition for the search statement.
[0220] The search module 42 requests the natural language understanding module 44 to perform semantic entity recognition on the search statement to obtain the semantic entities included in the search statement. Among them, the natural language understanding module 44 performs semantic entity recognition based on a natural language understanding model. Specifically, named entity recognition technology (NER) can be used to perform semantic entity recognition on the search statement to obtain the semantic entities included in the search statement. In the embodiments of the present application, the semantic entity can also be referred to as an entity.
[0221] Using named entity recognition technology, semantic entities related to time, location, and tags in the search statement can be recognized. Among them, semantic entities related to time and location are semantic entities unrelated to visual content; the tags of visual media files are data that can only be obtained through the natural image understanding of the model. Therefore, semantic entities related to tags are semantic entities related to visual content. In practical applications, based on practical experience, multiple tags that users are more concerned about can be counted, such as: "sky", "cat", "dog", "birthday", "child", etc. These tags are used to describe visual content. Subsequently, named entity recognition technology can match the keywords (or search terms) in the search statement with multiple pre-set tags to determine whether the keyword belongs to a semantic entity related to a tag.
[0222] Exemplarily, using named entity recognition technology to perform semantic entity recognition on the search statement "the sky photographed in Beijing during the National Day", it is determined that "National Day" belongs to a semantic entity related to time, "Beijing" belongs to a semantic entity related to location, and "sky" belongs to a semantic entity related to a tag.
[0223] S411: Return the semantic entity.
[0224] The natural language understanding module 44 returns the recognized semantic entity to the search module 42.
[0225] S412: Request semantic understanding of the search statement.
[0226] The search module 42 can send the search statement to the multimodal understanding module 43, and the multimodal understanding module 43 performs semantic understanding on the search statement to obtain the sentence semantic vector of the search statement (i.e., the first sentence semantic vector). For the specific semantic understanding process, reference can be made to the corresponding content in the above embodiments, which will not be elaborated here.
[0227] It should be additionally supplemented that when the search statement includes a semantic entity related to visual content (i.e., a semantic entity related to a label), the search module 42 can also send the semantic entity related to visual content to the multimodal understanding module 43, so that the multimodal understanding module 43 performs semantic understanding on the semantic entity related to visual content to obtain the entity semantic vector of this semantic entity. Continuing with the above example, "sky" belongs to the semantic entity related to visual content, and the multimodal understanding module 43 can perform semantic understanding on "sky" to obtain the entity semantic vector corresponding to "sky".
[0228] S413. Return the text vector.
[0229] The multimodal understanding module 43 can return the sentence semantic vector of the search statement to the search module 42.
[0230] S414. Perform recall respectively based on the attributes of the visual media and the visual semantic vector of the visual media.
[0231] The search module 42 includes different branches of search methods:
[0232] Exemplarily, the first branch is: performing recall based on the visual semantic vector of the visual media. Specifically, obtain the visual semantic vectors of each visual media among the multiple visual media stored in the mobile phone; calculate the vector similarity between the sentence semantic vector of the search statement and the visual semantic vectors of each visual media; according to the vector similarity, determine M candidate visual media (i.e., M candidate visual media files) whose visual semantic vectors match the sentence semantic vector from the multiple visual media stored in the mobile phone; where M is an integer greater than or equal to 1; these M visual media can be used as the visual media set recalled by the first branch.
[0233] Exemplarily, use the visual media with a vector similarity greater than the preset similarity threshold as the visual media whose visual semantic vector matches the sentence semantic vector.
[0234] Exemplarily, sort the multiple visual media according to the vector similarity from high to low, and use the top F (F≥1) visual media as the visual media whose visual semantic vector matches the sentence semantic vector.
[0235] Exemplarily, Q (Q≥1) visual media with vector similarity greater than a preset similarity threshold are determined from multiple visual media stored via the mobile phone; if Q is greater than or equal to the preset quantity threshold D, these Q visual media can be sorted in descending order of vector similarity; the top D visual media in the sorting are used as the visual media whose visual semantic vectors match the semantic vector of this sentence; if Q is less than the preset quantity threshold D, these Q visual media can be directly used as the visual media whose visual semantic vectors match the semantic vector of this sentence.
[0236] It should be noted that the multiple visual media stored via the mobile phone may include: visual media stored locally by the mobile phone and / or visual media stored in the cloud by the mobile phone (for example: the cloud storage space applied for by the mobile phone). To protect user privacy, all the multiple visual media stored via the mobile phone are stored in the mobile phone.
[0237] It should be noted that during the recall process of the first branch, the attributes of the visual media are not understood, and only the overall visual semantic information of the visual media can be understood.
[0238] Exemplarily, the second branch is: recall based on the attributes of the visual media.
[0239] Specifically, obtain the attributes of each visual media among the multiple visual media stored via the mobile phone; match the semantic subject in the search statement with the attributes of each visual media to determine the visual media (i.e., the second visual media) that matches this semantic subject; use the visual media that matches this semantic subject as the set of visual media recalled by the second branch. When the number of semantic subjects in the search statement is one, the set of visual media recalled by the second branch includes the visual media that matches this one semantic subject; when the number of semantic subjects in the search statement is multiple, the set of visual media recalled by the second branch includes the visual media that each of these multiple semantic subjects matches.
[0240] Exemplarily, for a semantic subject related to time, the corresponding time range (i.e., the time search range) of this semantic subject can be determined; match this time range with the acquisition time of each visual media as an attribute to determine the visual media whose acquisition time is within this time range; use the visual media whose acquisition time is within this time range as the visual media that matches this semantic subject. For example: if the semantic subject is "National Day", its corresponding time range is "from October 1st to October 7th", the acquisition time of Picture 1 is "October 2nd", and the acquisition time of Picture 2 is "October 8th", then, according to the above matching method, Picture 1 matches the semantic subject "National Day", and Picture 2 does not match the semantic subject "National Day".
[0241] Exemplarily, for a semantic entity related to a location, the geographical range corresponding to the semantic entity (i.e., the geographical search range) can be determined; the geographical range is matched with the attribute of the collection location of each visual medium to determine the visual media whose collection location is within the geographical range; the visual media whose collection location is within the geographical range is used as the visual media matched by the semantic entity. For example: the semantic entity is "Beijing", its corresponding geographical range is the whole Beijing City, the collection location of Picture 3 is "Xicheng District, Beijing City", and the collection location of Picture 4 is "Nankai District, Tianjin City". Then, according to the above matching method, Picture 3 matches the semantic entity "Beijing", and Picture 4 does not match the semantic entity "Beijing".
[0242] Exemplarily, for a semantic entity related to a tag, the main semantic vector of the semantic entity and the tag semantic vectors of the tags of each visual medium can be obtained; the vector similarity between the main semantic vector of the semantic entity and the tag semantic vectors of the tags of each visual medium is calculated; according to the vector similarity, the visual media matched by the semantic entity is determined. For example: the semantic entity is "human cub", and Picture 5 has a tag of "child". Through calculation, it is found that the main semantic vector of "human cub" is similar to the tag semantic vector of "child", that is, Picture 5 matches the semantic entity "human cub". In practical applications, after obtaining the set of visual media recalled by the first branch and the set of visual media recalled by the second branch, the candidate set of visual media can be determined according to the set of visual media recalled by the first branch and the set of visual media recalled by the second branch. In an optional implementation manner, the union or intersection of the set of visual media recalled by the first branch and the set of visual media recalled by the second branch can be used as the candidate set of visual media.
[0243] In practical applications, when a user searches for pictures and the like on a mobile phone, sometimes the user focuses on the visual semantic information of the pictures, sometimes on the attribute information such as the shooting location and shooting time of the pictures, and sometimes on both. Exemplarily, when the user searches for "pictures taken today", the user focuses on the shooting time of the pictures; when the user searches for "the sky taken today", the user focuses not only on the shooting time of the pictures, but also on the visual semantics of the pictures, that is, whether the picture content is the sky; when the user searches for "pictures taken while walking in Beijing", the user focuses on the shooting location of the pictures; when the user searches for "pictures of walking taken in Beijing", the user focuses not only on the shooting location attribute of the pictures, but also on the visual semantics of the pictures, that is, whether the picture content is a walking picture.
[0244] Taking the two search statements of "photos taken in Beijing this year" and "the sky taken in Beijing this year" as examples, referring to the foregoing introduction, the semantic proportion of visual content in the search statement of "the sky taken in Beijing this year" is greater than that in the search statement of "photos taken in Beijing this year". Obviously, for the search statement of "photos taken in Beijing this year", it is more suitable to use the visual media set recalled by the second branch as the candidate visual media set. Subsequently, the intersection of the visual media matching "this year" and the visual media matching "Beijing" in the candidate visual media set can be used as the final search result. The collection locations of the visual media in the final search result are all Beijing, and the collection time is this year, which meets the user's search requirements. If the visual media set recalled by the first branch is used as the candidate visual media set, the following situation may occur: There is a picture of "a girl holding a camera to take a photo" in the visual media set recalled by the first branch, but the shooting location of this photo is Shanghai and the shooting time is last year. Since the word "take" exists in the search statement of "photos taken in Beijing this year" and the action of "taking" exists in the picture of "a girl holding a camera to take a photo", there is a certain similarity between the sentence semantic vector of the search statement of "photos taken in Beijing this year" and the visual semantic vector of the picture of "a girl holding a camera to take a photo". That is to say, the picture of "a girl holding a camera to take a photo" may be recalled. That is, when the semantic proportion of visual content is relatively low, it is not appropriate to use the visual media set recalled by the first branch as the candidate visual media set.
[0245] To solve the above problems, in an optional implementation manner, the semantic proportion related to visual content in the search statement can be determined; according to the semantic proportion related to visual content in the search statement, the first extraction ratio corresponding to the visual media set recalled by the first branch and the second extraction ratio corresponding to the visual media set recalled by the second branch are determined; when the semantic proportion is greater than or equal to the preset proportion threshold, the first extraction ratio is greater than the second extraction ratio. Exemplarily, if the semantic proportion is S, the first extraction ratio is: β*S, and the second extraction ratio is: 1 - β*S, where the value of β can be set according to actual needs, and the embodiments of the present application do not make specific limitations on this.
[0246] In another alternative embodiment, the semantic proportion related to visual content in the search statement can be determined; when the semantic proportion related to visual content in the search statement is greater than or equal to a preset proportion threshold, the visual media set recalled by the first branch is used as the candidate visual media set; when the semantic proportion related to visual content in the search statement is less than the preset proportion threshold, the visual media set recalled by the second branch is used as the candidate visual media set. That the semantic proportion related to visual content in the search statement is greater than or equal to the preset proportion threshold indicates that the search statement belongs to a high-semantic search statement; that the semantic proportion related to visual content in the search statement is less than the preset proportion threshold indicates that the search statement belongs to a low-semantic search statement. The specific process will be described in detail in the embodiments below in conjunction with Figure 5 will be introduced in detail.
[0247] S415. Visual media filtering.
[0248] The search module 42 can perform semantic entity filtering, spatio-temporal filtering, and / or character relationship filtering on the candidate visual media set. The specific filtering method will be described in detail in the embodiments below in conjunction with Figure 5 will be introduced in detail.
[0249] S416. Visual media ranking.
[0250] The search module 42 ranks the candidate visual media set.
[0251] The specific ranking method will also be described in detail in the following embodiments.
[0252] S417. Return search results.
[0253] The search module 42 sends the search results obtained after ranking to the gallery service module 41.
[0254] Specifically, the search module 42 can select the top P (P≥1) visual media and their ranking information as the final search results.
[0255] S418. Display search results.
[0256] The gallery service module 41 can display the search results to the user. For example, through Figure 3A interface 305 in Figure 3A interface 308 in
[0257] such as Figure 3A interface 305 in Figure 3AThe interface 308 in it. In one example, the display order of pictures in the search results is related to the matching degree between the pictures and the search results. For example, the matching degree between the pictures displayed earlier and the search statement is greater than or equal to the matching degree between the pictures displayed later and the search statement. The calculation of the matching degree and the sorting method will be introduced in detail in the following embodiments.
[0258] Next, the search process executed by the search module 42 of the present application will be introduced in detail in conjunction with Figure 5 :
[0259] 501. Receive a search statement.
[0260] 502. Execute the search process corresponding to the search statement based on the visual semantic vector of the visual media.
[0261] The visual media set recalled by the first branch obtained by executing step 502 can be specifically referred to the search process corresponding to the first branch above.
[0262] Exemplarily, assume that the multimodal model can encode the visual media and the search statement into a k-dimensional vector space. The search statement is Q, and the corresponding sentence semantic vector is V Q ={α z}, z = 1, 2,..., Z, and a total of M picture sets R are recalled, and their vector representations are:
[0263]
[0264] Among them, the i-th row corresponds to the visual semantic vector of the i-th picture, and the value range of i is [1, Mm].
[0265] 503. Semantic entity recognition.
[0266] For the specific process of semantic entity recognition of the search statement, reference can be made to the corresponding content in the above embodiments, and details will not be repeated here.
[0267] 504. Perform a search based on the attributes of the visual media.
[0268] Specifically, match the semantic entity irrelevant to the time content in the search statement with the attributes of the visual media to obtain the visual media set recalled by the second branch. For details, reference can be made to the search process corresponding to the second branch above.
[0269] 505. Rewrite the search statement.
[0270] Specifically, delete the semantic entity irrelevant to the visual content in the search statement to obtain the rewritten search statement. The semantic entity irrelevant to the visual content specifically refers to: the semantic entity related to time and the semantic entity related to location.
[0271] Exemplary: For the search statement "sky photographed this year", where "today" is the semantic entity related to time, the rewritten search statement is "photographed sky".
[0272] In practical applications, after deleting the semantic entities unrelated to visual content, there may be some redundant stop words. For example, for the search statement "sky photographed in Beijing this year", where "this year" is the semantic entity related to time and "Beijing" is the semantic entity related to location, after deleting "this year" and "Beijing", the stop word "in" becomes a redundant word and thus also needs to be deleted. Specifically, delete the semantic entities unrelated to visual content in the search statement and their related stop words to obtain the rewritten search statement. Exemplarily, the rewritten search statement corresponding to the search statement "sky photographed in Beijing this year" is "photographed sky".
[0273] It should be noted that there is no order restriction in the execution of the above steps 502, 503, and 505. In an optional example, to improve efficiency, these three steps can be executed simultaneously.
[0274] 506. Perform the search process corresponding to the rewritten search statement based on the visual semantic vector of the visual media.
[0275] Performing step 506 obtains the set of visual media recalled by the third branch.
[0276] Specifically, obtain the visual semantic vectors of each visual media among the multiple visual media stored in the mobile phone; calculate the vector similarity between the sentence semantic vector of the rewritten search statement (i.e., the second sentence semantic vector) and the visual semantic vectors of each visual media; determine, based on the vector similarity, multiple visual media (i.e., reference visual media) whose visual semantic vectors match the sentence semantic vector from the multiple visual media stored in the mobile phone; these multiple visual media can be used as the set of visual media recalled by the third branch.
[0277] Among them, the specific implementation process of the step "determine, based on the vector similarity, multiple visual media whose visual semantic vectors match the sentence semantic vector from the multiple visual media stored in the mobile phone" can refer to the corresponding content in the above embodiments and will not be elaborated here.
[0278] Exemplarily, the rewritten search statement is Q′, and the corresponding vector V Q′ ={α′ z}, z = 1, 2,..., Z, and a total of M′ picture sets R′ are recalled, and its vector representation is:
[0279]
[0280] Among them, the i-th row corresponds to the visual semantic vector of the i-th picture, and the value range of i is [1, M′].
[0281] 507. Calculate the semantic proportion related to visual content in the search statement.
[0282] The following will introduce a method for determining the semantic proportion related to visual content:
[0283] 5071. Determine the representative visual semantic vector V I (i.e., the first representative visual semantic vector) corresponding to the visual media set recalled by the first branch and the representative visual semantic vector V I′ (i.e., the second representative visual semantic vector) corresponding to the visual media set recalled by the third branch.
[0284] 5072. Determine the sentence semantic vector V Q and the difference from the representative visual semantic vector V I to obtain a difference vector (V Q - V I )(i.e., the first difference vector).
[0285] 5073. Determine the sentence semantic vector V Q′ and the difference from the representative visual semantic vector V I′ to obtain a difference vector (V Q′ - V I′ )(i.e., the second difference vector).
[0286] 5074. Determine the semantic proportion related to visual content in the search statement according to the vector similarity between the difference vector (V Q - V I ) and the difference vector (V Q′ - V I′ ).
[0287] Among them, the semantic proportion related to visual content is positively correlated with this vector similarity.
[0288] In the above 5071, in one example, the vector similarity between the visual semantic vector of each visual media in the visual media set recalled by the first branch and the sentence semantic vector of the search statement can be obtained; the visual media in the visual media set recalled by the first branch are sorted according to the vector similarity from high to low; the average vector of the visual semantic vectors of the top T (T≥1) visual media in the sorting is used as the representative visual semantic vector V I .
[0289] Continuing with the above example, the representative visual semantic vector V I is:
[0290]
[0291] The vector similarity between each visual semantic vector of the visual media recalled by the third branch and the sentence semantic vector of the rewritten search statement can be obtained; the visual media in the visual media set recalled by the third branch are sorted according to the vector similarity from high to low; the average vector of the visual semantic vectors of the top H (H≥1) visual media in the sorting is used as the representative visual semantic vector V I′ 。
[0292] Continuing with the above example: the representative visual semantic vector V I′ is:
[0293]
[0294] The values of H and T above can be the same or different, and the embodiments of the present application do not make specific limitations on this. In another example, according to the vector similarity between the visual semantic vector of each visual media in the visual media set recalled by the first branch and the sentence semantic vector of the search statement, the visual media in the visual media set recalled by the first branch can be clustered using a clustering algorithm to obtain the clustering center point; the visual semantic vector of the clustering center point is used as the first representative visual semantic vector.
[0295] According to the vector similarity between the visual semantic vector of each visual media in the visual media set recalled by the third branch and the second sentence semantic vector of the rewritten search statement, the visual media in the visual media set recalled by the third branch can be clustered using a clustering algorithm to obtain the clustering center point; the visual semantic vector of the clustering center point is used as the second representative visual semantic vector.
[0296] In the embodiments of the present application, the representative visual semantic vector is used to represent the visual semantics of the entire set. The specific clustering algorithm can be selected according to actual needs, and the embodiments of the present application do not make any limitations on this.
[0297] In the above 5074, in an optional embodiment, the vector similarity between the difference vector (V Q -V i ) and the difference vector (V Q′ -V i′ ) can be directly used as the semantic proportion related to visual content in the search statement.
[0298] Among them, the larger the vector similarity between the difference vector (V Q -V i ) and the difference vector (V Q′ -V I′ ), the greater the semantic proportion of the visual content in the search statement; conversely, the smaller the semantic proportion of the visual content in the search statement.
[0299] It should be noted that in the embodiments of the present application, the calculation method of the semantic proportion related to visual content in the search statement draws on the word analogy feature of the distribution representation vector, that is, the additivity of word meanings is directly reflected in the additivity of the distribution representation vector.
[0300] When the semantic proportion related to visual content in the search statement is less than the preset proportion threshold, step 508 is executed subsequently.
[0301] When the semantic proportion related to visual content in the search statement is greater than or equal to the preset proportion threshold, steps 509 and 510 are executed subsequently.
[0302] 508. Determine the visual media recalled by the second branch as the candidate set.
[0303] 509. Determine the visual media recalled by the first branch as the candidate set.
[0304] 510. Visual media filtering.
[0305] To improve the search accuracy, one or more of the following processes can also be performed on the visual media recalled by the first branch: semantic entity filtering, spatio-temporal filtering, person name filtering, and person relationship filtering. Among them, semantic entity filtering refers to filtering the visual media recalled by the first branch based on the matching degree between the visual media recalled by the first branch and the semantic entity related to visual content in the search statement. Spatio-temporal filtering includes: time filtering and space (i.e., location) filtering. Person name filtering refers to filtering the visual media recalled by the first branch based on the matching degree between the person name attribute of the visual media recalled by the first branch and the semantic entity related to the person name in the search statement. Person relationship filtering refers to filtering the visual media recalled by the first branch based on the matching degree between the person relationship attribute of the visual media recalled by the first branch and the semantic entity related to the person relationship in the search statement. In practical applications, in addition to the above attributes such as the collection location, collection time, and visual media name stored in the mobile phone, the visual media may also include person name attributes and person relationship attributes manually input by the user. Therefore, in practical applications, named entity recognition technology can also be used to match the keywords in the search statement with a variety of preset person relationships to determine whether the keyword belongs to the semantic entity related to the person relationship.
[0306] When multiple processes such as semantic entity filtering, spatio-temporal filtering, person name filtering, and person relationship filtering need to be performed on the candidate set recalled by the first branch, the execution order of these multiple processes can be set according to actual needs, and the embodiments of the present application do not make specific limitations in this regard.
[0307] Exemplarily, as Figure 5 shown, the visual media filtering includes:
[0308] 5101. Semantic entity filtering.
[0309] 5102. Spatiotemporal filtering.
[0310] 5103. Character relationship filtering.
[0311] 5104. Person name filtering.
[0312] In the above 5101, generally, when a user searches, the images they hope to retrieve contain some visual content described in their search query, such as: sky, puppy, child, etc. That is to say, the semantic entities related to visual content in the search query represent the visual focus that the user is concerned about during the search.
[0313] However, when the first branch retrieves, it considers the matching degree between the visual content of visual media such as images and videos and the entire search query of the user, without fully considering the role of certain specific semantic entities (i.e., the semantic entities related to visual content) in the search query in the mobile phone gallery search scenario, resulting in the retrieval of some inaccurate photos.
[0314] For example: when a user searches for "photos of the Great Wall taken on National Day", the semantic entity related to visual content among them contains "the Great Wall". When the first branch makes a match, it considers the matching degree between the entire search query and the visual content of the image, which may lead to the first branch retrieving images that contain the "taking" behavior but do not contain the "Great Wall". This is because the image contains the "taking" behavior and the search query contains the word "taking", that is to say, there is a certain similarity between the visual semantic vector of the image and the sentence semantic vector of the search query, so the image has the possibility of being retrieved.
[0315] Another example: when a user searches for "last year when the child had a birthday holding a cake", the semantic entities related to visual content among them contain "child", "birthday", and "cake". When the first branch makes a match, it considers the matching degree between the entire search query and the image, which may lead to the model retrieving images of an adult having a birthday holding a cake. This is because the similarity between the visual semantic vector of the image of an adult having a birthday holding a cake and the sentence semantic vector of "last year when the child had a birthday holding a cake" is very high.
[0316] Therefore, in order to further improve the accuracy of searching for visual media on mobile phones, on the basis of the first branch retrieving visual media, the "semantic entity" can be strengthened, that is, the results retrieved by the first branch are fine-tuned by means of the matching degree between the visual media and the "semantic entity".
[0317] Specifically, the "semantic entity filtering" in the above 5101 may include the following steps:
[0318] 5101a. Determine the dimension to be matched according to the semantic entity related to the visual content in the search statement.
[0319] 5101b. Filter the M candidate visual media according to the matching degree of the visual content of the M candidate visual media in the dimension to be matched.
[0320] In the above 5101a, in one example, the semantic entity related to the visual content is used as the dimension to be matched. When there are multiple semantic entities related to the visual content in the search statement, these multiple semantic entities are respectively used as different dimensions to be matched to obtain multiple dimensions to be matched. Among them, the number of multiple dimensions to be matched is the same as the number of semantic entities related to the visual content.
[0321] Exemplarily, for the search statement "Last year, the child held a cake on his / her birthday", the semantic entities related to the visual content are "child", "birthday", and "cake". Then, the corresponding three dimensions to be matched are: "child", "birthday", and "cake".
[0322] In practical applications, in addition to using the semantic entity as the dimension to be matched, the search statement itself can also be used as the dimension to be matched. In this way, when filtering the semantic entity, the matching situation between the visual content of the candidate visual media and the entire search statement can also be considered to improve the rationality of the filtering. Specifically, the semantic entity related to the visual content and the search statement can be used as different dimensions to be matched to obtain multiple dimensions to be matched. When there are multiple semantic entities related to the visual content in the search statement, these multiple semantic entities and the search statement are respectively used as different dimensions to be matched. Among them, the number of multiple dimensions to be matched is one more than the number of semantic entities related to the visual content.
[0323] Exemplarily, for the search statement "Last year, the child held a cake on his / her birthday", the semantic entities related to the visual content are "child", "birthday", and "cake". Then, the corresponding four dimensions to be matched are: "child", "birthday", "cake", and "Last year, the child held a cake on his / her birthday".
[0324] In the above 5101b, the matching degree of the visual content of the candidate visual media in the dimension to be matched refers to the matching degree between the visual content of the candidate visual media and the dimension to be matched. The vector similarity between the visual semantic vector of the candidate visual media and the semantic vector to be matched in the dimension to be matched can be determined; among them, when the dimension to be matched is the semantic subject related to the visual content, the semantic vector to be matched in this dimension to be matched is the subject semantic vector of this semantic subject; when the dimension to be matched is a search statement, the semantic vector to be matched in this dimension to be matched is the sentence semantic vector of this search statement; according to this vector similarity, the matching degree of the visual content of the candidate visual media in the dimension to be matched is determined. Among them, the matching degree is positively correlated with the vector similarity.
[0325] When the number of dimensions to be matched is one, the candidate visual media with a matching degree less than or equal to the preset matching degree threshold can be filtered out according to the matching degrees of the visual contents of multiple candidate visual media in the dimension to be matched.
[0326] For example: the search statement is "photos of the sky taken on National Day", and the only semantic subject related to the visual content is "sky"; then, the recalled pictures that contain the "taking" behavior but do not contain "sky" have a relatively low matching degree with "sky" and will be filtered out.
[0327] When the number of dimensions to be matched is multiple, for each candidate visual media, the comprehensive matching degree of the candidate visual media is determined according to the matching degrees of the visual content of the candidate visual media in multiple dimensions to be matched; according to the comprehensive matching degrees of each candidate visual media, multiple candidate visual media are filtered. Specifically, from M candidate visual media, multiple first visual media with a comprehensive matching degree greater than or equal to the preset matching degree threshold (that is, meeting the preset requirements) are determined, which is equivalent to filtering out the candidate visual media with a comprehensive matching degree less than the preset matching degree threshold; or, the M candidate visual media are sorted in descending order of the comprehensive matching degree, and the top Z′ (Z′≥1) candidate visual media (that is, meeting the preset requirements) are used as multiple first visual media, which is equivalent to filtering out the (M - Z′) candidate visual media at the back.
[0328] In an optional implementation manner, any one of the following three methods can be used to determine the comprehensive matching degree of the candidate visual media:
[0329] Method 1: Sum the matching degrees of the visual content of the candidate visual media in multiple dimensions to be matched to obtain the comprehensive matching degree of the candidate visual media.
[0330] Method 2: Perform a weighted sum of the matching degrees of the visual content of the candidate visual media in multiple dimensions to be matched to obtain the comprehensive matching degree of the candidate visual media.
[0331] Among them, the weights of multiple dimensions to be matched can be configured in advance by the user.
[0332] Method 3: Using a machine learning model, determine the comprehensive matching degree of candidate visual media according to the matching degrees of the visual content of candidate visual media on multiple dimensions to be matched.
[0333] Among them, the machine learning model needs to be trained based on a dataset, and the purpose of its training is essentially to learn the weights of each dimension to be matched.
[0334] In the above Method 1, the contribution degrees of the matching degrees on different dimensions to the comprehensive matching degree are not distinguished, which may lead to the numerical values of the comprehensive matching degrees of multiple candidate visual media calculated finally being relatively close or equal, and then it is impossible to screen multiple candidate visual media.
[0335] Exemplarily, assume that there are Picture 1, Picture 2 and Picture 3 among multiple candidate visual media, and there are Dimension A, Dimension B and Dimension C among multiple dimensions to be matched. Calculate the matching degrees of the visual content of each picture on each dimension respectively, and the results are shown in Table 1.
[0336] Table 1:
[0337] Dimension A Dimension B Dimension C Picture 1 0.30 0.35 0.55 Picture 2 0.40 0.38 0.42 Picture 3 0.38 0.37 0.45
[0338] If calculated according to Method 1, the comprehensive matching scores of Picture 1, Picture 2 and Picture 3 are all 1.2, which will lead to the inability to screen the three pictures.
[0339] In the above Method 2, when there are too many semantic entities related to visual content to be concerned about (that is, the number of preset multiple tags is too large), it is difficult to accurately configure the weights of different dimensions.
[0340] In the above Method 3, the construction of the dataset is inseparable from user data. However, user data belongs to user privacy content, and users do not want their data to be reported to the cloud side.
[0341] In an optional implementation manner, to solve the above problems, the following steps can be adopted to determine the comprehensive matching degree:
[0342] S51: Determine the weights of each of the N dimensions to be matched.
[0343] Among them, the weight of the j-th dimension to be matched is positively correlated with the variation degree of the matching degree of the visual content of the M candidate visual media files on the j-th dimension to be matched; j is an integer, and the value of j ranges from 1 to N in sequence;
[0344] S52. According to the weights of the N dimensions to be matched, perform a weighted sum of the degrees of match of the visual content of the i-th candidate visual media file on each dimension to be matched among the N dimensions to be matched, so as to obtain the comprehensive degree of match of the i-th candidate visual media file.
[0345] Wherein, i is an integer, and the value of i sequentially ranges from 1 to M.
[0346] Taking the search statement "last year the child had a birthday holding a cake" as an example, the semantic entities related to the visual content include: child, birthday, cake. If multiple recalled photos all contain a cake, then the degree of variation of the degrees of match of the visual content of the multiple recalled photos on the dimension of "cake" will be relatively small, and the weight corresponding to the dimension of "cake" will be relatively small; if some of the multiple recalled photos contain a child and some do not, then the degree of variation of the degrees of match of the visual content of the multiple recalled photos on the dimension of "child" will be relatively large, and the weight corresponding to the dimension of "child" will be relatively large.
[0347] In one example, for each dimension to be matched, according to the information entropy of the degrees of match of the visual content of multiple candidate visual media on this dimension to be matched, determine the degree of variation of the degrees of match of the visual content of multiple candidate visual media on this dimension to be matched; wherein, the degree of variation is inversely proportional to the information entropy. The calculation method of information entropy will be introduced in detail in the following embodiments.
[0348] In this embodiment, the information entropy is used to measure the degree of variation of the degrees of match under each dimension to be matched, and based on this, determine the weight corresponding to this dimension to be matched.
[0349] In order to ensure that the degrees of match of the visual content of candidate visual media on different dimensions to be matched have a unified dimension, a normalization processing step can be performed. Specifically, for each candidate visual media, according to the vector similarity between the visual semantic vector of the candidate visual media and the semantic vector to be matched of the dimension to be matched, determine the initial degree of match of the candidate visual media on the dimension to be matched. Exemplarily, the vector similarity between the visual semantic vector of the candidate visual media and the semantic vector to be matched of the dimension to be matched can be used as the initial degree of match of the candidate visual media on the dimension to be matched. Perform normalization processing on the initial degrees of match of the visual content of multiple candidate visual media on the dimension to be matched to obtain the degrees of match of the visual content of multiple candidate visual media on this dimension to be matched.
[0350] The normalization process and the calculation processes of the weights and the comprehensive degree of match will be introduced in detail below:
[0351] Assume there are m candidate visual media and n dimensions to be matched. The initial matching degrees of the visual content of the m candidate visual media on each dimension to be matched among the n dimensions to be matched can be regarded as a data matrix:
[0352] X = (x ij ) m*n , i = 1, 2, …, m; j = 1, 2, …, n (5)
[0353] where x ij is the initial matching degree of the visual content of the i-th candidate visual media on the j-th dimension to be matched.
[0354] Continuing with the above example, there are multiple candidate visual media including Picture 1, Picture 2, and Picture 3, and multiple dimensions to be matched including Dimension A, Dimension B, and Dimension C, that is: m is 3 and n is 3.
[0355] Step 1: Perform normalization processing on the above data matrix.
[0356] The normalized matrix is:
[0357] R = (r ij ) m*n , i = 1, 2, …, m; j = 1, 2, …, n (6)
[0358] where
[0359]
[0360] where max(x j ) refers to the maximum value among the initial matching degrees of the visual content of multiple candidate visual media on the j-th dimension to be matched; min(x j ) refers to the minimum value among the initial matching degrees of the visual content of multiple candidate visual media on the j-th dimension to be matched.
[0361] In practical applications, other normalization methods can also be used, and the embodiments of this application do not make specific limitations on this.
[0362] Exemplarily, the results obtained by normalizing the example data in Table 1 above are shown in Table 2.
[0363] Table 2:
[0364] Dimension A Dimension B Dimension C Picture 1 0.00 0.00 1.00 Picture 2 1.00 1.00 0.00 Picture 3 0.80 0.67 0.23
[0365] Step 2: Calculate the information entropy corresponding to each dimension to be matched.
[0366] Formulas (8) and (9) can be used for calculation:
[0367]
[0368] Among them,
[0369]
[0370] Among them, e j refers to the information entropy corresponding to the j-th dimension to be matched.
[0371] Among them, since the domain of the ln(x) function is x > 0. In actual calculations, to avoid the situation where p ij in ln(p ij ) takes 0, ln(p ij ) in formula (8) can be replaced by ln(p ij +α), where α << 0.001.
[0372] The smaller the information entropy corresponding to the dimension to be matched, the greater the degree of variation in the matching degree of the visual content of multiple candidate visual media in this dimension to be matched, and the greater the amount of information provided. It can be considered that the role played by this dimension to be matched in the comprehensive evaluation is also greater.
[0373] Exemplarily, for the example data in Table 2 above, the information entropy calculated according to the above formula (8) and formula (9) is shown in Table 3:
[0374] Table 3:
[0375] Dimension A Dimension B Dimension C Information Entropy 0.69 0.67 0.48
[0376] Step 3: Calculate the weights corresponding to each dimension to be matched.
[0377] The formula (10) can be used to calculate the weights corresponding to each dimension to be matched:
[0378]
[0379] Among them, d j refers to the weight corresponding to the j-th dimension to be matched.
[0380] It can be seen that the above formula (10) is a monotonically decreasing function of information entropy. In this embodiment, the smaller the information entropy, the greater the degree of variation; the greater the degree of variation, the greater the weight. That is to say, the weight is negatively correlated with the information entropy.
[0381] To ensure that the sum of the weights corresponding to multiple dimensions to be matched is 1, the formula (11) can be used for the following calculation to obtain the final weight w j :
[0382]
[0383] Among them, w jRefers to the final weight corresponding to the j-th dimension to be matched.
[0384] In an alternative embodiment, other monotonically decreasing functions may also be used to calculate the above weights, and the embodiments of the present application do not make specific limitations thereto.
[0385] Exemplarily, for the example data in Table 3 above, the weights of different dimensions to be matched can be calculated according to the above formulas (10) and (11), as shown in Table 4:
[0386] Table 4:
[0387] Dimension A Dimension B Dimension C Weight 0.28 0.29 0.42
[0388] It should be noted that since rounding is introduced in the process of calculating the information entropy, the sum of the three dimensions in Table 4 above is not 1.
[0389] Step 4: Weighted summation.
[0390] Perform a weighted summation on the normalized matching degrees of any candidate visual media to obtain the comprehensive matching degree of the candidate visual media, which can be specifically calculated using formula (12):
[0391]
[0392] where s i Refers to the comprehensive matching degree of the i-th visual media.
[0393] Exemplarily, for the example data in Table 2 and Table 4 above, the comprehensive matching degrees of different pictures are calculated using the above formula (12), as shown in Table 5:
[0394] Table 5:
[0395] Picture 1 Picture 2 Picture 3 Comprehensive Matching Degree 0.42 0.58 0.52
[0396] In the above 5102, when the first branch recalls visual media, it only considers the visual semantic information of the visual media and does not consider the attribute information such as the time and location of the visual media. Therefore, it is necessary to perform time filtering, location filtering, etc. on the visual media recalled by the first branch. Specifically, when the user performs semantic search, if the search statement contains time information, pictures that do not meet the time limit need to be filtered out.
[0397] However, the user's description form of time information is rich and diverse and is fuzzy. For example, when the user searches for "the sky photographed in the afternoon", the "afternoon" in the user's search statement has no clear and standardized definition. When the user uses a fuzzy time expression in the search statement, the time window used for time filtering will affect the user experience.
[0398] Exemplary: The user starts traveling to other places on September 30, 2022 and arrives home on October 9, 2022. During this period, the user takes a lot of photos. One day in 2023, the user wants to view the beautiful scenery taken during this trip. Then the user is very likely to enter the search statement "Scenery taken during last year's National Day holiday". If directly based on "last year's National Day" in the search statement, the time window used for time filtering is set to "October 1, 2022 to October 7, 2022", then the scenic photos taken by the user on September 30, 2022, October 8, 2022, and October 9, 2022 will be filtered out, which obviously does not meet the user's expectations. To improve the rationality of time filtering, the embodiments of the present application provide a new time filtering method. Specifically, using a clustering algorithm, multiple candidate visual media are clustered according to the acquisition time of each of the multiple candidate visual media to obtain K (K≥1) clustering clusters.
[0399] In this way, pictures of the same series with relatively close acquisition times can be grouped into the same clustering cluster.
[0400] Continuing with the above example, the natural scenery picture taken by the user on September 30, 2022 and the natural scenery picture taken by the user on October 1, 2022 have relatively close shooting times and are grouped into the same clustering cluster through the above clustering algorithm.
[0401] The above clustering algorithm may include but is not limited to: K-Means clustering algorithm, Mean shift clustering algorithm, and density-based clustering algorithm.
[0402] Taking the density-based clustering algorithm as an example, in the scenario of semantic search including time information, it is unreasonable to set fixed first parameter ∈ and second parameter MinPts. For example, when the user searches for "Photos taken during the outing in 2022", the time range is 1 whole year; while when the user searches for "Photos taken during the outing in the morning", the time range is several hours. The same first parameter ∈ and second parameter MinPts should not be set in these two cases. To improve the rationality of clustering, the following steps can be used to determine the first parameter ∈ and the second parameter MinPts:
[0403] 51021. Determine the first parameter involved in the density-based clustering algorithm according to the time search range included in the search statement.
[0404] Among them, the first parameter is positively correlated with the duration corresponding to the time search range.
[0405] The time search range can be determined according to the time-related semantic entities in the search statement. Specifically, the time range corresponding to the time-related semantic entity (i.e., the time search range) can be returned through a mapping table. The mapping table can be constructed in advance as needed, and the specific form is not specifically limited in the embodiments of the present application.
[0406] Exemplarily, the time-related semantic entity extracted from the search statement "photos taken in spring" is "spring". Querying the mapping table, the corresponding time range is obtained as: February 1 - May 30; the time-related semantic entity extracted from the search statement "the sky taken in the morning" is "morning". Querying the mapping table, the corresponding event range is obtained as: 7:00 - 12:00; the time-related semantic entity extracted from the search statement "photos of going out for fun during National Day" is "National Day". Querying the mapping table, the corresponding time range is obtained as: October 1 - October 7; the time-related entity extracted from the search statement "the sky taken in Beijing this year" is "this year". Querying the mapping table, the corresponding time range is obtained as "January 1, 2023 to December 31, 2023".
[0407] The first parameter can be determined according to the duration corresponding to the time search range. Exemplarily, the start time of the time search range (which can be understood as the start timestamp) is T start and the end time (which can be understood as the end timestamp) is T end , and the duration of the time search range is: T end -T start , and the following formula can be used to calculate the first parameter:
[0408] ∈=α*(T end -T start ) (13)
[0409] where α is a coefficient that can be adjusted manually, and its value can be set according to actual needs, which is not specifically limited in the present application.
[0410] 51022. Determine the second parameter involved in the density-based clustering algorithm according to the ratio of the number of multiple candidate visual media to the time search range.
[0411] The following formula can be used to calculate the second parameter:
[0412]
[0413] where N is the total number of visual media in the candidate set whose acquisition time is within the time search range, N≥1; where n is the number of times the time search range repeats between T 1 and T 2 . T 1is the acquisition time of the earliest acquired visual media among multiple visual media stored via the mobile phone; T 2 is the acquisition time of the latest acquired visual media among multiple visual media stored via the mobile phone. Exemplarily, the search statement is "the sky photographed during the National Day", and its time search range is from October 1st to October 7th, with a duration of 7 days; the acquisition time of the earliest acquired visual media stored in the mobile phone is August 1st, 2020; the acquisition time of the latest acquired visual media stored in the mobile phone is October 20th, 2023; then, from August 1st, 2020 to October 20th, 2023, the time search range from October 1st to October 7th repeats 4 times (i.e., once a year).
[0414] where, (T end - T start ) * n can be understood as the total duration corresponding to the time search range.
[0415] where, β is a coefficient that can be adjusted manually, and its magnitude can be set according to actual needs. This application does not make specific limitations on this.
[0416] In this embodiment, according to the duration defined by the time search range included in the search statement, the first parameter ∈ and the second parameter MinPts involved in the density-based clustering algorithm are dynamically adjusted, so as to ensure the rationality of clustering and further improve the accuracy of the final search result.
[0417] When the search statement includes a time search range, determine the acquisition time range corresponding to each of the K (K≥1) clustering clusters; filter out the clustering clusters whose acquisition time range does not overlap with the time search range, and retain the clustering clusters whose acquisition time range overlaps with the time search range.
[0418] The overlap between the acquisition time range and the time search range can be partial overlap or full overlap. Whether it is partial overlap or full overlap between the two, it belongs to having an overlapping part.
[0419] In this way, G (G≥1) clustering clusters are selected from the K clustering clusters. The acquisition time of the earliest acquired visual media among these G clustering clusters is before the start time of the time search range, and / or, the acquisition time of the latest acquired visual media among these G clustering clusters is after the end time of the time search range. Note: There is no intersection between the G clustering clusters, and there is also no intersection between the acquisition time ranges of the G clustering clusters themselves.
[0420] The acquisition time range corresponding to a clustering cluster is from the acquisition time T min T1 of the earliest acquired visual media in this clustering cluster to the acquisition time T max of the latest acquired visual media in this clustering cluster, that is: [Tmin , T max .
[0421] Specifically, the time search range is from T start , T end , and the acquisition time range corresponding to the clustering cluster is from T min , T max . When [T min , T max and [T start , T end have an overlapping part, retain this clustering cluster; when [T min , T max and [T start , T end do not have an overlapping part, filter out this clustering cluster. For example: the time search range is from October 1, 2022 to October 7, 2022, and the acquisition time range of the clustering cluster is from September 30, 2022 to October 1, 2022. These two ranges have an overlapping part (i.e., October 1, 2022), and this clustering cluster is retained.
[0422] Continuing with the above example, the natural scenery pictures taken by the user on September 30, 2022 and the natural scenery pictures taken by the user on October 1, 2022 are close in shooting time. Through the above clustering algorithm, they are grouped into the same clustering cluster. Assuming this clustering cluster contains only these two pictures, then the acquisition time range of this clustering cluster is: from September 30, 2022 to October 1, 2022, and it has an overlapping part with the time search range: from October 1, 2022 to October 7, 2022. Therefore, this clustering cluster will be retained, that is to say, the natural scenery pictures taken by the user on September 30, 2022 will not be filtered out.
[0423] Similarly, the natural scenery pictures taken by the user on October 8, 2022 and the natural scenery pictures taken by the user on October 7, 2022 are close in shooting time. Through the above clustering algorithm, they are grouped into the same clustering cluster. Assuming this clustering cluster contains only these two pictures, then the acquisition time range of this clustering cluster is the pictures from October 7, 2022 to October 8, 2022, and it has an overlapping part with the time search range: from October 1, 2022 to October 7, 2022. Therefore, this clustering cluster will be retained. That is to say, the natural scenery pictures taken by the user on October 8, 2022 will not be filtered out.
[0424] In this way, for the search statement "scenery taken during last year's National Day holiday", the acquisition time of the earliest acquired visual media among the multiple clustering clusters obtained by screening is September 30, 2022, and the acquisition time of the latest acquired visual media is October 8, 2022.
[0425] It can be seen that adopting the time filtering method provided by the embodiments of the present application can ensure that a series of photos with relatively close acquisition times are displayed to the user, guarantee the coherence of the search results, and improve the user's search experience.
[0426] In addition, when the search statement also includes a semantic entity related to a location, location filtering can be further performed on the G clustered clusters selected. Specifically, visual media in the clustered clusters with acquisition locations that do not match the semantic entity related to the location in the search statement can be filtered out. Exemplarily, the acquisition location "Prince Kung's Mansion, Xicheng District, Beijing" and the semantic entity "Sujiatuo Town, Haidian District, Beijing" can be considered to match (belonging to municipal-level matching); the acquisition location "Tsinghua University, Haidian District, Beijing" and the semantic entity "Sujiatuo Town, Haidian District, Beijing" can be considered to match (belonging to district-level matching); the acquisition location "Shanghai" and the semantic entity "Beijing" can be considered not to match.
[0427] In the above embodiments, clustering is performed first, then time filtering, and finally location filtering. Of course, in practical applications, location filtering can also be performed first, then clustering, and finally time filtering. The specific execution order of these three steps can be set according to actual needs, and the embodiments of the present application do not make specific limitations on this.
[0428] 5103. Filtering of the relationship between people.
[0429] When the search statement includes a semantic entity related to the relationship between people, visual media in each clustered cluster with a people relationship attribute that does not match the semantic entity are filtered out. Exemplarily, if the search statement includes "good friends" and the name attribute of Picture B is "colleague", and these two do not match, then Picture B is filtered out.
[0430] 5104. Filtering of names.
[0431] When the search statement includes a semantic entity related to a name, visual media in each clustered cluster with a name attribute that does not match the semantic entity are filtered out. Exemplarily, if the search statement includes "Zhang San" and the name attribute of Picture A is "Li Si", and these two do not match, then Picture A is filtered out.
[0432] It should be added that the semantic entity related to the name and the semantic entity related to the relationship between people in the search statement can also be identified through named entity recognition technology. The name attribute and the people relationship attribute of the visual media are manually added by the user for the visual media in advance.
[0433] 511. Sorting.
[0434] For the candidate set obtained in the above step 508, the visual media in the candidate set can be sorted according to the acquisition time of the visual media in the candidate set from earliest to latest, so as to obtain the display order of the visual media. The acquisition time of the visual media with a higher display order is earlier than that of the visual media with a lower display order. Subsequently, the mobile phone can display the candidate set according to the display order of the visual media in the candidate set.
[0435] For the filtered candidate set (including the above G clustering clusters) obtained in the above step 510, one of the following methods can be used for sorting:
[0436] Method 1: Sort the visual media in the G clustering clusters from highest to lowest according to the target matching degree of the visual media in the G clustering clusters, so as to obtain the display order of the visual media in the G clustering clusters. The target matching degree can be the matching degree between the visual content of the visual media and the search statement in the above text or the comprehensive matching degree in the above text (the specific calculation method can refer to the corresponding content in the above embodiments). Subsequently, the mobile phone can display the visual media in the G clustering clusters according to the display order of the visual media in the G clustering clusters.
[0437] Method 2: Sort the G clustering clusters according to the start time of the acquisition time range of the G clustering clusters from earliest to latest, so as to obtain the display order of the G clustering clusters (that is, the inter-cluster sorting). Exemplarily, as Figure 6 shown in the interface 601, the acquisition time range corresponding to the clustering cluster A is from September 30, 2022 to October 1, 2022; the acquisition time range corresponding to the clustering cluster B is from October 3, 2022 to October 5, 2022; then, the display order of the clustering cluster A is prior to the display order of the clustering cluster B. For each clustering cluster, sort the visual media within the cluster from highest to lowest according to the target matching degree of the visual media within the cluster, so as to obtain the display order between the visual media within the cluster (intra-cluster sorting); or, for each clustering cluster, sort the visual media within the cluster according to the acquisition time of the visual media within the cluster from earliest to latest, so as to obtain the display order of the visual media within the cluster. Subsequently, the mobile phone displays the visual media in the G clustering clusters according to the inter-cluster sorting and the intra-cluster sorting.
[0438] It should be added that, in order to better protect the user privacy and security, meet the principle of minimizing user data, and avoid reporting user data to the cloud side as much as possible, the above entire search process is completed on the terminal side.
[0439] In addition, the present application provides an electronic device, including: a memory, a processor, and a display, wherein the memory is used to store a program; the processor is coupled to the memory and the display, and is used to execute the program stored in the memory to implement the above visual media search method.
[0440] The embodiments of the present application also provide a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a computer, one or more steps in any of the above visual media search methods can be implemented.
[0441] The computer-readable storage medium may be a non-transitory computer-readable storage medium. For example, the non-transitory computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0442] Another embodiment of the present application also provides a computer program product containing instructions. When the computer program product is executed by a computer, one or more steps in any of the above methods can be implemented.
[0443] Among them, the electronic device, the computer-readable storage medium, and the computer program product provided in this embodiment are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above, and will not be elaborated here.
[0444] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces. The indirect coupling or communication connection of the device or unit may be in an electrical, mechanical or other form.
[0445] The unit described as a separated component may or may not be physically separated. The component displayed as a unit may be a physical unit or multiple physical units, that is, it may be located in one place, or may be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0446] In addition, each functional unit in the various embodiments of the present application may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0447] When an integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods of the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0448] The above content is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A visual media search method applicable to an electronic device, characterized in that, it includes: displaying a first interface; the first interface includes a search box; receiving a search operation on a search statement input into the search box; wherein, the search statement includes a time search range and a first keyword, the start time of the time search range is a first time point, and the end time is a second time point; displaying a first search result, the first search result corresponding to a first visual media file; wherein, the first visual media file matches the first keyword; the acquisition time of the first visual media file is before the first time point or after the second time point.
2. The method according to claim 1, characterized in that, after receiving the search operation on the search statement input into the search box, it includes: displaying a second search result, the second search result corresponding to a second visual media file; the second visual media file matches the first keyword; the acquisition time of the second visual media file is between the first time point and the second time point, and the difference between the acquisition time of the first visual media file and the acquisition time of the second visual media file is less than a first threshold.
3. The method according to claim 1, characterized in that, displaying the first interface includes: responding to an operation of opening a gallery application, displaying the first interface; or, responding to an operation of opening the negative first screen triggered on the main screen of the electronic device, displaying the first interface; or, responding to a pull-down search operation triggered on the main screen of the electronic device, displaying the first interface.
4. The method according to claim 1, characterized in that, the displaying of the first search result includes: displaying a first picture, the first picture being an equal-proportion thumbnail of the first visual media file or an equal-proportion thumbnail of any image frame.
5. The method according to claim 4, characterized in that, it further includes: displaying a third picture; the third picture being an equal-proportion thumbnail of a third visual media file or an equal-proportion thumbnail of any image frame; the third visual media file matches the first keyword; the first picture and the third picture are displayed on a second interface; the first picture is displayed in front of the third picture; the matching degree of the visual media file corresponding to the first picture with the first keyword is a first matching degree; the matching degree of the visual media file corresponding to the second picture with the first keyword is a second matching degree; the first matching degree is greater than the second matching degree.
6. The method according to any one of claims 1 to 5, characterized in that, the visual content of the first visual media file matches the first keyword; the visual content is data that needs to be obtained through a natural picture understanding model.
7. The method according to any one of claims 1 to 5, characterized in that, the first keyword and the first visual media file are matched through a contrastive text-image pre-trained CLIP model.
8. A visual media search method applicable to an electronic device, characterized in that, it includes: displaying a first interface; the first interface includes a search box; Receive a search operation for a search statement input into the search box, where the search statement includes a time search range; Use a clustering algorithm to cluster the multiple first visual media files according to the acquisition times of the multiple first visual media files to obtain K clustering clusters, where the acquisition time ranges of the K clustering clusters are different; where K≥1 and is an integer; Determine a target clustering cluster according to whether the acquisition time range overlaps partially with the time search range; Determine the search result of the search statement according to the target clustering cluster; Display the search result.
9. The method according to claim 8, wherein, The start time of the acquisition time range of each clustering cluster is the acquisition time of the earliest acquired visual media file in the clustering cluster, and the end time is the acquisition time of the latest acquired visual media file in the clustering cluster; There is no intersection between the acquisition time ranges of the K clustering clusters.
10. The method according to claim 9, wherein, The start time of the time search range is the first time point, and the end time is the second time point; Determining a target clustering cluster according to whether the acquisition time range overlaps partially with the time search range includes: When the start time of the acquisition time range of the kth clustering cluster among the K clustering clusters is between the first time point and the second time point and the end time is not between the first time point and the second time point, it is determined that the acquisition time range of the kth clustering cluster overlaps partially with the time search range; When the end time of the acquisition time range of the kth clustering cluster is between the first time point and the second time point and the start time is not between the first time point and the second time point, it is determined that the acquisition time range of the kth clustering cluster overlaps partially with the time search range; When the start time of the acquisition time range of the kth clustering cluster is not between the first time point and the second time point and the end time is not between the first time point and the second time point, it is determined that the acquisition time range of the kth clustering cluster overlaps partially with the time search range; k is an integer, 1≤k≤K; Determine the clustering clusters among the K clustering clusters whose acquisition time ranges overlap partially with the time search range as the target clustering clusters.
11. The method according to claim 9, wherein, The search result includes multiple target clustering clusters; The method further includes: Determine the display order of the multiple target clustering clusters according to the order of the start times of the acquisition time ranges of the multiple target clustering clusters; Displaying the search result includes: Display the multiple target clustering clusters according to the display order of the multiple target clustering clusters.
12. The method according to any one of claims 8 to 11, wherein, further includes: Perform semantic understanding on the search statement to obtain a first sentence semantic vector; Obtain the visual semantic vectors of multiple visual media files; The visual semantic vector of each visual media file is obtained by performing semantic understanding on the image or image frame of the visual media file using a natural picture understanding model; Determine M candidate visual media files from the multiple visual media files, where the visual semantic vectors match the first sentence semantic vector; where M is an integer greater than 1; Determine the multiple first visual media files according to the M candidate visual media files.
13. The method according to claim 12, characterized in that, Using a clustering algorithm, cluster the multiple first visual media files according to the acquisition time of the multiple first visual media files to obtain K clustering clusters, including: If the semantic ratio related to visual content in the search statement is greater than or equal to a preset ratio threshold, use a clustering algorithm to cluster the multiple first visual media files according to the acquisition time of the multiple first visual media files to obtain K clustering clusters; the visual content is data that needs to be obtained through a natural picture understanding model; The method further includes: Obtain the attributes of the multiple visual media files; Determine multiple second visual media whose attributes match the semantic subject in the search statement from the multiple visual media files; If the semantic ratio related to visual content in the search statement is less than the preset ratio threshold, determine the search result of the search statement according to the multiple second visual media.
14. The method according to claim 13, characterized in that, further includes: Remove the semantic subject unrelated to visual content in the search statement to obtain a rewritten search statement; Perform semantic understanding on the rewritten search statement to obtain a second sentence semantic vector; Determine multiple reference visual media whose visual semantic vectors match the second sentence semantic vector from the multiple visual media files; Determine the first representative visual semantic vector corresponding to the multiple first visual media files and the second representative visual semantic vector corresponding to the multiple reference visual media; Determine the semantic ratio related to visual content in the search statement according to the vector similarity between the first difference vector and the second difference vector; The first difference vector represents the difference between the first sentence semantic vector and the first representative visual semantic vector; The second difference vector represents the difference between the second sentence semantic vector and the second representative visual semantic vector.
15. The method according to claim 14, characterized in that, Removing the semantic subject unrelated to visual content in the search statement to obtain a rewritten search statement includes: Removing the semantic subject related to time and / or the semantic subject related to location in the search statement to obtain a rewritten search statement.
16. The method according to claim 12, characterized in that, Determining the multiple first visual media files according to the M candidate visual media files includes: Determine N dimensions to be matched corresponding to the search statement according to the semantic subject related to visual content in the search statement; N is greater than 1 and is an integer; Obtain the weights of the N dimensions to be matched; where the weight of the jth dimension to be matched is positively correlated with the degree of variation of the matching degree of the visual content of the M candidate visual media files in the jth dimension to be matched; j is an integer, and the values of j range from 1 to N in sequence; According to the weights of the N dimensions to be matched, perform a weighted sum of the degrees of match of the visual content of the i-th candidate visual media file on each dimension to be matched among the N dimensions to be matched, to obtain the comprehensive degree of match of the i-th candidate visual media file; i is an integer, and the values of i sequentially range from 1 to M; Determine the multiple first visual media files whose comprehensive degrees of match meet the preset requirements from the M candidate visual media files.
17. The method according to claim 16, wherein, further comprising: Match the search terms in the search statement with a plurality of preset tags to determine whether the search terms belong to semantic entities related to the tags; The tags are used to describe visual content; Determine the semantic entities related to visual content in the search statement as the semantic entities related to visual content in the search statement.
18. The method according to any one of claims 8 to 11, wherein, The clustering algorithm includes: a density-based clustering algorithm; The method further comprises: According to the time search range, determine a first parameter involved in the density-based clustering algorithm; the first parameter is used to describe the neighborhood radius of data points; the first parameter is positively correlated with the duration corresponding to the time search range.
19. The method according to claim 18, wherein, further comprising: Determine the number of visual media files among the multiple first visual media files whose acquisition times are within the time search range; According to the ratio of the number of visual media files to the total duration corresponding to the time search range, determine a second parameter involved in the density-based clustering algorithm; the second parameter is used to describe the minimum number of data points in the neighborhood of data points.
20. The method according to any one of claims 8 to 11, wherein, Determining the search result of the search statement according to the target clustering cluster includes: If the search statement includes a semantic entity related to a location, filter the visual media files in the target clustering cluster according to the semantic entity related to the location and the location attribute of the visual media files in the target clustering cluster; If the search statement includes a semantic entity related to a person relationship, filter the visual media files in the target clustering cluster according to the semantic entity related to the person relationship and the person relationship attribute of the visual media files in the target clustering cluster; and / or If the search statement includes a semantic entity related to a person name, filter the visual media files in the target clustering cluster according to the semantic entity related to the person name and the person name attribute of the visual media files in the target clustering cluster.
21. An electronic device, wherein, comprising: A memory, a processor, and a display, wherein, The memory is used to store programs; The display is used to display a search page; The processor is coupled to the memory and the display, and is used to execute the program stored in the memory to implement the method according to any one of claims 1 to 20.
22. A computer-readable storage medium storing a computer program, characterized in that, when the computer program is executed by a computer, it can implement the method described in any one of claims 1 to 20.